Inference.ai (operated by Distribyte Inc., doing business as Inference.ai) runs a GPU cloud and AI infrastructure platform aimed at teams that need compute for training and running models without paying hyperscaler markups. Founded in 2023 in Palo Alto, California, by CEO John Yue and CTO Michael Yu — who previously co-founded and led Toronto-based Bifrost Cloud — the company raised a $4 million seed round co-led by Maple VC.
The platform now bundles four products: Engine, bare-metal GPU access sold hourly, fractional, or reserved; Ghost, a pre-wired Linux VM with frontier models and coding agents like Claude Code ready to deploy to Discord, Telegram, or WhatsApp; Maestro, a FinOps layer that tracks AI spend by team or agent and can auto-route requests to the cheapest endpoint meeting an SLA; and Academy, a training arm for AI infrastructure skills. The pitch is a full stack from raw GPU to a working agent, priced below what most cloud providers charge for the same silicon.
Inference.ai targets AI teams and startups that need GPU capacity without negotiating directly with AWS, Azure, or Google Cloud, plus finance and platform teams trying to control runaway inference spend. Pricing for GPU access is on request rather than published — the site directs prospective customers to contact sales for rates on Engine, Ghost, and Maestro.







