← All use cases

211 tok/s peak MoE class

Production inference

Serverless tokens or warm dedicated GPUs — one gateway, OpenAI-compatible.

Start on serverless: pick any of 15+ open models, pay per token, stream over SSE. Cached prompt prefixes bill at cached rates automatically.

When traffic steadies, move the same model id to a dedicated B200/B300 deployment with autoscaling and scale-to-zero. No client changes — just a hotter endpoint behind the same URL.

Keep exploring