211 tok/s peak MoE class
Production inference
Serverless tokens or warm dedicated GPUs — one gateway, OpenAI-compatible.
Start on serverless: pick any of 15+ open models, pay per token, stream over SSE. Cached prompt prefixes bill at cached rates automatically.
When traffic steadies, move the same model id to a dedicated B200/B300 deployment with autoscaling and scale-to-zero. No client changes — just a hotter endpoint behind the same URL.
- ✓P50 cold start <800ms on B200 NVFP4
- ✓Standard / Priority / Fast tiers
- ✓Function calling + structured JSON
- ✓Logs, metrics, traces per call
Keep exploring
Fine-tuning that ships
SFT → DPO → RL, then deploy the checkpoint in seconds.
Agents that stay up
Low-latency tool loops, streaming, and spend guardrails for autonomous workloads.
Batch at half price
Evals, embeddings, re-ranking, and backfills — async, 50% off.
Enterprise RAG that cites
Embed, retrieve, rerank, answer — with cached prefixes and audit trails.
Multimodal in production
Vision chat, speech-to-text, and image gen behind the same key.