60/min beta rate envelope
Agents that stay up
Low-latency tool loops, streaming, and spend guardrails for autonomous workloads.
Agents are token furnaces: cached prefixes, batch-tolerant traffic, and bursty tool calls. Run the hot loop on Flash-class models, escalate hard reasoning to frontier weights — same key, same gateway.
Guardrails matter more than latency: hard spend limits (402s), per-key scopes, IP allowlists, and a full audit trail of everything the agent did.
- ✓Streaming + tool calls native
- ✓Cached prefixes up to 97% off
- ✓Per-key scopes + IP allowlist
- ✓MCP config + workspace skill
Keep exploring
Production inference
Serverless tokens or warm dedicated GPUs — one gateway, OpenAI-compatible.
Fine-tuning that ships
SFT → DPO → RL, then deploy the checkpoint in seconds.
Batch at half price
Evals, embeddings, re-ranking, and backfills — async, 50% off.
Enterprise RAG that cites
Embed, retrieve, rerank, answer — with cached prefixes and audit trails.
Multimodal in production
Vision chat, speech-to-text, and image gen behind the same key.