-50% vs serverless rate
Batch at half price
Evals, embeddings, re-ranking, and backfills — async, 50% off.
Not everything needs streaming. Submit up to 100k requests per job; results land for download on completion, billed at half the serverless rate.
Same models, same gateway, same metering — batch spend shows up in the same ledger and analytics as live traffic.
- ✓50% off serverless rates
- ✓100k requests per job
- ✓Evals + embeddings + rerank
- ✓Unified ledger + analytics
Keep exploring
Production inference
Serverless tokens or warm dedicated GPUs — one gateway, OpenAI-compatible.
Fine-tuning that ships
SFT → DPO → RL, then deploy the checkpoint in seconds.
Agents that stay up
Low-latency tool loops, streaming, and spend guardrails for autonomous workloads.
Enterprise RAG that cites
Embed, retrieve, rerank, answer — with cached prefixes and audit trails.
Multimodal in production
Vision chat, speech-to-text, and image gen behind the same key.