$0.10 per 1M embed tokens
Enterprise RAG that cites
Embed, retrieve, rerank, answer — with cached prefixes and audit trails.
Chunk to ~512 tokens, embed with Qwen3 Embedding 8B, pull top-20 by cosine, then rerank to top-5 with the cross-encoder. Rerank quality is where most RAG wins hide.
Answer with a Flash-class model over shared system prompts — repeat questions bill at cached rates. Every retrieval is logged, so regulated teams can show exactly what the model saw.
- ✓$0.10/1M embeddings + $0.20/1M rerank
- ✓Cached corpus prefixes up to 97% off
- ✓Request logs for compliance review
- ✓Batch backfills at 50% off
Keep exploring
Production inference
Serverless tokens or warm dedicated GPUs — one gateway, OpenAI-compatible.
Fine-tuning that ships
SFT → DPO → RL, then deploy the checkpoint in seconds.
Agents that stay up
Low-latency tool loops, streaming, and spend guardrails for autonomous workloads.
Batch at half price
Evals, embeddings, re-ranking, and backfills — async, 50% off.
Multimodal in production
Vision chat, speech-to-text, and image gen behind the same key.