4 modalities, one API
Multimodal in production
Vision chat, speech-to-text, and image gen behind the same key.
One gateway serves text, vision, audio, and image models: describe uploads with Kimi K3 or GLM Flash, transcribe calls with Whisper V3, generate and edit with FLUX Kontext.
The playground takes images and text files today; the API takes OpenAI-format image_url parts with per-request metering like everything else.
- ✓Vision + STT + image gen, one key
- ✓1M-ctx multimodal flagships
- ✓Playground attachments built in
- ✓Same ledger, same guardrails
Keep exploring
Production inference
Serverless tokens or warm dedicated GPUs — one gateway, OpenAI-compatible.
Fine-tuning that ships
SFT → DPO → RL, then deploy the checkpoint in seconds.
Agents that stay up
Low-latency tool loops, streaming, and spend guardrails for autonomous workloads.
Batch at half price
Evals, embeddings, re-ranking, and backfills — async, 50% off.
Enterprise RAG that cites
Embed, retrieve, rerank, answer — with cached prefixes and audit trails.