← All use cases

4 modalities, one API

Multimodal in production

Vision chat, speech-to-text, and image gen behind the same key.

One gateway serves text, vision, audio, and image models: describe uploads with Kimi K3 or GLM Flash, transcribe calls with Whisper V3, generate and edit with FLUX Kontext.

The playground takes images and text files today; the API takes OpenAI-format image_url parts with per-request metering like everything else.

Keep exploring