[ public beta ] — $1 free credits on every new account
runaiicloud
ModelsGPUsPricingDocsConnectCompareEnterprisePlayground
Log inGet started

[ runaii cloud — serverless · dedicated · training ]

Own your
inference.

The open-model cloud for teams who read pricing pages: serverless tokens, dedicated Blackwell GPUs, and training — behind one API your SDK already speaks.

Start building — $1 freeRead the docs

no card · no sales call · how we benchmarked the incumbents

runaii — zshexit 0 · 312ms
curl https://api.runaii.cloud/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $RUNAII_API_KEY" \
  -d '{
    "model": "runaii/glm-5.3",
    "messages": [{"role": "user", "content": "Say hello in Spanish"}]
  }'
liveMaking requestMaking request
<0ms
P50 cold start, B200
0 tok/s
peak MoE throughput
0+
serverless models
$0
free, no card
$0.000/M
cached input · up to 97% off

gateway live·billed to the micro·status →

[ featured — glm-5.3-flash ]

Frontier speed. $0.03 cached.

The flash-class workload everyone benchmarks — we ship it with 1M context, 412 tok/s on our fleet, and cache-hit input at under a nickel. One key, every model.

Try it — $1 freeModel card →

input

$0.15

/ 1M tokens

cached

$0.030

/ 1M tokens

output

$0.50

/ 1M tokens

B — fleet speeds

Tokens per second, on our fleet

Median output speed, Standard tier. Same architecture, honest numbers.

GLM 5.3 Flash412 tok/s1M ctx$0.15/1M · ⚡0.030 cached
DeepSeek V4 Flash385 tok/s1M ctx$0.22/1M · ⚡0.022 cached
OpenAI gpt-oss-120b260 tok/s131K ctx$0.15/1M
MiniMax M3190 tok/s512K ctx$0.30/1M
Llama 3.3 70B Instruct175 tok/s131K ctx$0.23/1M
GLM-5.3118 tok/s1M ctx$1.40/1M · ⚡0.260 cached
Qwen3.8-Max104 tok/s262K ctx$2.00/1M
Kimi K396 tok/s1M ctx$3.00/1M

how we benchmarked the incumbents →

[ open models — one api, one key ]

all 15 →
GLGLM-5.3$1.40/1M· ⚡$0.260 cachedGFGLM 5.3 Flash$0.15/1M· ⚡$0.030 cachedKimi K3$3.00/1MDeepSeek V4 Pro$1.32/1M· ⚡$0.044 cachedDeepSeek V4 Flash$0.22/1M· ⚡$0.022 cachedQMQwen3.8-Max$2.00/1MQFQwen3.8 Flash$0.40/1MOpenAI gpt-oss-120b$0.15/1MOpenAI gpt-oss-20b$0.05/1MMiniMax M3$0.30/1ML3Llama 3.3 70B Instruct$0.23/1MQEQwen3 Embedding 8B$0.10/1MQRQwen3 Reranker 8B$0.20/1MWhisper V3 Large$0.25/1MFLUX.1 Kontext Pro$0.04/1MGLGLM-5.3$1.40/1M· ⚡$0.260 cachedGFGLM 5.3 Flash$0.15/1M· ⚡$0.030 cachedKimi K3$3.00/1MDeepSeek V4 Pro$1.32/1M· ⚡$0.044 cachedDeepSeek V4 Flash$0.22/1M· ⚡$0.022 cachedQMQwen3.8-Max$2.00/1MQFQwen3.8 Flash$0.40/1MOpenAI gpt-oss-120b$0.15/1MOpenAI gpt-oss-20b$0.05/1MMiniMax M3$0.30/1ML3Llama 3.3 70B Instruct$0.23/1MQEQwen3 Embedding 8B$0.10/1MQRQwen3 Reranker 8B$0.20/1MWhisper V3 Large$0.25/1MFLUX.1 Kontext Pro$0.04/1M
runaii/glm-5.3 $1.40/$4.40/1M/runaii/glm-5.3-flash $0.15/$0.50/1M/runaii/kimi-k3 $3.00/$15.00/1M/runaii/deepseek-v4-pro $1.32/$3.96/1M/runaii/qwen3.8-max $2.00/$6.00/1M/runaii/gpt-oss-120b $0.15/$0.60/1M/runaii/minimax-m3 $0.30/$1.20/1M/runaii/llama-3.3-70b $0.23/$0.85/1M/runaii/whisper-v3 $0.25/M/1M/runaii/flux-kontext $0.04/img/1M/runaii/glm-5.3 $1.40/$4.40/1M/runaii/glm-5.3-flash $0.15/$0.50/1M/runaii/kimi-k3 $3.00/$15.00/1M/runaii/deepseek-v4-pro $1.32/$3.96/1M/runaii/qwen3.8-max $2.00/$6.00/1M/runaii/gpt-oss-120b $0.15/$0.60/1M/runaii/minimax-m3 $0.30/$1.20/1M/runaii/llama-3.3-70b $0.23/$0.85/1M/runaii/whisper-v3 $0.25/M/1M/runaii/flux-kontext $0.04/img/1M/
A — 01/03

The stack

Three primitives. One control plane. No replatforming between stages.

01 / Serverless

Per-token inference. Zero capacity planning.

Fifteen open models behind one OpenAI-compatible URL. Standard, Priority, and Fast tiers; cached prefixes bill for almost nothing.

1M ctx · streaming · tools →

02 / Deployments

Your weights on B200s, billed per second.

Reserve Blackwell-class GPUs with autoscaling and scale-to-zero. Your VPC, your region pin, your control plane.

B200 · B300 · GB300 →

03 / Training

LoRA to RL, then straight to prod.

Managed SFT/DPO per 1M tokens, dedicated RL per GPU-second. Every checkpoint deploys to inference in seconds.

SFT · DPO · RL →

[ metered in real time ]

Watch every token land.

Per-request receipts: tokens in/out, cache hit rate, latency, and the exact micros charged — streamed to your console the moment a call finishes.

Open the consolereceipts doc →
runaii — live usagestreaming

req/min

0

tok/s

0

p50

0ms

14:20 200 glm-5.3-flash 1.4k tok cached 62% $0.00041 182ms

15:23 200 kimi-k3 8.1k tok tools×3 $0.01180 1.1s

16:26 200 glm-5.3-flash 512 tok stream $0.00033 96ms

17:29 429 rate-limited — retry 60s $0.00000 —

18:32 200 qwen3.8-max 2.2k tok cached 31% $0.00290 640ms

[ try it — watch a request happen ]

Type a prompt. See the bill.

A working preview of the request flow — route, stream, meter. The same path your code takes, minus the API key.

POST /v1/chat/completions · glm-5.3-flashdemo
$ Explain vector search in one line
— hit Run to watch tokens stream —
Making request—

Visual demo of the request flow — route, stream, meter. Run it against your own key in the playground →

Silicon, priced per second

all gpus →
8× NVIDIA B300 288GBFP8Max quality for frontier dense models$96/hr $0.0267/s
8× NVIDIA B200 180GBNVFP4Best tokens/$ for large MoE models$80/hr $0.0222/s
4× NVIDIA B300 288GBNVFP4Smallest frontier-capable footprint$48/hr $0.0133/s
4× NVIDIA B200 180GBNVFP4Great for ≤70B dense models$40/hr $0.0111/s
16× NVIDIA GB300FP8For 1T+ parameter models$184/hr $0.0511/s
B — receipts

Us vs everybody

full 7-way table →

Serverless APIs

15+ open models, one URL

Modal: SDK-first · Baseten: curated set

Dedicated GPUs

B200/B300/GB300, per-second

RunPod: 30+ SKUs · Koyeb: A100/L40S

Training loop

Checkpoint → endpoint in seconds

Fireworks: the same loop, closed walls

Start cost

$1 free, no card, no sales call

Baseten: sales-led · Koyeb: credits via program

Model library

all 15 →
GLM-5.3NEWFrontier agentic model. Top-tier coding, tool use, and long-context reas…$1.40/$4.40⚡ $0.260/M cached1M ctxGLM 5.3 FlashNEWUltra-fast multimodal workhorse for chat, extraction, and routing.…$0.15/$0.50⚡ $0.030/M cached1M ctxKimi K3NEWFrontier open-weights model for coding, reasoning, and long context.…$3.00/$15.001M ctxDeepSeek V4 ProSWE-bench leader. Strong math, code, and agentic tool use.…$1.32/$3.96⚡ $0.044/M cached1M ctxDeepSeek V4 FlashSpeed-tuned V4 for high-volume pipelines.…$0.22/$0.66⚡ $0.022/M cached1M ctxQwen3.8-MaxQwen flagship. Multilingual strength and dense world knowledge.…$2.00/$6.00262K ctx

Ship tonight

cookbook →

example_01.py

Streaming chat in 9 lines

OpenAI SDK, new base URL, token-by-token SSE.

model="runaii/glm-5.3-flash"

example_02.py

RAG that cites sources

Embed → retrieve → rerank → answer, cached.

model="runaii/qwen3-embedding-8b"

example_03.py

Tool-calling agent loop

Declare tools, loop 8 deep, cap the spend.

model="runaii/kimi-k3"

-60%

model bill after switching

We moved our main agent to an open model and nobody noticed — except the bill dropped 60%.
Platform team — beta partner

<1s

cold start on B200s

Cold starts under a second on B200s. Evals went from overnight to over lunch.
ML lead — inference beta

1 day

base-URL migration to prod

One base URL change and we were compatible. The playground sold the team in a day.
Founding engineer — early access
customer stories →

Start with $1.
Stay for the bill.

Get started →Pricing
runaiicloud

Serverless inference, dedicated GPUs, and training for open models. OpenAI- and Anthropic-compatible APIs.

© 2026 runaii

Platform

Model libraryGPUsPricingCompare providersSavings calculatorDocsServerlessDeploymentsTrainingBatch API

Developers

PlaygroundCookbookCLIAgents / MCPResearch notesUI/UX systemUse casesTutorialsModel advisorBlogCustomersFAQ

Company

EnterpriseStartupsAboutCareersPartnersTrust centerSLAStatusChangelogrunaii chatSupportAPI keysTermsPrivacy