runaiicloud
ModelsGPUsPricingDocsConnectCompareEnterprisePlayground
Log inGet started

Pricing · $1 free on signup

Pay for tokens.
Not for vibes.

Serverless per-token inference, per-second dedicated GPUs, and training priced per 1M tokens. Every price on one page — pick a tier to see it live.

Serverless inference

Per 1M tokens · Standard tier · OpenAI + Anthropic compatible

Browse model cards →
ModelModalityInput / 1MCached / 1MOutput / 1MContext
GLM-5.3NEWLLM$1.40$0.260$4.401M
GLM 5.3 FlashNEWVision$0.15$0.030$0.501M
Kimi K3NEWVision$3.00—$15.001M
DeepSeek V4 ProLLM$1.32$0.044$3.961M
DeepSeek V4 FlashLLM$0.22$0.022$0.661M
Qwen3.8-MaxLLM$2.00—$6.00262K
Qwen3.8 FlashVision$0.40—$1.60262K
OpenAI gpt-oss-120bLLM$0.15—$0.60131K
OpenAI gpt-oss-20bLLM$0.05—$0.20131K
MiniMax M3LLM$0.30—$1.20512K
Llama 3.3 70B InstructLLM$0.23—$0.85131K
Qwen3 Embedding 8BEmbedding$0.10——41K
Qwen3 Reranker 8BRerank$0.20——41K
Whisper V3 LargeAudio$0.25———
FLUX.1 Kontext ProImage$0.04———

Batch API: 50% off all serverless rates · 24h turnaround · same gateway.

Picked your model? Start with $1 free.

Paste one key into your app — billed per token from the table above.

Get started →Read the docs

Dedicated GPUs

Per-second billing · autoscale + scale-to-zero · your weights, your VPC.

Per-GPU deep dive →

Full Precision

Popular

$96/hr

NVIDIA B300 288GB ×8

FP8 · Max quality for frontier dense models

Deploy this preset

Throughput

Popular

$80/hr

NVIDIA B200 180GB ×8

NVFP4 · Best tokens/$ for large MoE models

Deploy this preset

Minimal

Popular

$48/hr

NVIDIA B300 288GB ×4

NVFP4 · Smallest frontier-capable footprint

Deploy this preset

Efficient

Popular

$40/hr

NVIDIA B200 180GB ×4

NVFP4 · Great for ≤70B dense models

Deploy this preset

Frontier Max

$184/hr

NVIDIA GB300 ×16

FP8 · For 1T+ parameter models

Deploy this preset

Training

Managed SFT/DPO priced per 1M training tokens (dataset tokens × epochs). RL jobs bill per GPU-second at on-demand rates. Checkpoints deploy to inference at base-model prices.

Model sizeLoRA SFTLoRA DPOFull SFTFull DPO
Up to 16B params e.g. gpt-oss-20b$0.50$1.00$1.00$2.00
16–80B params e.g. llama-3.3-70b$3.00$6.00$6.00$12.00
80–300B params e.g. qwen3.8-235b-class$6.00$12.00$12.00$24.00
300B+ params e.g. deepseek-v4, kimi-k3$10.00$20.00$20.00$40.00

Per 1M training tokens · VLM fine-tuning billed the same way · open training console →

Regions

  • Global — Recommended — lowest cost, any-traffic routing×1.00
  • US only — Data residency: United States×1.00
  • EU only — Data residency: European Union×1.15
  • APAC only — Data residency: Asia-Pacific×1.15

Startups

Early-stage teams get $500 in credits, locked beta pricing, and a direct line to engineering.

Startup program →

[ credits ]

Buy credits once.
Use them anywhere.

One balance across everything we meter — serverless tokens, per-second GPU deployments, and training runs. The ledger splits every call into a line item, so you always know exactly what it cost.

$1 free on signup · no card · balance never expires

Top up in console →See how metering works
  • Serverless tokens

    15+ models, one key. Cached prefixes up to 97% off.

  • Dedicated GPU-seconds

    B200/B300/GB300, autoscale, scale-to-zero. Same balance.

  • Training runs

    SFT/DPO per 1M tokens, RL per GPU-second. Ships to prod in seconds.

  • Never expires

    Sit on your balance for a year — it stays yours.

Plans that scale with you

Usage is always metered separately. Plans buy limits, support, and guarantees.

Starter

$0 + usage

For developers getting a first agent into production.

  • ✓$1 welcome credits
  • ✓15+ serverless models
  • ✓60 req/min rate envelope
  • ✓Community support (Discord)
  • ✓7-day log retention
  • ✓1 workspace seat
Start with $1 free

Team

Popular

$250/mo + usage

For startups scaling inference with guardrails.

  • ✓Everything in Starter
  • ✓$100 monthly credit allowance
  • ✓Priority tier + higher rate limits
  • ✓Unlimited seats + roles
  • ✓30-day log retention
  • ✓Spend limits + daily caps + alerts
  • ✓Slack support
Talk to us

Enterprise

Custom

For regulated scale: residency, VPC, and contracts.

  • ✓Everything in Team
  • ✓VPC / self-hosted + hybrid
  • ✓SSO/SAML + audit log + DPA
  • ✓Reserved B200/B300/GB300
  • ✓99.99% uptime SLA
  • ✓Forward-deployed engineer
Contact sales

FAQ

Do credits expire?

No — your balance never expires. Top up once and spend it across serverless tokens, GPU deployments, and training runs whenever you're ready.

What does $1 free credits get me?

About 700K input tokens on GLM-5.3, or ~6.6M on GLM Flash. Enough to prototype a full agent loop before paying.

How does cached-input pricing work?

Repeat prompt prefixes (system prompts, RAG context) are billed at the cached rate — up to 97% off on DeepSeek V4 Pro. Cache hits are automatic.

Serverless vs dedicated — which do I need?

Serverless for variable traffic and prototyping. Dedicated (B200/B300/GB300, per-second) when you need guaranteed throughput, custom weights, or VPC isolation.

Do you offer batch discounts?

Yes — async batch inference is 50% off serverless rates, same as the Fireworks model we benchmarked against.

How does region pinning affect price?

Global routing is cheapest. US-only is the same price; EU/APAC add 15% for data-residency guarantees.

Start with $1 free

No card required. One API key, every model.

Claim creditsRead the docs

Do the math — savings calculator

Estimate what moving your monthly inference spend to an open-model mix saves you. Rows above are the real prices.

$900

closed APIs

$25

runaii mix

$875

you keep / mo

Estimate using a Flash-heavy open mix; exact numbers on this page. Cached prefixes + batch cut it further. Start with $1 free →

runaiicloud

Serverless inference, dedicated GPUs, and training for open models. OpenAI- and Anthropic-compatible APIs.

© 2026 runaii

Platform

Model libraryGPUsPricingCompare providersSavings calculatorDocsServerlessDeploymentsTrainingBatch API

Developers

PlaygroundCookbookCLIAgents / MCPResearch notesUI/UX systemUse casesTutorialsModel advisorBlogCustomersFAQ

Company

EnterpriseStartupsAboutCareersPartnersTrust centerSLAStatusChangelogrunaii chatSupportAPI keysTermsPrivacy