runaii
cloud
Models
GPUs
Pricing
Docs
Connect
Compare
Enterprise
Playground
☀
Log in
Get started
☰
Blog
Notes from the fleet
2026-09-14
Article
Billing at the micro: how runaii.cloud meters every token
Reserve → settle → refund, exactly-once ledger events, and a parity view that proves balances are honest. The inside of our per-token billing engine.
billing
engineering
ledger
2026-09-14
Article
Your prompts are cached — here is what that saves you
Prompt caching on GLM-5.3, Kimi K3, and DeepSeek V4 bills cache hits at up to 10x less. How the split works, and how to structure prompts to hit it.
pricing
caching
tips
2026-09-13
Article
Serving GLM-5.3-flash at 412 tok/s on RTX PRO 6000
Inside the lazarus fleet — SGLang tuning, spot VMs, checkpoint-wake, and why scale-to-zero makes cheap inference possible.
infrastructure
sglang
gpu
2026-09-09
Release
v0.9.0 — Public beta
Real auth: signup/login, protected console, $1 welcome credits
2026-09-04
Tutorial
Ship RAG with Qwen3 embeddings + rerank
Multilingual embeddings, cross-encoder reranking, and a Flash model to answer — the whole pipeline on one gateway.
2026-09-02
Release
v0.8.0 — Console expansion
Teams/roles/invites, audit log, IP allowlists, spend limits
2026-08-28
Tutorial
Tool-calling agents on Kimi K3
A two-tool agent loop with spend guardrails — the pattern behind every support bot and coding assistant.
2026-08-24
Release
v0.7.0 — Marketing site
Pricing with live tier switcher, docs, research notes, model library 2.0
2026-08-20
Tutorial
Run evals at 50% off with Batch
Nightly SWE-bench-style evals without the daytime bill. Submit async, download results.
2026-08-14
Tutorial
Migrate from OpenAI in one diff
The drop-in migration Fireworks proved and we copied: same SDK, new base URL.
2026-08-10
Release
v0.6.0 — Catalog
15 serverless models, B200/B300/GB300 presets, region premiums