runaiicloud
ModelsGPUsPricingDocsConnectCompareEnterprisePlayground
Log inGet started

Research notes · Sept 2026

We studied the best.
Then remixed them.

Competitive recon across Modal (developer love), Baseten (performance + enterprise), and Fireworks (the training loop) — plus our internal pricing catalog. Here's what each does brilliantly, and what runaii kept.

modal.combaseten.cofireworks.aiSynthesized

Modal

  • —DX wins: stay in Python, ship to cloud. Composable primitives for logic + hardware in one file.
  • —Speed story: sub-second cold starts, instant autoscale, sub-10ms proxy overhead, streaming/WebSocket native.
  • —Breadth: inference + training + batch + sandboxes + notebooks under one SDK — agents are first-class.
  • —GTM: $30/mo free compute, examples gallery, case-study wall (Decagon, Runway, Suno, Quora).
Kept for runaii: One-line deploy, scale-to-zero with burst, examples-first docs, generous free tier ($1 credits to start).

Baseten

  • —Performance is the brand: custom kernels, latest decoding, advanced caching baked into the Inference Stack.
  • —Enterprise trust: cross-cloud HA, 99.99% uptime, single-tenant + self-hosted + hybrid, SOC2/HIPAA.
  • —DevEx for iteration: deploy → optimize → manage loop, model library with Try-It, savings calculator.
  • —Humans scale you: Forward Deployed Engineers from prototype to production.
Kept for runaii: NVFP4/FP8 precision presets, region-pinned deployments (Global/US/EU/APAC), performance transparency, white-glove onboarding.

Fireworks

  • —The loop is the moat: guided → config-led → custom-trainer spectrum; every checkpoint deploys in seconds.
  • —Tiered inference: Serverless (per-token) / On-Demand (dedicated) / Reserved (guaranteed) maps to real budgets.
  • —Library as storefront: model cards with $/1M + context + modality, one-line code switch.
  • —Proof over promises: customer wall (Cursor, Vercel, Notion, Quora) + changelog + demos + cookbooks.
Kept for runaii: Serverless/Priority/Fast per-token tiers, model cards with cached-input pricing, training→deploy in seconds, OpenAI + Anthropic compatibility.

Round 2 recon — JarvisLabs · Koyeb · RunPod

Second sweep, Sept 2026. (Note: runaware.ai doesn't resolve — covered RunPod, the closest AI-developer-cloud match, instead.)

JarvisLabs

GPU-VM roots: root SSH, templates (PyTorch/ComfyUI live in ~2s), per-minute billing, pause = pay storage only. Managed Endpoints (they pick GPU/precision) vs Serverless (your vLLM/SGLang). Wins on price transparency ($3.99 H200/hr on the hero) and per-GPU pages (/gpu/*). CLI + Claude Code skills. Lesson for us: /gpus pages + per-second pricing display + pause/resume semantics on deployments.

Koyeb

Serverless containers for anything (not just models): git-push deploys, 50+ locations, scale-to-zero, one-click app catalog (/deploy), tutorials + community changelog, startup credits up to $30k. Footer IA worth copying: Product / Resources (docs, tutorials, community, API, status) / Company. Lesson: /cookbook + /cli + /agents trio and a Deploy-style quickstart catalog.

RunPod

The AI developer cloud: Pods (30+ SKUs, spot + reserved) + Serverless (FlashBoot sub-200ms) + Clusters, one account, no replatforming. Hub playground per model, use-case pages (/use-cases/*), per-product pages, case-study wall, Trust Center with SOC 2/ISO/HIPAA. Lesson: lifecycle story (experiment → production, no migration) and use-case SEO pages — our next content sprint.

Side-by-side table →

Decisions this research locked

1. One API, two compat layers

Fireworks proved drop-in wins; we ship OpenAI + Anthropic on the same gateway.

2. Three primitives, not ten products

Modal's sprawl confuses; Baseten's stack clarifies. Serverless / Deployments / Training is the IA.

3. Pricing on the card

Fireworks' $/1M + context per card converts. We add cached-input pricing for honesty.

4. Dark console, light marketing mix → dark-first

Fireworks/Modal dark premium feels faster; our console stays OLED-dark, marketing goes cinematic dark.

5. Motion with restraint

Orbs + reveal + marquee + terminal typing; everything disabled under prefers-reduced-motion.

The runaii interface language

The design system that keeps every page consistent: emerald on OLED, Geist type, translucent glass surfaces, and GPU-only motion (transform/opacity/filter) that respects prefers-reduced-motion.

Tokens

Emerald brand ramp, ink typography scale, status colors — CSS variables only.

Type

Geist variable fonts, self-hosted; display role via weight + tight tracking.

Surfaces

Glass cards: 46% translucency, 16px blur, hairline borders, inner highlights.

Motion

300–700ms reveals, spring hovers, view-transition page fades. No animation libraries.

Live homepage + product-page reads (Sept 2026), internal model/hardware catalog recon, and prior zcode / chat-session notes. Re-verify pricing before quoting — providers move fast.

Modal — homepage + products

Python-native DX, sub-second cold starts, 0→1000+ GPU autoscale, sandboxes, observability, $30/mo free compute.

https://modal.com

Baseten — inference platform

Fastest runtimes, cross-cloud HA, kernels + decoding + caching research, cloud/VPC/hybrid, FDE support.

https://www.baseten.co

Fireworks — own your model

Training→inference loop, serverless / on-demand / reserved tiers, model library cards, customer proof wall.

https://fireworks.ai

Internal catalog recon

src/lib/catalog.ts — Fireworks + Modal/Baseten pricing and hardware presets (B200/B300/GB300).

/models
See the UI/UX system ↑Browse the model library
runaiicloud

Serverless inference, dedicated GPUs, and training for open models. OpenAI- and Anthropic-compatible APIs.

© 2026 runaii

Platform

Model libraryGPUsPricingCompare providersSavings calculatorDocsServerlessDeploymentsTrainingBatch API

Developers

PlaygroundCookbookCLIAgents / MCPResearch notesUI/UX systemUse casesTutorialsModel advisorBlogCustomersFAQ

Company

EnterpriseStartupsAboutCareersPartnersTrust centerSLAStatusChangelogrunaii chatSupportAPI keysTermsPrivacy