runaiicloud
ModelsGPUsPricingDocsConnectCompareEnterprisePlayground
Log inGet started
← All GPUs
DIEHBMINTERPOSERNVSWITCHSUBSTRATESUBSTRATE · POWER & THERMALINTERPOSER · NVLink-C2C FABRICHBMHBMHBMHBMGB300FP8 · 16×NVSWITCH

Frontier Max · FP8

16× NVIDIA GB300

$184/hr $0.0511/sec · $134,320/mo at 24/7

For 1T+ parameter models. 1T+ parameter models.

VRAM pool

288GB ×16 = 4.6TB

Position

The frontier ceiling

Deploy this →Configure sizeCompare providers →

What you're renting

Compute die

NVIDIA GB300 — FP8 tensor cores, tuned for 1t+ parameter models.

HBM stack

288GB ×16 = 4.6TB of high-bandwidth attached memory — the whole model fits, no offload.

Interposer / NVLink

NVLink-C2C fabric so the 16 GPUs act like one pool.

Per-second metering

$184/hr → 0.0511/sec. Pause the deployment and the meter hits $0.

Scale-to-zero

Autoscale + region pins (GLOBAL / US / EU / APAC) on the control plane.

Reserve it in minutesDeploy →

Other presets

8× NVIDIA B300 288GB

$96/hr · FP8

Max quality for frontier dense models

8× NVIDIA B200 180GB

$80/hr · NVFP4

Best tokens/$ for large MoE models

4× NVIDIA B300 288GB

$48/hr · NVFP4

Smallest frontier-capable footprint

4× NVIDIA B200 180GB

$40/hr · NVFP4

Great for ≤70B dense models

runaiicloud

Serverless inference, dedicated GPUs, and training for open models. OpenAI- and Anthropic-compatible APIs.

© 2026 runaii

Platform

Model libraryGPUsPricingCompare providersSavings calculatorDocsServerlessDeploymentsTrainingBatch API

Developers

PlaygroundCookbookCLIAgents / MCPResearch notesUI/UX systemUse casesTutorialsModel advisorBlogCustomersFAQ

Company

EnterpriseStartupsAboutCareersPartnersTrust centerSLAStatusChangelogrunaii chatSupportAPI keysTermsPrivacy