Full Precision
Popular8× NVIDIA B300 288GB
$96/hr $0.0267/sec
288GB ×8 = 2.3TB · FP8
Max quality for frontier dense models — Frontier dense models, 1M-ctx serving.
Hardware
Blackwell-class capacity and multi-cloud GPU VMs, per-second billing, autoscale to zero. Every preset and instance launchable in one click.
Full Precision
Popular$96/hr $0.0267/sec
288GB ×8 = 2.3TB · FP8
Max quality for frontier dense models — Frontier dense models, 1M-ctx serving.
Throughput
Popular$80/hr $0.0222/sec
180GB ×8 = 1.4TB · NVFP4
Best tokens/$ for large MoE models — Large MoE throughput kings.
Minimal
Popular$48/hr $0.0133/sec
288GB ×4 = 1.1TB · NVFP4
Smallest frontier-capable footprint — Smallest frontier-capable footprint.
Efficient
Popular$40/hr $0.0111/sec
180GB ×4 = 720GB · NVFP4
Great for ≤70B dense models — ≤70B dense models, evals.
Frontier Max
$184/hr $0.0511/sec
288GB ×16 = 4.6TB · FP8
For 1T+ parameter models — 1T+ parameter models.
Full VMs with your choice of OS image (Ubuntu, Debian, CUDA, PyTorch), boot-disk type with published IOPS scaling, local NVMe scratch, and Spot pricing (−35%).
1× T4
$0.44/hr spot $0.29
8 vCPU · 30GB RAM · 16GB VRAM
n1-standard-8
Cheap inference, transcoding, and CUDA experiments
1× L4
Popular$0.87/hr spot $0.57
8 vCPU · 32GB RAM · 24GB VRAM
g2-standard-8
Best price/perf for LLM inference up to ~13B and video
4× L4
$3.6/hr spot $2.34
24 vCPU · 96GB RAM · 96GB VRAM
g2-standard-48
Multi-worker inference farms on one VM
1× RTX PRO 6000
$2.4/hr spot $1.56
24 vCPU · 96GB RAM · 96GB VRAM
g4-standard-48
96GB Blackwell workstation-class VRAM, single GPU
1× A100
$2.93/hr spot $1.90
12 vCPU · 85GB RAM · 80GB VRAM
a2-ultragpu-1g
Ampere workhorse for training and dense inference
8× A100
$23.46/hr spot $15.25
96 vCPU · 680GB RAM · 640GB VRAM
a2-ultragpu-8g
Full-node A100 training
8× H100
$27.35/hr spot $17.78
208 vCPU · 880GB RAM · 640GB VRAM
a3-highgpu-8g
Hopper full node — frontier fine-tunes
1× T4
$0.526/hr spot $0.34
4 vCPU · 16GB RAM · 16GB VRAM
g4dn.xlarge
Low-cost T4 inference
1× L4
$0.8/hr spot $0.52
4 vCPU · 16GB RAM · 24GB VRAM
g6.xlarge
Ada-Lovelace generation L4 inference
8× A100
$32.77/hr spot $21.30
96 vCPU · 1152GB RAM · 320GB VRAM
p4d.24xlarge
Full-node Ampere with 1.6TB local NVMe
8× H100
$98.32/hr spot $63.91
192 vCPU · 2048GB RAM · 640GB VRAM
p5.48xlarge
Hopper full node
1× T4
$0.526/hr spot $0.34
4 vCPU · 28GB RAM · 16GB VRAM
Standard_NC4as_T4_v3
Low-cost T4 inference
1× A100
$0.74/hr spot $0.48
4 vCPU · 64GB RAM · 80GB VRAM
Standard_NC4ads_A100_v4
Single-A100 80GB inference
8× A100
$27.2/hr spot $17.68
96 vCPU · 880GB RAM · 640GB VRAM
Standard_ND96asr_A100_v4
Full-node A100 training
8× H100
$98.32/hr spot $63.91
96 vCPU · 1900GB RAM · 640GB VRAM
Standard_ND96isr_H100_v5
Hopper full node
Deployments bill per second with scale-to-zero. Start with $1 free.