Full Precision · FP8
$96/hr $0.0267/sec · $70,080/mo at 24/7
Max quality for frontier dense models. Frontier dense models, 1M-ctx serving.
VRAM pool
288GB ×8 = 2.3TB
Position
vs H100 8×: ~2× tokens/$ on MoE
Compute die
NVIDIA B300 288GB — FP8 tensor cores, tuned for frontier dense models, 1m-ctx serving.
HBM stack
288GB ×8 = 2.3TB of high-bandwidth attached memory — the whole model fits, no offload.
Interposer / NVLink
NVLink-C2C fabric so the 8 GPUs act like one pool.
Per-second metering
$96/hr → 0.0267/sec. Pause the deployment and the meter hits $0.
Scale-to-zero
Autoscale + region pins (GLOBAL / US / EU / APAC) on the control plane.