Efficient · NVFP4
$40/hr $0.0111/sec · $29,200/mo at 24/7
Great for ≤70B dense models. ≤70B dense models, evals.
VRAM pool
180GB ×4 = 720GB
Position
vs L40S 4×: ~2.5× inference speed
Compute die
NVIDIA B200 180GB — NVFP4 tensor cores, tuned for ≤70b dense models, evals.
HBM stack
180GB ×4 = 720GB of high-bandwidth attached memory — the whole model fits, no offload.
Interposer / NVLink
NVLink-C2C fabric so the 4 GPUs act like one pool.
Per-second metering
$40/hr → 0.0111/sec. Pause the deployment and the meter hits $0.
Scale-to-zero
Autoscale + region pins (GLOBAL / US / EU / APAC) on the control plane.