HGX H200
The HGX H200 is the memory-dense Hopper node: 141 GB of HBM3e per GPU at 4.8 TB/s, which is what long-context and KV-cache-heavy serving actually runs out of first.
Specification
- GPU memory
- 141 GB HBM3e · 4.8 TB/s
- Intra-node
- NVLink 4 · 900 GB/s
- Inter-node
- 8×400G IB NDR · 3.2 Tbps
- Node config
- 8 GPU · 30 TB NVMe · 2 TB RAM
- Best for
- Long-context inference, KV-heavy
H200 pods run 8×400 Gb/s InfiniBand NDR per node — 3.2 Tbps — as a rail-optimized non-blocking fat-tree with SHARP in-network reduction, at 1:1 oversubscription.
Why this machine
Most inference fleets are not FLOP-bound, they are memory-bound. H200 gives you 76% more HBM per GPU than an H100 at the same 8-GPU node shape and the same 3.2 Tbps of InfiniBand NDR, which usually means fewer nodes for the same context length and concurrency.
Serving is also where virtualization hurts most visibly: a neighbor VM shows up in your p99, not your average. On bare metal there is no neighbor. Nodes, leaf switches and storage lanes are physically dedicated, so tail latency is a property of your own load.
What it is for
Long-context inference
KV-heavy serving where context length, not arithmetic, sets the node count. More HBM per GPU means fewer nodes for the same window.
Production 70B+ fleets
Bare-metal p99s with no virtualization jitter, and 20 TB of egress included per node every month.
Fine-tuning blocks
Whole 8-GPU nodes on short reserved terms, provisioned in under four minutes inside an existing reservation and returned without a ticket.
HGX H200 questions
How much does an HGX H200 node cost?
From $24 per 8-GPU node-hour on a yearly reservation, or from $34 monthly. Ingress is $0, inter-node traffic is $0, and 20 TB of egress per node-month is included.
H200 or H100 for inference?
H200 carries 141 GB of HBM3e at 4.8 TB/s against the H100’s 80 GB of HBM3 at 3.35 TB/s, for $24 versus $15 per node-hour. If context length or model size is forcing you onto more nodes, H200 is usually cheaper in total. If the model fits comfortably in 80 GB, H100 wins on price.
Can I mix H200 and H100 in one cluster?
Capacity is allocated in whole nodes on a dedicated fabric, so a mixed reservation is a scoping question for the quote rather than a product limit. Send the workload and timeline and it comes back in the spec.
What storage is attached?
30 TB of Gen5 NVMe per node at roughly 55 GB/s for local scratch, plus managed WEKA dedicated to your cluster at up to 720 GB/s aggregate read per pod, POSIX and GPUDirect Storage.