Blackwell Ultra
HGX B300
The HGX B300 is Blackwell Ultra: 288 GB of HBM3e per GPU at 8 TB/s, the densest memory per GPU we rack, on the same 800G Quantum-X800 XDR fabric as B200.
Specification
- GPU memory
- 288 GB HBM3e · 8 TB/s
- Intra-node
- NVLink 5 · 1.8 TB/s per GPU
- Inter-node
- Quantum-X800 XDR 800G
- Node config
- 8 GPU · 30 TB NVMe · 2 TB RAM
- Best for
- Memory-bound training, FP4 mega-inference
B300 pods run Quantum-X800 InfiniBand XDR at 800G per link with NVLink 5 at 1.8 TB/s per GPU inside the node, 1:1 on every tier.
Why this machine
B300 exists for the jobs that stall on capacity rather than arithmetic — long-context serving at scale, memory-bound training, and mega-inference where the KV cache decides how many nodes you need. At 288 GB per GPU it carries 60% more HBM than a B200 for roughly 15% more per node-hour, which is usually the cheaper end state when memory is the binding constraint.
Everything else matches the rest of the line: no hypervisor, dedicated leaf switches, root on the machine, a 72-hour acceptance test you sign off, and at least 2% hot spares on the same fabric with a 15-minute automated replacement SLA.
What it is for
Memory-bound training
Runs where activation and optimizer state, not FLOPs, set the node count. Fewer, denser nodes shorten the critical path.
FP4 mega-inference
Very large models served at low precision with the full HBM budget available and no virtualization layer taking a cut.
Long-context serving
KV caches that would force a B200 fleet onto more nodes fit in fewer B300 nodes, on the same XDR fabric.
HGX B300 questions
How much does an HGX B300 node cost?
From $45 per 8-GPU node-hour on a yearly reservation, or from $64 monthly, for the whole node: 8 GPUs, 30 TB of NVMe and 2 TB of RAM.
When is B300 worth the premium over B200?
When memory capacity is the constraint. B300 carries 288 GB of HBM3e at 8 TB/s against 180 GB at 7.7 TB/s, for $45 against $39 per node-hour. If a B200 fleet needs materially more nodes to hold the same working set, B300 is cheaper overall.
Is B300 a rack-scale system?
No. B300 is an 8-GPU HGX node like B200 and H200, allocated in whole nodes. GB200, GB300 and VR200 are the NVL72 rack-scale systems, quoted per rack.
What happens if a node fails mid-run?
Every reserved cluster carries at least 2% hot spares on the same fabric, replaced automatically under a 15-minute SLA. Node failure runs about 0.4% monthly and trailing-90-day cluster goodput has been 97.4%.