Your GPUs.Your fabric.Nothing in between.

TensorBay builds dedicated bare-metal AI clusters — no hypervisor, no noisy neighbors, no oversubscribed fabric. Every cluster ships with a published acceptance test: NCCL all-reduce numbers you verify before you pay.

Hypervisor overhead
0%
InfiniBand per node
3.2 Tbps
Cluster goodput, 90 days
97.4%
Power across 5 sites
1 GW
HGX B200 — RACK ELEVATIONSN TB-0834-COL
42U — 2000 MM0102030405060708IB QUANTUM-X800 — LEAF AIB QUANTUM-X800 — LEAF B800G XDRDIRECT-TO-CHIP LIQUID LOOP — 130 KW
  • HGX B200Available
  • HGX H200Available
  • HGX H100Available
  • HGX B300Available
  • GB200 NVL72Reserve
  • GB300 NVL72On request
  • VR200Early buy · Mar 2027

On the factory floor right now

  • Tier-1 retail bankNLP workloads · reserved capacity
  • Global energy services firmSeismic ML · training runs
  • Food & protein producerSupply-chain forecasting · fine-tunes
  • Digital payments platformFraud detection · real-time inference
  • Consumer credit bankCredit-risk models · fine-tuning
  • Metropolitan governmentCitizen-service LLMs · inference
  • National education networkEducational AI · inference fleet
  • Agentic AI startupAgentic AI R&D · reserved H100
  • Regional AI cloud providerAI cloud services · reserved nodes

Zero preemptions in six months. The acceptance report matched what we measured on day one — that's never happened with any other provider.

Head of Infrastructure, frontier training lab

01Hardware

The line-up. Priced in the open.

Every SKU below is a whole machine — full HBM, full NVLink, full NIC bandwidth. Reserved rates published for 12- and 36-month terms; capacity is allocated in whole nodes. If a price says "on request", it's supply, not secrecy.

GPU memory
141 GB HBM3e · 4.8 TB/s
Intra-node
NVLink 4 · 900 GB/s
Inter-node
8×400G IB NDR · 3.2 Tbps
Node config
8 GPU · 30 TB NVMe · 2 TB RAM
Best for
Long-context inference, KV-heavy

Reserved · yearly · from

$24/node-hr

HGX H100

Workhorse
GPU memory
80 GB HBM3 · 3.35 TB/s
Intra-node
NVLink 4 · 900 GB/s
Inter-node
8×400G IB NDR · 3.2 Tbps
Node config
8 GPU · 15 TB NVMe · 2 TB RAM
Best for
Cost-efficient training ≤1K GPUs

Reserved · yearly · from

$15/node-hr

HGX B300

Blackwell Ultra
GPU memory
288 GB HBM3e · 8 TB/s
Intra-node
NVLink 5 · 1.8 TB/s per GPU
Inter-node
Quantum-X800 XDR 800G
Node config
8 GPU · 30 TB NVMe · 2 TB RAM
Best for
Memory-bound training, FP4 mega-inference

Reserved · yearly · from

$45.00/node-hr

GB200 NVL72

Rack-scale
GPU memory
186 GB HBM3e/GPU · 13.4 TB per rack
Intra-node
72-GPU coherent NVLink 5 domain
Inter-node
Quantum-X800 XDR between racks
Node config
Whole or half rack · liquid-cooled
Best for
400B+ dense / MoE, trillion-param

Reserved · yearly

On request

GB300 NVL72

Rack-scale
GPU memory
288 GB HBM3e/GPU · 20.7 TB per rack
Intra-node
72-GPU coherent NVLink 5 domain
Inter-node
Quantum-X800 XDR between racks
Node config
Whole or half rack · liquid-cooled
Best for
Frontier MoE, reasoning-heavy inference

Reserved · yearly

On request

VR200 NVL72

Early buy — delivery March 2027
GPU memory
288 GB HBM4 · 22 TB/s
Intra-node
NVLink 6 · 3.6 TB/s per GPU
Inter-node
ConnectX-9 · 1.6T XDR
Node config
NVL72 rack · liquid-cooled
Best for
Rubin-generation reserve capacity

Reserved · yearly

On request

02Why bare metal

The hypervisor tax is real. We don't charge it.

Virtualized GPU clouds skim 3–8% of your FLOPs before your job starts and add jitter your all-reduce can feel. TensorBay hands you the machine — firmware to fabric.

  • F-01

    No hypervisor, no tax

    Jobs run on the metal. Measured 5–7% higher sustained MFU versus virtualized instances on identical HGX H100 hardware — and zero p99 latency spikes from neighbor VMs.

  • F-02

    Single-tenant by physics, not policy

    Your cluster is your hardware: dedicated nodes, dedicated leaf switches, dedicated storage lanes. Isolation you can trace on a wiring diagram, not a compliance PDF.

  • F-03

    Acceptance-tested before handover

    Every cluster passes a 72-hour burn-in: NCCL all-reduce sweeps, HBM bandwidth checks, thermals under sustained load. You get the report. Sign-off is yours, not ours.

  • F-04

    Root on everything

    Custom kernels, your own drivers, NIC firmware tuning, RDMA settings. Bring Slurm, Kubernetes, or bare SSH — we pre-stage images but never lock the BIOS.

  • F-05

    Hot spares in-rack

    Every reserved cluster includes ≥2% hot-spare nodes on the same fabric. Node failure means a reschedule, not a support ticket: 15-minute replacement SLA, automated.

  • F-06

    Deterministic pricing

    Per-minute billing, published rates, zero ingress, 20 TB of egress included per node-month, no "IB surcharge" surprise. The number on this page is the number on the invoice.

03The factory

A factory, not a marketplace.

We don't broker other people's GPUs. TensorBay designs, builds, and operates its own halls — liquid-cooled, high-density, and instrumented down to per-GPU power draw. When output is measured in tokens, the plant matters.

  1. 01Power

    1 GW

    Contracted across 5 sites — US, Philippines, Brazil, Norway and Switzerland — anchored on hydro-rich grids.

  2. 02Cooling

    PUE 1.12

    Direct-to-chip liquid cooling on 130 kW racks. Every watt saved is a watt on silicon.

  3. 03Machine hall

    24,000+ GPUs

    New capacity commissioned in 6-week blocks — 72-hour burn-in included, report published.

  4. 04Yield

    97.4%

    Cluster goodput, trailing 90 days. 0.4% monthly node failure rate, hot-swapped from in-rack spares.

Factories publish yield. So do we — live, at status.tensorbay.com

04Fabric & storage

The fabric is the product.

Training throughput dies in the network and the filesystem long before it dies in the GPU. Ours are specified like the rest of the machine — exactly, in writing.

The fabric is the product.
LayerSpecification
Training fabric — Hopper8×400 Gb/s InfiniBand NDR per node (3.2 Tbps), rail-optimized non-blocking fat-tree, SHARP in-network reduction
Training fabric — BlackwellQuantum-X800 InfiniBand XDR, 800G per link; GB200 racks add a 72-GPU coherent NVLink 5 domain at 1.8 TB/s per GPU
Oversubscription1:1, all tiers. No blended east-west fabric, no shared spine with other tenants
Local scratch30 TB NVMe Gen5 per node, ~55 GB/s read — checkpoint staging without touching the network
Parallel filesystemManaged WEKA, dedicated per cluster: up to 720 GB/s aggregate read per pod, POSIX + GPUDirect Storage
Object storageS3-compatible, NVMe-cached, co-located with compute — $0.055/GB-month hot tier
Data transferIngress $0 · 20 TB/node-month egress included, then $1.25/TB · inter-node $0. Free 100G Direct Connect on yearly reservations
Front-end networkDual 100 GbE per node, DDoS-protected, BYO-IP supported

05Pricing

List prices. Like a factory quotes.

Reserved capacity, priced in the open: monthly and yearly terms with node-replacement and goodput SLAs written into the contract, not the FAQ. On-demand and spot are on the roadmap — reserve is how you buy today.

Per node-hourMonthly termYearly term
HGX B200from $56from $39
HGX H200from $34from $24
HGX H100from $22from $15
HGX B300from $64from $45.00
GB200 NVL72On request
GB300 NVL72On request
VR200 NVL72Early buy · delivery Mar 2027
SLA99.9% monthly + 15-min node replacement + goodput credits
Scale128–10,000+ GPUs · dedicated fabric per tenant
BillingPer-minute metering inside the reservation · $0 ingress

Storage

  • Local NVMe scratch — 30 TB per nodeIncluded
  • Parallel filesystem — managed WEKAdedicated per cluster$0.09/GB-mo
  • Object storage — hot, S3-compatibleNVMe-cached, co-located$0.055/GB-mo
  • Object storage — archive$0.015/GB-mo
  • Snapshots$0.021/GB-mo

Data transfer

  • Ingressalways$0
  • Egress included20 TB/node-mo
  • Egress overageall sites$1.25/TB
  • Inter-node / east-westInfiniBand fabric$0

Networking

  • East-west InfiniBand fabricdedicated per clusterIncluded
  • Private VLAN / VPCIncluded
  • Additional public IPv4IPv6 free$4/IP-mo
  • BYO-IP — /24 v4 or /48 v6$200 setupFree
  • Managed firewall$4/node-mo
  • DDoS protectionIncluded
  • Direct Connect 10G / 100G100G free on yearly terms$450 / $1,800-mo

Prices per 8-GPU HGX node-hour; rack-scale systems quoted per rack. Prices exclude applicable taxes. Managed Slurm and Kubernetes control planes: $0. Storage and Direct Connect priced above; nothing else exists to charge for. Long-term and multi-year contracts: on request.

06Use cases

Built for jobs that notice the difference.

  • Foundation & frontier training

    512–10K+ GPU reserved clusters on dedicated XDR fabric with in-rack spares and goodput SLAs. When a 6-week run costs seven figures, 97% goodput versus 91% is the whole margin.

    → Reserved B200 / GB200 / GB300 pods
  • Fine-tuning & research blocks

    Whole 8-GPU nodes on short reserved blocks — per-minute metering inside the reservation, WEKA-backed datasets that persist between runs. Provisioned in under 4 minutes, returned without a ticket.

    → H100 / H200 nodes from $15/hr
  • Production inference fleets

    Memory-dense H200 and B300 nodes for long-context and 70B+ serving; bare-metal p99s with no virtualization jitter and 20 TB of included egress per node, every month.

    → H200 from $24/hr · B300 from $45/hr

07Security

Isolation you can audit.

Single-tenant hardware is the security model — everything else is attestation of it. We publish report scope because an acronym without one is marketing.

  • SOC 2 Type II

    Scope: bare-metal compute, fabric, and managed storage

  • ISO 27001:2022

    Certified — all datacenter sites

  • HIPAA

    BAA available on dedicated clusters

  • GDPR

    EU region with in-country data residency

  • Physically dedicated nodes, leaf switches, and storage per tenant
  • Full-disk encryption at rest; cryptographic erase + NVMe sanitize on release
  • Certified media destruction on decommission
  • SSO/SAML, hardware-key MFA, immutable audit logs
  • Trust portal with pen-test summaries: trust.tensorbay.com

08FAQ

The questions that come up on every call.

Short answers, same numbers as the rest of this page. If something here matters to your contract, it goes in the contract.

  • How much does a bare-metal GPU cluster cost?

    Rates are published per 8-GPU node-hour on a yearly term: HGX H100 from $15, HGX H200 from $24, HGX B200 from $39, HGX B300 from $45. Monthly terms are higher. GB200, GB300 and VR200 NVL72 racks are quoted per rack — that is a supply constraint, not a pricing one.

  • Can I rent a single GPU, or a fraction of one?

    No. Capacity is allocated in whole 8-GPU nodes, or whole NVL72 racks for rack-scale systems. Slicing a node means a hypervisor, and the hypervisor is the thing we removed.

  • Do you offer on-demand or spot instances?

    Not today. Reserved monthly and yearly terms are how you buy; on-demand and spot are on the roadmap. Inside a reservation, metering is per minute.

  • What is the hypervisor tax, and what does bare metal actually recover?

    Virtualized GPU clouds skim 3–8% of your FLOPs before the job starts and add jitter your all-reduce can feel. On identical HGX H100 hardware we measure 5–7% higher sustained MFU than virtualized instances, and no p99 latency spikes from neighbor VMs.

  • What is the acceptance test?

    Every cluster runs a 72-hour burn-in before handover: NCCL all-reduce sweeps, HBM bandwidth checks and thermals under sustained load. You get the report and you sign off — not us. The quote tells you the exact report format your cluster must pass.

  • What happens when a node fails?

    Every reserved cluster carries at least 2% hot-spare nodes on the same fabric. Replacement is automated with a 15-minute SLA, so a failure is a reschedule rather than a support ticket. Node failure runs about 0.4% monthly and cluster goodput has been 97.4% over the trailing 90 days.

  • Is the InfiniBand fabric shared with other tenants?

    No. Nodes, leaf switches and storage lanes are physically dedicated per tenant at 1:1 oversubscription on every tier — there is no blended east-west fabric and no shared spine. Hopper pods run 8×400 Gb/s NDR per node; Blackwell pods run Quantum-X800 XDR at 800G per link.

  • Do you charge for data transfer?

    Ingress is $0, inter-node traffic is $0, and every node includes 20 TB of egress per month, then $1.25/TB. Yearly reservations include a 100G Direct Connect. Managed Slurm and Kubernetes control planes are $0.

  • Where is the capacity, and can I pin data to a region?

    1 GW is contracted across five sites — the US, the Philippines, Brazil, Norway and Switzerland — anchored on hydro-rich grids. The EU region supports in-country data residency for GDPR, and a HIPAA BAA is available on dedicated clusters.

  • How quickly can I get a cluster?

    Send GPU count, interconnect and timeline and you get a firm quote, a delivery date and the acceptance-test format within 48 hours. Nodes inside an existing reservation provision in under 4 minutes; new capacity is commissioned in 6-week blocks with burn-in included.

09Commission your cluster

Spec your cluster. We'll send the acceptance test with the quote.

Tell us GPU count, interconnect, and timeline — get a firm quote, a delivery date, and the exact burn-in report format your cluster must pass, within 48 hours.

20 TB egress included per nodePer-minute billingLive goodput at status.tensorbay.com