All GPUs

NVIDIA B200 price: $5.50/hr

NVIDIA's Blackwell GPU: 180GB of HBM3e per GPU, FP4 Tensor Cores at 9 PFLOPS dense, and eight GPUs per HGX node. On OpenRelay it runs your container as a pod.

See all GPUs

What the rate buys

Per GPU-hour
$5.50
billed while the pod runs
8-GPU node, per hour
$44.00
one pod, one node
One GPU, 24/7
$3,960
per month, 720 hours

The B200 runs as a pod: your container image on 1 to 8 whole GPUs of a single node, with SSH access and an HTTPS endpoint. Every pod gets a persistent volume at /workspace, 10 to 1,024 GB, that adds nothing to the hourly rate and survives a stop or restart. A stopped pod is not billed. The details are in the pods documentation.

B200 specifications

SpecificationB200
ArchitectureNVIDIA Blackwell
Memory180GB HBM3e
FP4 Tensor Core9 PFLOPS dense (18 sparse)
FP8 and FP6 Tensor Core4.5 PFLOPS dense
FP16 and BF16 Tensor Core2.25 PFLOPS dense
FP6437 TFLOPS
NVLink5th generation, 1.8 TB/s per GPU
Compute capability10.0 (CUDA 12.8 or later)

Source: NVIDIA HGX B200 specifications, per GPU (board totals divided by 8), read 2026-10-01.

What the B200 is for

Models that outgrow an H100

180GB is 2.25 times an H100's 80GB, so a 70B model in FP16, about 140GB of weights, fits on one GPU instead of two.

FP4 inference

FP4 Tensor Cores run at 9 PFLOPS dense, twice the FP8 rate. Hopper has no FP4 path at all.

FP64 alongside AI

The B200 keeps 37 TFLOPS of FP64, which the B300 gives up, so simulation code that needs double precision belongs here.

Your own serving stack

A pod runs the container you bring, so vLLM, SGLang, TensorRT-LLM or your own server run as you configure them.

Frequently asked questions

How much does it cost to rent an NVIDIA B200?

$5.50 per GPU-hour, billed while the pod is running; a stopped pod is not billed. A full node of eight is $44.00 an hour, about $31,680 a month around the clock. The /workspace volume is included in the rate.

Why 180GB and not the 192GB other sources quote?

192GB is the figure for the B200 die. Every shipping HGX B200 board carries 180GB per GPU: NVIDIA lists 1.4TB for the eight-GPU board. 180GB is what your container sees.

Should I pick the B200 or the H100?

The B200 has 180GB against 80GB, 4.5 PFLOPS of dense FP8 against the H100's 1.98, and FP4, which the H100 does not have. The H100 costs less per hour and is the better buy when the model fits in 80GB.

Will my CUDA container run on the B200?

Yes, if it is built against CUDA 12.8 or later. That is the release that added the B200's architecture (sm_100); an image built for an older CUDA cannot compile kernels for it.

How do I get B200 capacity?

Book a call. B200 nodes are allocated with our team: we size the GPU count and the term with you, and you run pods on them like on any other GPU we sell.

B200 capacity, sized with you

We allocate B200 nodes with you: how many GPUs, for how long, and on what terms. Book a call and we will size it.

See all GPUs