All GPUs

NVIDIA B300 price: $7.49/hr

NVIDIA's Blackwell Ultra GPU: 288GB of HBM3e per GPU and 13.5 PFLOPS of dense FP4, eight GPUs per HGX node. On OpenRelay it runs your container as a pod.

See all GPUs

What the rate buys

Per GPU-hour
$7.49
billed while the pod runs
8-GPU node, per hour
$59.92
one pod, one node
One GPU, 24/7
$5,393
per month, 720 hours

The B300 runs as a pod: your container image on 1 to 8 whole GPUs of a single node, with SSH access and an HTTPS endpoint. Every pod gets a persistent volume at /workspace, 10 to 1,024 GB, that adds nothing to the hourly rate and survives a stop or restart. A stopped pod is not billed. The details are in the pods documentation.

B300 specifications

SpecificationB300
ArchitectureNVIDIA Blackwell Ultra
Memory288GB HBM3e
FP4 Tensor Core13.5 PFLOPS dense (18 sparse)
FP8 and FP6 Tensor Core4.5 PFLOPS dense
FP16 and BF16 Tensor Core2.25 PFLOPS dense
FP641.25 TFLOPS
NVLink5th generation, 1.8 TB/s per GPU
Compute capability10.3 (CUDA 12.9 or later)

Source: NVIDIA HGX B300 specifications, per GPU (board totals divided by 8), read 2026-10-01.

What the B300 is for

The largest models per GPU

288GB holds a 120B-parameter model in FP16, about 240GB of weights, on a single GPU.

4-bit inference

13.5 PFLOPS of dense FP4 is 1.5 times the B200, for models served at 4 bits.

Long context

The extra memory over the B200 goes to the KV cache, which is what grows with context length and batch size.

Not for FP64

Blackwell Ultra gives up double precision: 1.25 TFLOPS of FP64 against the B200's 37. Code that needs FP64 belongs on a B200 or an MI355X.

Frequently asked questions

How much does it cost to rent an NVIDIA B300?

$7.49 per GPU-hour, billed while the pod is running; a stopped pod is not billed. A full node of eight is $59.92 an hour, about $43,142 a month around the clock. The /workspace volume is included in the rate.

Should I pick the B300 or the B200?

The B300 has 288GB against 180GB and 13.5 PFLOPS of dense FP4 against 9. FP8 and FP16 run at the same rate on both. The B200 keeps 37 TFLOPS of FP64 where the B300 has 1.25, and costs less per hour.

Why 288GB when NVIDIA lists 2.1TB for the HGX B300?

288GB per GPU is NVIDIA's own per-GPU figure, and its enterprise reference architecture gives 2.3TB per eight-GPU node. The 2.1TB on the HGX page does not match its per-GPU spec. 288GB is what your container sees.

Will my CUDA container run on the B300?

Yes, if it is built against CUDA 12.9 or later. That is the release that added the B300's architecture (sm_103); an image built for an older CUDA cannot compile kernels for it.

How do I get B300 capacity?

Book a call. B300 nodes are allocated with our team: we size the GPU count and the term with you, and you run pods on them like on any other GPU we sell.

B300 capacity, sized with you

We allocate B300 nodes with you: how many GPUs, for how long, and on what terms. Book a call and we will size it.

See all GPUs