All GPUs

AMD Instinct MI355X price: $5.50/hr

AMD's CDNA 4 accelerator: 288GB of HBM3E per GPU, 8 TB/s of memory bandwidth and 10.1 PFLOPS of MXFP4. On OpenRelay it runs your ROCm container as a pod, from one GPU to a full node of eight.

What the rate buys

Per GPU-hour
$5.50
billed while the pod runs
8-GPU node, per hour
$44.00
one pod, one node
One GPU, 24/7
$3,960
per month, 720 hours

The MI355X runs as a pod: your container image on 1 to 8 whole GPUs of a single node, with SSH access and an HTTPS endpoint. Every pod gets a persistent volume at /workspace, 10 to 1,024 GB, that adds nothing to the hourly rate and survives a stop or restart. A stopped pod is not billed. The details are in the pods documentation.

MI355X specifications

SpecificationMI355X
ArchitectureAMD CDNA 4
Memory288GB HBM3E
Memory bandwidth8 TB/s
MXFP4 and MXFP610.1 PFLOPS
MXFP85 PFLOPS
FP16 matrix2.5 PFLOPS (5 with structured sparsity)
FP6478.6 TFLOPS
Compute units256 (16,384 stream processors)
Infinity Fabric7 links, 153 GB/s each
Board power1,400W

Source: AMD Instinct MI355X product page, read 2026-10-01.

What the MI355X is for

Models too big for an 80GB GPU

A 405B-parameter model in FP8 is about 405GB of weights: two MI355X hold it with room left for the KV cache. An 80GB H100 needs six for the weights alone.

4-bit and 6-bit inference

MXFP4 and MXFP6 run at 10.1 PFLOPS, twice the MXFP8 rate, for models quantized below 8 bits.

Long context

288GB per GPU leaves the space a long KV cache needs, the part of serving that grows with context length rather than with model size.

FP64 work

78.6 TFLOPS of double precision, for simulation and scientific code that cannot drop to lower precision.

Frequently asked questions

How much does it cost to rent an AMD Instinct MI355X?

$5.50 per GPU-hour, billed while the pod is running; a stopped pod is not billed. A full node of eight is $44.00 an hour, about $31,680 a month around the clock. The /workspace volume is included in the rate.

Will my CUDA container run on the MI355X?

No. The MI355X runs AMD's ROCm stack on the amdgpu driver, so the image has to be a ROCm image: a CUDA image starts, but its container cannot see the GPU. rocm/pytorch and rocm/dev-ubuntu-22.04 at ROCm 7.2.4 are validated on our hardware, and the dashboard blocks a CUDA image on an AMD GPU before you deploy.

Does my PyTorch code need changes?

Usually not. PyTorch's ROCm build answers the same torch.cuda calls, so code that does not ship its own CUDA kernels runs as it is on a ROCm image. Extensions with hand-written CUDA kernels need a HIP port first.

Do I get the whole GPU?

Yes. A pod gets whole GPUs of one node, from 1 to 8, and compute partitions are not sold. Each GPU is granted to your container on its own, so rocm-smi inside the pod lists exactly the GPUs you pay for.

What does an MI355X pod not do yet?

There is no InfiniBand or RDMA on AMD pods, so a job cannot span nodes over a fast fabric. A full node of eight GPUs is the largest MI355X pod.

Run your container on the MI355X

$5.50/hr per GPU, from one GPU to a full node of eight.