Serverless GPU, and when an hourly GPU is cheaper
A serverless GPU runs your container when requests arrive and bills the seconds it runs. OpenRelay does not sell one; this page covers when they pay off and what we sell for the same job.
- GPUs from $0.18/hr
- Billed only while running
- Per-token API, no GPU to run
How serverless GPUs work.
Products sold as serverless GPU share three traits.
Workers start on demand
A request arrives, the platform starts your container on a free GPU, and the container loads its model.
Scale to zero
When requests stop, workers shut down, and you pay for the seconds they ran.
Cold starts
A request that finds no running worker waits for one to start; a worker kept warm avoids that and bills while it waits.
What OpenRelay sells instead.
One takes the GPU off your bill. The other is a GPU you rent by the hour and stop when you are done.
What we do not do.
The serverless features OpenRelay does not have.
Start on a request
Nothing starts a VM or Pod when traffic arrives; your code or your team starts and stops it.
Scale with traffic
You run as many VMs and Pods as you start, and no more.
Hold a stopped Pod's GPUs
A Pod restarts on the node that holds its volume, and if those GPUs are taken the start is refused until they free up.
When per-second billing wins.
It comes down to k: what a serverless GPU costs per hour of runtime, over the hourly rate for the same GPU. Serverless is cheaper while the GPU would be busy less than 1/k of the time.
| k | Serverless is cheaper below | Busy hours in 720 |
|---|---|---|
| 1.5x | 67% busy | 480 hours |
| 2x | 50% busy | 360 hours |
| 3x | 33% busy | 240 hours |
| 4x | 25% busy | 180 hours |
To find k, multiply a platform's per-second price by 3,600 and divide by the hourly rate for the same GPU below. Stopping an hourly GPU while it is idle cuts its bill to the hours it runs, which lowers the bar further.
Hourly rates.
Read from the catalog when this page loads. Pods run on the GPUs marked Pod; the rest rent as VMs.
All GPUs| GPU | VRAM | Runs as | Per hour | 720 hours |
|---|---|---|---|---|
| NVIDIA RTX 3090 | 24GB GDDR6X | VM | $0.18/hr | $130 |
| NVIDIA RTX 4090 | 24GB GDDR6X | VM | $0.35/hr | $252 |
| NVIDIA RTX 5090 | 32GB GDDR7 | VM | $0.60/hr | $432 |
| NVIDIA A100 80GB | 80GB HBM2e | VM | $1.10/hr | $792 |
| NVIDIA RTX Pro 6000 | 96GB GDDR7 | VM | $1.55/hr | $1,116 |
| NVIDIA H100 SXM | 80GB HBM3 | VM | $2.60/hr | $1,872 |
| NVIDIA B200 | 180GB HBM3e | Pod | $5.50/hr | $3,960 |
| AMD Instinct MI355X | 288GB HBM3E | Pod | $5.50/hr | $3,960 |
| NVIDIA B300 | 288GB HBM3e | Pod | $7.49/hr | $5,393 |
Serverless GPUs, answered.
What is a serverless GPU?
A GPU you do not rent directly. You give the platform a container; it runs the container on a GPU while requests arrive, stops it when they do not, and bills the seconds it ran.
Does OpenRelay offer serverless GPUs?
No. Nothing here starts a GPU in response to a request or scales a container to zero. We sell a per-token inference API for catalog models, and GPU VMs and Pods that bill while running and stop billing when stopped.
When is a serverless GPU cheaper than an hourly GPU?
When the GPU would sit idle most of the time. At a serverless rate twice the hourly rate for the same GPU, serverless wins below 50% utilization and loses above it.
What causes a cold start?
Starting the container and loading the model weights into GPU memory. Bigger images and bigger models take longer, which is why platforms offer warm workers that bill for as long as they stay warm.
How do I avoid paying for an idle GPU on OpenRelay?
If your model is in the catalog, use the per-token API. If not, stop the VM or Pod when it is idle, from the dashboard, with orl vms stop, or with POST /v1/vms/{id}/stop, and start it when you need it again. Stopped time is not billed.
Can I serve my own model behind an API here?
Yes, on a GPU VM or a Pod: run an inference server such as vLLM or SGLang and reach it over HTTPS. It bills while it runs. For a model served on capacity reserved for your organization, ask about a dedicated endpoint.
Start with the API or a GPU.
The API bills tokens. A GPU bills the hours it runs.