Serverless GPU, and when an hourly GPU is cheaper

A serverless GPU runs your container when requests arrive and bills the seconds it runs. OpenRelay does not sell one; this page covers when they pay off and what we sell for the same job.

  • GPUs from $0.18/hr
  • Billed only while running
  • Per-token API, no GPU to run

How serverless GPUs work.

Products sold as serverless GPU share three traits.

Workers start on demand

A request arrives, the platform starts your container on a free GPU, and the container loads its model.

Scale to zero

When requests stop, workers shut down, and you pay for the seconds they ran.

Cold starts

A request that finds no running worker waits for one to start; a worker kept warm avoids that and bills while it waits.

What we do not do.

The serverless features OpenRelay does not have.

Start on a request

Nothing starts a VM or Pod when traffic arrives; your code or your team starts and stops it.

Scale with traffic

You run as many VMs and Pods as you start, and no more.

Hold a stopped Pod's GPUs

A Pod restarts on the node that holds its volume, and if those GPUs are taken the start is refused until they free up.

When per-second billing wins.

It comes down to k: what a serverless GPU costs per hour of runtime, over the hourly rate for the same GPU. Serverless is cheaper while the GPU would be busy less than 1/k of the time.

kServerless is cheaper belowBusy hours in 720
1.5x67% busy480 hours
2x50% busy360 hours
3x33% busy240 hours
4x25% busy180 hours

To find k, multiply a platform's per-second price by 3,600 and divide by the hourly rate for the same GPU below. Stopping an hourly GPU while it is idle cuts its bill to the hours it runs, which lowers the bar further.

Hourly rates.

Read from the catalog when this page loads. Pods run on the GPUs marked Pod; the rest rent as VMs.

All GPUs
GPUVRAMRuns asPer hour720 hours
NVIDIA RTX 309024GB GDDR6XVM$0.18/hr$130
NVIDIA RTX 409024GB GDDR6XVM$0.35/hr$252
NVIDIA RTX 509032GB GDDR7VM$0.60/hr$432
NVIDIA A100 80GB80GB HBM2eVM$1.10/hr$792
NVIDIA RTX Pro 600096GB GDDR7VM$1.55/hr$1,116
NVIDIA H100 SXM80GB HBM3VM$2.60/hr$1,872
NVIDIA B200180GB HBM3ePod$5.50/hr$3,960
AMD Instinct MI355X288GB HBM3EPod$5.50/hr$3,960
NVIDIA B300288GB HBM3ePod$7.49/hr$5,393

Serverless GPUs, answered.

What is a serverless GPU?

A GPU you do not rent directly. You give the platform a container; it runs the container on a GPU while requests arrive, stops it when they do not, and bills the seconds it ran.

Does OpenRelay offer serverless GPUs?

No. Nothing here starts a GPU in response to a request or scales a container to zero. We sell a per-token inference API for catalog models, and GPU VMs and Pods that bill while running and stop billing when stopped.

When is a serverless GPU cheaper than an hourly GPU?

When the GPU would sit idle most of the time. At a serverless rate twice the hourly rate for the same GPU, serverless wins below 50% utilization and loses above it.

What causes a cold start?

Starting the container and loading the model weights into GPU memory. Bigger images and bigger models take longer, which is why platforms offer warm workers that bill for as long as they stay warm.

How do I avoid paying for an idle GPU on OpenRelay?

If your model is in the catalog, use the per-token API. If not, stop the VM or Pod when it is idle, from the dashboard, with orl vms stop, or with POST /v1/vms/{id}/stop, and start it when you need it again. Stopped time is not billed.

Can I serve my own model behind an API here?

Yes, on a GPU VM or a Pod: run an inference server such as vLLM or SGLang and reach it over HTTPS. It bills while it runs. For a model served on capacity reserved for your organization, ask about a dedicated endpoint.

Start with the API or a GPU.

The API bills tokens. A GPU bills the hours it runs.