Serverless inference for open models
Send a request, get tokens back, pay for the tokens. The models are already deployed, so there is no instance to size, no endpoint to warm up and no charge while your app is idle.
- From $0.05 per 1M input tokens
- No hourly charge
- OpenAI-compatible
- Standard batch at 50% off
What serverless means here.
Three properties, true of every model in the catalog.
Nothing to deploy
Catalog models are running before you call them; you pick one by id.
Billed per token
Every response's usage object carries the token counts you are billed on.
Idle is free
No hourly charge and no reserved capacity, so a quiet hour costs nothing.
Per-token rates.
Read from the catalog when this page loads. A standard batch bills each request at half.
Inference pricing| Model | Context | Input / 1M | Output / 1M | Use it for |
|---|---|---|---|---|
| GPT-OSS 120Bopenrelay/gpt-oss-120b | 128K | $0.15 | $0.60 | Agents, tool use, complex reasoning |
| GPT-OSS 20Bopenrelay/gpt-oss-20b | 128K | $0.05 | $0.20 | High-throughput, low-latency tasks |
| Gemma 4 31B NVFP4 32Kopenrelay/gemma-4-31b-32k | 32K | $0.20 | $0.90 | Low-latency, high-volume chat |
| Gemma 4 31Bopenrelay/gemma-4-31b | 32K | $0.99 | $1.49 | Vision-grounded chat |
| GLM 5.2openrelay/glm-5.2 | 1M | $1.82 | $5.72 | Reasoning, coding, bilingual agents |
| GLM 5.3 Flashopenrelay/glm-5.3-flash | 256K | $0.195 | $0.65 | Agentic coding, long context, vision |
| DeepSeek V4-Flashopenrelay/deepseek-v4-flash | 128K | $0.22 | $0.66 | Reasoning, agents, long-output generation |
| Qwen3.8 27Bopenrelay/qwen3.8-27b | 256K | $0.42 | $3.00 | Reasoning, agents, tool use |
| DeepSeek-OCR 2openrelay/deepseek-ocr-2 | 8K | $0.039 | $0.039 | Document to markdown, OCR, data extraction |
| BGE-M3openrelay/bge-m3 | 8K | $0.013 | Input only | Embeddings |
When your own GPU is cheaper.
A rented GPU beats per-token pricing once you keep it busy enough. This is how busy, from the same catalog rates.
| Model | Runs on one | GPU for 720 hours | Break-even output |
|---|---|---|---|
| GPT-OSS 20B$0.20 per 1M output on the API | RTX 4090 | $252 | 486 tokens/s |
| GPT-OSS 120B$0.60 per 1M output on the API | H100 SXM | $1,872 | 1,204 tokens/s |
Break-even is the output rate, averaged over all 720 hours, at which the GPU costs what the API charges for output alone. Prompts bill on the API too, which lowers that bar; idle hours raise the rate the GPU must hold while busy.
What it does not cover.
Serverless here means catalog models on our capacity.
Your own weights
The shared API serves the catalog only. Run your model on a GPU VM or a Pod, or ask about a dedicated endpoint.
Scale to zero for your model
A VM or Pod bills while it runs and stops billing when you stop it; nothing starts it on a request. Serverless GPU has the math.
Streaming works the same way.
Set stream=True. The last chunk carries the usage you are billed on.
Streaming docsfrom openai import OpenAI
client = OpenAI(base_url="https://inference.openrelay.inc/v1", api_key="or_••••••••")
stream = client.chat.completions.create(
model="openrelay/gpt-oss-120b",
messages=[{"role": "user", "content": "Write a haiku about GPUs."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")
elif chunk.usage:
print("\n", chunk.usage) # the last chunk: what this request billsServerless inference, answered.
What is serverless inference?
Running a model through an API where the provider owns the servers. You pay per request or per token, and nothing bills between requests. The alternative is renting the GPU and paying for every hour it is up, busy or not.
Is there a cold start?
Not on catalog models. They stay deployed, so a request goes to a model that is already loaded. Cold starts come with platforms that scale your own container to zero and load it again on the next request.
How is serverless inference billed here?
Per token, from a prepaid balance. Each model has an input and an output rate, plus a cached-input rate where it has one: GPT-OSS 20B is $0.05 in and $0.20 out per 1M tokens. A failed request is not billed.
Can I deploy my own model as a serverless endpoint?
No. The shared API serves the catalog, and we do not host custom models that scale to zero. Run your model on a GPU VM or a Pod, billed while it runs, or ask about a dedicated endpoint.
Serverless or batch?
Use the realtime API when someone is waiting for the answer. Send the rest through the Batch API: same models, results within 24 hours, half the token rate on the standard tier.
Which SDKs work?
OpenAI SDKs for chat completions and embeddings, and Anthropic SDKs on models that support the Messages API. Change the base URL and the key; the rest of your code stays.
Try it on your own prompts.
Make a deposit, create a key, and point your SDK at inference.openrelay.inc.