Serverless inference for open models

Send a request, get tokens back, pay for the tokens. The models are already deployed, so there is no instance to size, no endpoint to warm up and no charge while your app is idle.

  • From $0.05 per 1M input tokens
  • No hourly charge
  • OpenAI-compatible
  • Standard batch at 50% off

What serverless means here.

Three properties, true of every model in the catalog.

Nothing to deploy

Catalog models are running before you call them; you pick one by id.

Billed per token

Every response's usage object carries the token counts you are billed on.

Idle is free

No hourly charge and no reserved capacity, so a quiet hour costs nothing.

Per-token rates.

Read from the catalog when this page loads. A standard batch bills each request at half.

Inference pricing
ModelContextInput / 1MOutput / 1MUse it for
GPT-OSS 120Bopenrelay/gpt-oss-120b128K$0.15$0.60Agents, tool use, complex reasoning
GPT-OSS 20Bopenrelay/gpt-oss-20b128K$0.05$0.20High-throughput, low-latency tasks
Gemma 4 31B NVFP4 32Kopenrelay/gemma-4-31b-32k32K$0.20$0.90Low-latency, high-volume chat
Gemma 4 31Bopenrelay/gemma-4-31b32K$0.99$1.49Vision-grounded chat
GLM 5.2openrelay/glm-5.21M$1.82$5.72Reasoning, coding, bilingual agents
GLM 5.3 Flashopenrelay/glm-5.3-flash256K$0.195$0.65Agentic coding, long context, vision
DeepSeek V4-Flashopenrelay/deepseek-v4-flash128K$0.22$0.66Reasoning, agents, long-output generation
Qwen3.8 27Bopenrelay/qwen3.8-27b256K$0.42$3.00Reasoning, agents, tool use
DeepSeek-OCR 2openrelay/deepseek-ocr-28K$0.039$0.039Document to markdown, OCR, data extraction
BGE-M3openrelay/bge-m38K$0.013Input onlyEmbeddings

When your own GPU is cheaper.

A rented GPU beats per-token pricing once you keep it busy enough. This is how busy, from the same catalog rates.

ModelRuns on oneGPU for 720 hoursBreak-even output
GPT-OSS 20B$0.20 per 1M output on the APIRTX 4090$252486 tokens/s
GPT-OSS 120B$0.60 per 1M output on the APIH100 SXM$1,8721,204 tokens/s

Break-even is the output rate, averaged over all 720 hours, at which the GPU costs what the API charges for output alone. Prompts bill on the API too, which lowers that bar; idle hours raise the rate the GPU must hold while busy.

What it does not cover.

Serverless here means catalog models on our capacity.

Your own weights

The shared API serves the catalog only. Run your model on a GPU VM or a Pod, or ask about a dedicated endpoint.

Scale to zero for your model

A VM or Pod bills while it runs and stops billing when you stop it; nothing starts it on a request. Serverless GPU has the math.

Streaming works the same way.

Set stream=True. The last chunk carries the usage you are billed on.

Streaming docs
stream.pypython
from openai import OpenAI

client = OpenAI(base_url="https://inference.openrelay.inc/v1", api_key="or_••••••••")

stream = client.chat.completions.create(
    model="openrelay/gpt-oss-120b",
    messages=[{"role": "user", "content": "Write a haiku about GPUs."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="")
    elif chunk.usage:
        print("\n", chunk.usage)  # the last chunk: what this request bills

Serverless inference, answered.

What is serverless inference?

Running a model through an API where the provider owns the servers. You pay per request or per token, and nothing bills between requests. The alternative is renting the GPU and paying for every hour it is up, busy or not.

Is there a cold start?

Not on catalog models. They stay deployed, so a request goes to a model that is already loaded. Cold starts come with platforms that scale your own container to zero and load it again on the next request.

How is serverless inference billed here?

Per token, from a prepaid balance. Each model has an input and an output rate, plus a cached-input rate where it has one: GPT-OSS 20B is $0.05 in and $0.20 out per 1M tokens. A failed request is not billed.

Can I deploy my own model as a serverless endpoint?

No. The shared API serves the catalog, and we do not host custom models that scale to zero. Run your model on a GPU VM or a Pod, billed while it runs, or ask about a dedicated endpoint.

Serverless or batch?

Use the realtime API when someone is waiting for the answer. Send the rest through the Batch API: same models, results within 24 hours, half the token rate on the standard tier.

Which SDKs work?

OpenAI SDKs for chat completions and embeddings, and Anthropic SDKs on models that support the Messages API. Change the base URL and the key; the rest of your code stays.

Try it on your own prompts.

Make a deposit, create a key, and point your SDK at inference.openrelay.inc.