An AI inference platform for open models

OpenRelay serves open-weight models behind an OpenAI-compatible API and bills per token. When you need your own model or runtime, the same account rents GPUs by the hour.

  • OpenAI-compatible API
  • From $0.05 per 1M input tokens
  • Standard batch at 50% off
  • GPUs from $0.18/hr

The models and their rates.

Read from the catalog when this page loads. A standard batch bills each request at half.

Inference pricing
ModelContextInput / 1MOutput / 1MUse it for
GPT-OSS 120Bopenrelay/gpt-oss-120b128K$0.15$0.60Agents, tool use, complex reasoning
GPT-OSS 20Bopenrelay/gpt-oss-20b128K$0.05$0.20High-throughput, low-latency tasks
Gemma 4 31B NVFP4 32Kopenrelay/gemma-4-31b-32k32K$0.20$0.90Low-latency, high-volume chat
Gemma 4 31Bopenrelay/gemma-4-31b32K$0.99$1.49Vision-grounded chat
GLM 5.2openrelay/glm-5.21M$1.82$5.72Reasoning, coding, bilingual agents
GLM 5.3 Flashopenrelay/glm-5.3-flash256K$0.195$0.65Agentic coding, long context, vision
DeepSeek V4-Flashopenrelay/deepseek-v4-flash128K$0.22$0.66Reasoning, agents, long-output generation
Qwen3.8 27Bopenrelay/qwen3.8-27b256K$0.42$3.00Reasoning, agents, tool use
DeepSeek-OCR 2openrelay/deepseek-ocr-28K$0.039$0.039Document to markdown, OCR, data extraction
BGE-M3openrelay/bge-m38K$0.013Input onlyEmbeddings

What the gateway does per request.

Every call to inference.openrelay.inc passes through one gateway before it reaches a model.

How routing works

Authenticates the key

An or_ key, sent as a Bearer token or, for Anthropic clients, as x-api-key.

Speaks your SDK's format

OpenAI Chat Completions and Embeddings, and Anthropic Messages on models that support it.

Retries before the first byte

Where a model has more than one backend, a failure before the response starts moves the request to the next.

Prices cached prompts separately

Prompt tokens served from cache bill at the model's cached-input rate.

Bills what it reports

Every response carries the token counts you pay for. A failed request is not billed.

Checks the balance first

Requests draw on a prepaid balance and are refused once it falls below the floor.

Which one fits.

Pick by who is waiting for the answer and whose model it is.

WorkloadUseYou pay for
Chat, agents and RAG on a catalog modelPer-token APIInput and output tokens
Evals, labeling, backfills nobody waits onBatch APITokens at half the rate on the standard tier
Steady traffic that needs capacity of its ownDedicated endpointThe endpoint's own price
Your own weights, fine-tune or runtimeGPU VM or PodAn hourly rate while it runs

We do not sell scale-to-zero hosting for your own model. The serverless GPU page covers when it pays and what to use here instead.

Two lines change.

Existing OpenAI code runs against OpenRelay with a new base URL and an OpenRelay key.

Quickstart
chat.pypython
from openai import OpenAI

client = OpenAI(
    base_url="https://inference.openrelay.inc/v1",
    api_key="or_••••••••",
)

reply = client.chat.completions.create(
    model="openrelay/gpt-oss-120b",
    messages=[{"role": "user", "content": "Summarize this ticket in one line."}],
)
print(reply.choices[0].message.content)

AI inference platforms, answered.

What is an AI inference platform?

The layer between a trained model and the application calling it: GPUs with the model loaded, a server that schedules requests onto them, an API in front and metering behind. Using one means you send requests while someone else keeps the model deployed.

Which models does OpenRelay serve?

Models on the API include GPT-OSS 120B, GPT-OSS 20B, Gemma 4 31B NVFP4 32K, Gemma 4 31B, GLM 5.2, GLM 5.3 Flash, DeepSeek V4-Flash, Qwen3.8 27B, DeepSeek-OCR 2, and BGE-M3. The table above has each one's context window and rates, and the inference catalog has the details.

How is OpenRelay priced?

API calls bill per token at each model's input and output rates, from $0.05 per 1M input tokens, plus a cached-input rate on models that have one. Standard batches bill at half the token rate, expedited batches at the full rate. GPU VMs and Pods bill an hourly rate, from $0.18/hr, for the time they run. All of it draws on one prepaid balance.

Can I serve my own model?

Not on the shared API, which serves the catalog. Run your weights on a GPU VM or a Pod with the serving stack you choose, or ask us about a dedicated endpoint.

Do I have to change my code?

Not for OpenAI clients: set the base URL to https://inference.openrelay.inc/v1 and use an OpenRelay key. Anthropic clients work the same way on models that support the Messages API.

What does an idle hour cost?

Nothing on the API, which has no hourly charge. A GPU VM or Pod bills while it runs, so stop it when it sits idle: stopped time is not billed.

One key for every model.

Make a deposit, call any model in the catalog, and rent a GPU from the same account when you need your own.