An AI inference platform for open models
OpenRelay serves open-weight models behind an OpenAI-compatible API and bills per token. When you need your own model or runtime, the same account rents GPUs by the hour.
- OpenAI-compatible API
- From $0.05 per 1M input tokens
- Standard batch at 50% off
- GPUs from $0.18/hr
Four ways to run a model.
They differ in who runs the model and what you pay for. One account and one API key cover all four.
Per-token API
Call a catalog model, pay per token, and pay nothing while idle.
Serverless inferenceBatch API
Upload a JSONL file of requests and collect the results within 24 hours; standard batches bill at half the token rate.
Batch inferenceDedicated endpoint
A model on capacity reserved for your organization, with its own private model id and price.
Contact usYour own server
Run your weights and serving stack on a GPU VM or a Pod, billed hourly while it runs.
GPU pricingThe models and their rates.
Read from the catalog when this page loads. A standard batch bills each request at half.
Inference pricing| Model | Context | Input / 1M | Output / 1M | Use it for |
|---|---|---|---|---|
| GPT-OSS 120Bopenrelay/gpt-oss-120b | 128K | $0.15 | $0.60 | Agents, tool use, complex reasoning |
| GPT-OSS 20Bopenrelay/gpt-oss-20b | 128K | $0.05 | $0.20 | High-throughput, low-latency tasks |
| Gemma 4 31B NVFP4 32Kopenrelay/gemma-4-31b-32k | 32K | $0.20 | $0.90 | Low-latency, high-volume chat |
| Gemma 4 31Bopenrelay/gemma-4-31b | 32K | $0.99 | $1.49 | Vision-grounded chat |
| GLM 5.2openrelay/glm-5.2 | 1M | $1.82 | $5.72 | Reasoning, coding, bilingual agents |
| GLM 5.3 Flashopenrelay/glm-5.3-flash | 256K | $0.195 | $0.65 | Agentic coding, long context, vision |
| DeepSeek V4-Flashopenrelay/deepseek-v4-flash | 128K | $0.22 | $0.66 | Reasoning, agents, long-output generation |
| Qwen3.8 27Bopenrelay/qwen3.8-27b | 256K | $0.42 | $3.00 | Reasoning, agents, tool use |
| DeepSeek-OCR 2openrelay/deepseek-ocr-2 | 8K | $0.039 | $0.039 | Document to markdown, OCR, data extraction |
| BGE-M3openrelay/bge-m3 | 8K | $0.013 | Input only | Embeddings |
What the gateway does per request.
Every call to inference.openrelay.inc passes through one gateway before it reaches a model.
How routing worksAuthenticates the key
An or_ key, sent as a Bearer token or, for Anthropic clients, as x-api-key.
Speaks your SDK's format
OpenAI Chat Completions and Embeddings, and Anthropic Messages on models that support it.
Retries before the first byte
Where a model has more than one backend, a failure before the response starts moves the request to the next.
Prices cached prompts separately
Prompt tokens served from cache bill at the model's cached-input rate.
Bills what it reports
Every response carries the token counts you pay for. A failed request is not billed.
Checks the balance first
Requests draw on a prepaid balance and are refused once it falls below the floor.
Which one fits.
Pick by who is waiting for the answer and whose model it is.
| Workload | Use | You pay for |
|---|---|---|
| Chat, agents and RAG on a catalog model | Per-token API | Input and output tokens |
| Evals, labeling, backfills nobody waits on | Batch API | Tokens at half the rate on the standard tier |
| Steady traffic that needs capacity of its own | Dedicated endpoint | The endpoint's own price |
| Your own weights, fine-tune or runtime | GPU VM or Pod | An hourly rate while it runs |
We do not sell scale-to-zero hosting for your own model. The serverless GPU page covers when it pays and what to use here instead.
Two lines change.
Existing OpenAI code runs against OpenRelay with a new base URL and an OpenRelay key.
Quickstartfrom openai import OpenAI
client = OpenAI(
base_url="https://inference.openrelay.inc/v1",
api_key="or_••••••••",
)
reply = client.chat.completions.create(
model="openrelay/gpt-oss-120b",
messages=[{"role": "user", "content": "Summarize this ticket in one line."}],
)
print(reply.choices[0].message.content)AI inference platforms, answered.
What is an AI inference platform?
The layer between a trained model and the application calling it: GPUs with the model loaded, a server that schedules requests onto them, an API in front and metering behind. Using one means you send requests while someone else keeps the model deployed.
Which models does OpenRelay serve?
Models on the API include GPT-OSS 120B, GPT-OSS 20B, Gemma 4 31B NVFP4 32K, Gemma 4 31B, GLM 5.2, GLM 5.3 Flash, DeepSeek V4-Flash, Qwen3.8 27B, DeepSeek-OCR 2, and BGE-M3. The table above has each one's context window and rates, and the inference catalog has the details.
How is OpenRelay priced?
API calls bill per token at each model's input and output rates, from $0.05 per 1M input tokens, plus a cached-input rate on models that have one. Standard batches bill at half the token rate, expedited batches at the full rate. GPU VMs and Pods bill an hourly rate, from $0.18/hr, for the time they run. All of it draws on one prepaid balance.
Can I serve my own model?
Not on the shared API, which serves the catalog. Run your weights on a GPU VM or a Pod with the serving stack you choose, or ask us about a dedicated endpoint.
Do I have to change my code?
Not for OpenAI clients: set the base URL to https://inference.openrelay.inc/v1 and use an OpenRelay key. Anthropic clients work the same way on models that support the Messages API.
What does an idle hour cost?
Nothing on the API, which has no hourly charge. A GPU VM or Pod bills while it runs, so stop it when it sits idle: stopped time is not billed.
One key for every model.
Make a deposit, call any model in the catalog, and rent a GPU from the same account when you need your own.