DeepSeek V4-Flash, behind an OpenAI-compatible endpoint
DeepSeek's V4-Flash is a 284B mixture-of-experts that activates 13B parameters per token, with reasoning and tool calling under an MIT license. Call it with the OpenAI SDK and pay per token.
- 284B MoE, 13B active
- 128K context
- Tool calling
- MIT license
Per token, nothing else.
Input / output per 1M tokens
$0.22 / $0.66
openrelay/deepseek-v4-flashFull catalog on the inference pricing page. Add a card and make a deposit to start.
from openai import OpenAI
client = OpenAI(
base_url="https://inference.openrelay.inc/v1",
api_key="or_••••••••", # same SDK, new base URL
)
resp = client.chat.completions.create(
model="openrelay/deepseek-v4-flash",
messages=[{"role": "user", "content": "What is the weather in Lisbon?"}],
tools=[{"type": "function", "function": {
"name": "get_weather",
"parameters": {"type": "object",
"properties": {"city": {"type": "string"}}},
}}],
)
print(resp.choices[0].message.tool_calls)The spec sheet.
What the model is, what it takes in, and the surface it serves on.
- Model id
- openrelay/deepseek-v4-flash
- Parameters
- 284B total, 13B active per token (mixture-of-experts)
- Checkpoint
- DeepSeek-V4-Flash-0731, the official release that replaced the preview, served at NVFP4
- Context
- 128K tokens (131,072) on this endpoint, prompt and completion combined. The model supports up to 1M.
- Modality
- Text in, text out
- Endpoint
- POST /v1/chat/completions (streaming and non-streaming), tool calling, reasoning
- License
- MIT, open weights
Where DeepSeek V4-Flash earns its place.
What this model is actually for, versus the rest of the catalog.
13B active out of 284B
Each token routes through 13B of the 284B parameters, so compute per token tracks the active count rather than the total.
Tuned for agent work
DeepSeek built the 0731 release around agentic capability, and its published results are coding-agent and tool-use benchmarks. Tool calls and reasoning both run through chat completions.
MIT weights
The checkpoint is MIT-licensed on Hugging Face. Build against the API now and keep the option to serve the same weights yourself.
DeepSeek V4-Flash, answered.
Is there a DeepSeek V4-Flash API?
Yes: openrelay/deepseek-v4-flash, on the OpenAI-compatible chat completions endpoint at inference.openrelay.inc/v1. The OpenAI SDK works with a new base URL and an OpenRelay key.
What does DeepSeek V4-Flash cost?
The rate on this page is read from the catalog when the page renders. Input and output tokens bill separately, and a reasoning trace counts as output.
How much context can I send?
131,072 tokens (128K) on this endpoint, prompt and completion combined. DeepSeek trained V4-Flash for up to 1M; this deployment runs a shorter window.
Can I self-host DeepSeek V4-Flash?
The weights are MIT-licensed, so yes. DeepSeek's own vLLM example serves the 0731 release on a node of four GB300 GPUs, which makes self-hosting a multi-GPU job. The API covers you until volume justifies that hardware.
First request in five lines.
Point the OpenAI SDK at inference.openrelay.inc/v1 and run DeepSeek V4-Flash per token. No contract, no minimums.