Models/ DeepSeek

DeepSeek V4-Flash, behind an OpenAI-compatible endpoint

DeepSeek's V4-Flash is a 284B mixture-of-experts that activates 13B parameters per token, with reasoning and tool calling under an MIT license. Call it with the OpenAI SDK and pay per token.

  • 284B MoE, 13B active
  • 128K context
  • Tool calling
  • MIT license

Per token, nothing else.

Input / output per 1M tokens

$0.22 / $0.66

openrelay/deepseek-v4-flash

Full catalog on the inference pricing page. Add a card and make a deposit to start.

tools.py · function callingpython
from openai import OpenAI

client = OpenAI(
    base_url="https://inference.openrelay.inc/v1",
    api_key="or_••••••••",          # same SDK, new base URL
)

resp = client.chat.completions.create(
    model="openrelay/deepseek-v4-flash",
    messages=[{"role": "user", "content": "What is the weather in Lisbon?"}],
    tools=[{"type": "function", "function": {
        "name": "get_weather",
        "parameters": {"type": "object",
                       "properties": {"city": {"type": "string"}}},
    }}],
)
print(resp.choices[0].message.tool_calls)

The spec sheet.

What the model is, what it takes in, and the surface it serves on.

Model id
openrelay/deepseek-v4-flash
Parameters
284B total, 13B active per token (mixture-of-experts)
Checkpoint
DeepSeek-V4-Flash-0731, the official release that replaced the preview, served at NVFP4
Context
128K tokens (131,072) on this endpoint, prompt and completion combined. The model supports up to 1M.
Modality
Text in, text out
Endpoint
POST /v1/chat/completions (streaming and non-streaming), tool calling, reasoning
License
MIT, open weights

Where DeepSeek V4-Flash earns its place.

What this model is actually for, versus the rest of the catalog.

13B active out of 284B

Each token routes through 13B of the 284B parameters, so compute per token tracks the active count rather than the total.

Tuned for agent work

DeepSeek built the 0731 release around agentic capability, and its published results are coding-agent and tool-use benchmarks. Tool calls and reasoning both run through chat completions.

MIT weights

The checkpoint is MIT-licensed on Hugging Face. Build against the API now and keep the option to serve the same weights yourself.

DeepSeek V4-Flash, answered.

Is there a DeepSeek V4-Flash API?

Yes: openrelay/deepseek-v4-flash, on the OpenAI-compatible chat completions endpoint at inference.openrelay.inc/v1. The OpenAI SDK works with a new base URL and an OpenRelay key.

What does DeepSeek V4-Flash cost?

The rate on this page is read from the catalog when the page renders. Input and output tokens bill separately, and a reasoning trace counts as output.

How much context can I send?

131,072 tokens (128K) on this endpoint, prompt and completion combined. DeepSeek trained V4-Flash for up to 1M; this deployment runs a shorter window.

Can I self-host DeepSeek V4-Flash?

The weights are MIT-licensed, so yes. DeepSeek's own vLLM example serves the 0731 release on a node of four GB300 GPUs, which makes self-hosting a multi-GPU job. The API covers you until volume justifies that hardware.

First request in five lines.

Point the OpenAI SDK at inference.openrelay.inc/v1 and run DeepSeek V4-Flash per token. No contract, no minimums.