Qwen 3.8 27B, dense, with a 256K window
Qwen3.8-27B is a dense 27B model from Alibaba's Qwen3.8 generation: Apache 2.0, a 262,144-token native context, and thinking on by default. This endpoint serves it text in, text out.
- Dense 27B
- 256K context
- Tool calling
- Apache 2.0
Per token, nothing else.
Input / output per 1M tokens
$0.42 / $3.00
openrelay/qwen3.8-27bFull catalog on the inference pricing page. Add a card and make a deposit to start.
from openai import OpenAI
client = OpenAI(
base_url="https://inference.openrelay.inc/v1",
api_key="or_••••••••", # same SDK, new base URL
)
resp = client.chat.completions.create(
model="openrelay/qwen3.8-27b",
messages=[{"role": "user", "content":
"Why is this wrong? def mean(xs): return sum(xs) / len(xs) - 1"}],
max_tokens=8192, # room for the reasoning trace and the answer
)
print(resp.choices[0].message.content)The spec sheet.
What the model is, what it takes in, and the surface it serves on.
- Model id
- openrelay/qwen3.8-27b
- Parameters
- 27B, dense
- Architecture
- Gated DeltaNet linear attention with a gated attention layer after every three, 64 layers
- Modality
- Text in, text out. The checkpoint includes Qwen's vision encoder; image input is not enabled on this endpoint.
- Context
- 256K tokens (262,144), prompt and completion combined
- Endpoint
- POST /v1/chat/completions with tool calling and reasoning
- Precision
- FP8
- License
- Apache 2.0, open weights
Where Qwen3.8 27B earns its place.
What this model is actually for, versus the rest of the catalog.
Thinking on by default
Qwen3.8 writes a reasoning trace before it answers unless told otherwise, and that trace bills as output. Leave room for it in max_tokens.
Small enough to take with you
At 27B parameters the model fits on a single data-center GPU, and the weights are Apache 2.0. Leaving the API later means one GPU, not a cluster.
The full native window
262,144 tokens is Qwen's native context for this model, and the endpoint serves all of it: room for long tool loops and large code context.
Qwen3.8 27B, answered.
Is there a Qwen 3.8 27B API?
Yes: openrelay/qwen3.8-27b, on the OpenAI-compatible chat completions endpoint at inference.openrelay.inc/v1, with tool calling. The OpenAI SDK works with a new base URL and an OpenRelay key.
What does Qwen3.8-27B cost?
The rate on this page is read from the catalog when the page renders. Input and output tokens bill separately, and the reasoning trace counts as output.
Does this endpoint accept images?
No. Qwen3.8-27B is a vision-language model, but this endpoint is configured for text only. For image input, use GLM 5.3 Flash or Gemma 4 31B.
What license are the weights under?
Apache 2.0, which allows commercial use. The same weights are on Hugging Face if you later want to serve them yourself.
First request in five lines.
Point the OpenAI SDK at inference.openrelay.inc/v1 and run Qwen3.8 27B per token. No contract, no minimums.