Qwen 4, served as an API
Alibaba's next Qwen generation, announced at its Apsara conference and reported in four tiers from Max down to a 27B. Every model here takes the same OpenAI SDK call, so adopting it would be a one-string change.
- Confirmed: in training
- Reported: four tiers
- Preview architecture is public
- Per token, no contract
jk. It is not out yet.
Alibaba has not shipped Qwen 4. Its own announcement says the model is in training and gives no date, sizes, benchmarks, or price. If you searched for a Qwen 4 API, nobody is serving one.
Below is every rumor we could source, what Alibaba has actually confirmed, and the model you can call in the meantime. When the real one ships, this page becomes its page at this URL.
The Qwen 4 rumor mill.
Each row labelled by how much it is worth. Last checked October 4, 2026.
- Release status
- ConfirmedUnreleased. Alibaba's Apsara press release of September 22, 2026 says Qwen 4 is in training, and gives no date, sizes, license, or price. No Qwen 4 weights are on Hugging Face. Alibaba, September 22, 2026
- Lineup
- ReportedCoverage of the Apsara briefing describes four Qwen 4 tiers: Max, Plus, Flash, and a 27B. Alibaba's English press release names none of them, and none has a spec sheet. Pandaily
- Architecture
- ConfirmedAlready public. Alibaba's model card for Qwen3.8-Flash-Next calls it an experimental preview of the architecture that will underpin Qwen 4: 125B parameters with 6B active, hybrid linear and sparse attention, and n-gram embeddings. Qwen3.8-Flash-Next model card
- After Qwen 4
- ConfirmedAlibaba's roadmap projects the Qwen 4.5 and Qwen 5 series at 5 to 10 trillion parameters. That describes the generations after this one, not Qwen 4. Alibaba, September 22, 2026
- Timing
- No signalAlibaba has published no date or window. A release date quoted anywhere else is inferred from past cadence, not announced.
- License
- ConfirmedNot announced for any Qwen 4 tier. The current family is mixed: Qwen3.8-27B is Apache 2.0, while Qwen3.8-Flash-Next ships under the Qwen Community License. Qwen3.8-27B model card
- How it would serve here
- ConfirmedSelf-hosted on our own GPUs, the same path Qwen3.8 27B runs on today, for any tier that ships with open weights. No upstream API provider sits between us and the weights.
- Price
- No signalUnknown until we have the model and can measure what it costs to serve. Anyone quoting one is guessing.
What you can call right now.
The Alibaba models live on the API today, at the rates they bill at.
Full inference catalogWhy it shows up here early.
Not a promise about a date. A description of the path a new open model takes to get behind this API.
We already run Qwen
Qwen3.8 27B serves on MI355X GPUs we operate, from Alibaba's official FP8 checkpoint. A new Qwen lands on a serving path that exists.
The architecture already runs in vLLM
Qwen3.8-Flash-Next, the public preview of the Qwen 4 design, ships with vLLM and SGLang recipes. Engine support is the slow part of serving a new architecture, and for this one it is already public.
Nothing to migrate when it arrives
Same base URL, same OpenAI SDK, same API key. Moving from Qwen3.8 27B to a Qwen 4 model means changing the model string.
One email, the day it is callable.
Leave an address and we send exactly one message when Qwen 4 is serving here, with the model id and the rate. Nothing else goes to this list.
Qwen 4, answered.
Is Qwen 4 out?
No. Alibaba said on September 22, 2026 that Qwen 4 is in training, and there are no weights, API, or price yet. Qwen3.8 27B is the Qwen model on OpenRelay today.
What did Alibaba announce at Apsara?
That Qwen 4 is in training, and that the Qwen 4.5 and Qwen 5 series are projected to reach 5 to 10 trillion parameters. Coverage of the briefing adds four Qwen 4 tiers (Max, Plus, Flash, and a 27B) that the English press release does not list.
When is the Qwen 4 release date?
Unannounced. Alibaba gave no date or window, so any date you see is someone's estimate. When a Qwen 4 model serves here, this URL becomes its page and the notify list gets one email.
Will OpenRelay serve Qwen 4?
Only a tier with open weights, because we serve models we run ourselves. We already self-host Qwen3.8 27B, so the serving path exists, but Alibaba has not said which Qwen 4 tiers will get open weights.
What is Qwen3.8-Flash-Next?
An open-weight model Alibaba published in August 2026 and describes as an experimental preview of the Qwen 4 architecture, with 125B parameters and 6B active. It is not Qwen 4, and it does not serve here.
What should I use until then?
Qwen3.8 27B, on the card above. It thinks by default, calls tools, and takes a 256K context through the standard OpenAI SDK call.
The account outlasts the model.
One API key, one base URL, per-token billing. Whatever Alibaba ships next, adopting it is a one-line change rather than a new vendor.