Models/ Zhipu

GLM 5.5, served as an API

Zhipu's next flagship is the most-anticipated open-weight release of the year, with a reported trillion-plus parameters and a 1M-token context window. Point the OpenAI SDK at our base URL and call it per token, no contract and no minimums.

  • Reported: 1T+ parameters
  • Reported: 1M-token context
  • Open weights expected
  • Per token, no contract

jk. It is not out yet.

Zhipu has not shipped GLM 5.5. There are no weights, no benchmarks, and no date, and the August 2026 window an analyst note pointed at has come and gone. If you searched for a GLM 5.5 API, nobody is serving one, here or anywhere.

Below is every rumor we could source, what Zhipu has actually confirmed, and the model you can call in the meantime. When the real one ships, this page becomes its page at this URL.

The GLM 5.5 rumor mill.

Each row labelled by how much it is worth. Last checked September 24, 2026.

Release status
ConfirmedUnreleased. The current GLM release is GLM 5.3, out August 14, 2026 on the same 743B base as GLM 5.2 with stronger post-training, alongside GLM 5.3 Flash. Z.ai release history
The 5.4 question
Unverified5.4 appears to be skipped. Coverage of Zhipu's roadmap jumps from 5.3 straight to 5.5, so a search for a GLM 5.4 API is most likely a search for this model.
Timing
ReportedA JPMorgan research note relayed by Reuters on June 25, 2026 put GLM 5.5 in an August 2026 window. That window has passed with no release, so treat the date as stale rather than imminent. (Reuters, June 25, 2026)
Size and context
ReportedMore than one trillion total parameters and a 1M-token context window, which would be a roughly 35% step up from GLM 5.2's 744B. No benchmarks and no official spec have been published. (Reuters, via a JPMorgan note)
Vision
ReportedZhipu co-founder Jie Tang ran a public developer poll on June 29, 2026 asking what the next GLM must have, and vision was the near-unanimous answer across 1,400+ replies. GLM 5.3 shipped text-only, so it is still open. Poll coverage
License
UnverifiedOpen weights on Hugging Face, expected to follow the MIT-licensed pattern of GLM 5 and GLM 5.2. Widely assumed, not confirmed by Zhipu.
How it would serve here
ConfirmedSelf-hosted on our own GPUs, the same path GLM 5.3 Flash runs on today. No upstream API provider sits between us and the weights.
Price
No signalUnknown until we have the model and can measure what it costs to serve. Anyone quoting one is guessing.

Why it shows up here early.

Not a promise about a date. A description of the path a new open model takes to get behind this API.

We run the weights, not someone else's API

GLM serves here on hardware we operate. When new weights land there is no partner to negotiate with and no upstream provider to wait on, which is most of the delay in getting a new open model behind an API.

Open weights make this a days problem

The work between a Hugging Face release and a served endpoint is engine support and capacity, both of which we already have standing for the GLM family. We have done it for every GLM release so far.

Nothing to migrate when it arrives

Same base URL, same OpenAI SDK, same API key. Adopting the next GLM means changing the model string in one line, which is also why it is worth being on the list.

One email, the day it is callable.

Leave an address and we send exactly one message when GLM 5.5 is serving here, with the model id and the rate. Nothing else goes to this list.

GLM 5.5, answered.

Is GLM 5.5 out?

No. As of September 2026 Zhipu has not released it, and nobody is serving it. GLM 5.3 and GLM 5.3 Flash are the current releases, and GLM 5.3 Flash and GLM 5.2 both run on OpenRelay today.

What happened to GLM 5.4?

It looks skipped. Zhipu shipped 5.2, then 5.3 and 5.3 Flash in August 2026, and the reporting on what comes next points at 5.5 rather than 5.4. If you searched for a GLM 5.4 API, this is the page for the model you are probably looking for.

When is the GLM 5.5 release date?

Unannounced. A JPMorgan note relayed by Reuters in June 2026 pointed at an August window, which passed without a release, so that date is stale. Zhipu announces on its own channels and publishes weights to Hugging Face; when that happens this page becomes the live model page and the notify list gets an email.

Will OpenRelay serve GLM 5.5?

Assuming it ships with open weights like every GLM before it, yes, and quickly. We self-host the GLM family on our own GPUs rather than reselling another provider's endpoint, so standing up a new release is our own engineering work rather than a partnership.

What should I use until then?

GLM 5.3 Flash for agentic coding, long context, and vision, or GLM 5.2 when you need the 1M-token window. Both are on the cards above at their live rates, both take the same OpenAI SDK call, and moving to the next GLM later means editing one string.

Why does this page exist if the model does not?

Because people search for the next model before it ships, and landing on nothing is worse than landing on a sourced answer. Every rumor above is labelled by what it is worth and carries its source, the model you can actually call is one click away, and this URL becomes the real model page at launch rather than redirecting you somewhere else.

The account outlasts the model.

One API key, one base URL, per-token billing. Whatever Zhipu ships next, adopting it is a one-line change rather than a new vendor.