GLM 5.3 Flash, text and images through one endpoint
Z.ai's GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series: 320B parameters, 18B active per token, MIT-licensed. Send text or images to chat completions and pay per token.
- 320B MoE, 18B active
- Image input
- 256K context
- MIT license
Per token, nothing else.
Input / output per 1M tokens
$0.195 / $0.65
openrelay/glm-5.3-flashFull catalog on the inference pricing page. Add a card and make a deposit to start.
from openai import OpenAI
client = OpenAI(
base_url="https://inference.openrelay.inc/v1",
api_key="or_••••••••", # same SDK, new base URL
)
import base64
img = base64.b64encode(open("screenshot.png", "rb").read()).decode()
resp = client.chat.completions.create(
model="openrelay/glm-5.3-flash",
messages=[{"role": "user", "content": [
{"type": "text", "text": "Which button submits this form?"},
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{img}"}},
]}],
)
print(resp.choices[0].message.content)The spec sheet.
What the model is, what it takes in, and the surface it serves on.
- Model id
- openrelay/glm-5.3-flash
- Parameters
- 320B total, 18B active per token (mixture-of-experts)
- Architecture
- Hybrid sparse and linear attention, which Z.ai designed to cut long-context serving cost
- Modality
- Text and image in, text out
- Context
- 256K tokens (262,144), prompt and completion combined
- Endpoint
- POST /v1/chat/completions with optional image parts, tool calling, reasoning
- Precision
- FP8
- License
- MIT, open weights
Where GLM 5.3 Flash earns its place.
What this model is actually for, versus the rest of the catalog.
Vision in the same weights
Images go in as message parts on the same model id. No separate vision model and no second endpoint to route between.
Built for long prompts
The hybrid attention design targets the cost of long context, which is where agent transcripts and repository-sized code prompts end up. This endpoint takes up to 256K tokens.
Next to GLM 5.2
GLM 5.2 in the same catalog serves a 1M-token window and the Anthropic Messages API. 5.3 Flash adds image input at 18B active. Switching between them is a model id.
GLM 5.3 Flash, answered.
Is GLM 5.3 Flash multimodal?
Yes. Z.ai trained it natively multimodal, and openrelay/glm-5.3-flash accepts image parts in chat completions messages. Output is text.
What does the GLM 5.3 Flash API cost?
The rate on this page is read from the catalog when the page renders. Input and output tokens bill separately, and a cached prompt prefix bills at a lower input rate.
GLM 5.3 Flash or GLM 5.2?
GLM 5.2 for the 1M-token window or the Anthropic Messages API. GLM 5.3 Flash for image input. 5.3 Flash speaks chat completions only.
How long a prompt can I send?
Up to 262,144 tokens, prompt and completion combined.
First request in five lines.
Point the OpenAI SDK at inference.openrelay.inc/v1 and run GLM 5.3 Flash per token. No contract, no minimums.