GLM-5.3-Flash
Z.ai's open-weights multimodal efficiency model. Cheap enough and capable enough to run a personal assistant on all month.
Quality
Modality
multimodal
Context
1M tokens
Access
open weights
Fabian's Take
"This is what I run my OpenClaw assistant on, and it's incredibly capable and fast for the price. It manages my email and calendar and does the occasional piece of research, and the whole thing costs me roughly CHF 5 a month through OpenRouter. It has less personality than some other models, but it's still pleasant to work with."
GLM-5.3-Flash is Z.ai’s efficiency tier, released on 26 August 2026 and the first natively multimodal model in the GLM-5 series. It takes text, images, video, and files in and returns text, with tool calling and structured output. The weights are on Hugging Face under an MIT license.
The case for it
It’s the model I run my OpenClaw assistant on. That assistant handles my email and my calendar and does the occasional piece of research, and it costs me about CHF 5 a month through OpenRouter. For a thing that works every day, that is a remarkable number, and it’s the whole argument for this model.
What makes it possible is the price: $0.15 per million input tokens and $0.50 per million output, roughly a tenth of what the GLM-5.3 flagship lists at. An assistant burns a lot of tokens on tool calls that no human ever reads, and at these rates you stop having to care.
One setting to change
Thinking is always on and the effort level defaults to maximum, which cannot be fully disabled. For hard problems that’s what you want. For “move this meeting” it isn’t, and you’ll pay for reasoning nobody needed. Turn reasoning_effort down for routine work and leave it high for the jobs that deserve it.
Where it sits
Z.ai’s own benchmark tables put Flash ahead of Claude Opus 4.8 on several agentic, automation, and knowledge-work lines, and behind it on some coding and hard-reasoning ones. Those are vendor-reported numbers, so treat the direction as informative and the decimal places as marketing. The honest summary is that it’s close enough to the expensive models on the kind of tool-driven work an assistant does, and far enough below them that you wouldn’t hand it your hardest problem.
It also has noticeably less personality than the frontier chat models. That matters if you’re talking to it all day and matters not at all if it’s quietly filing your email.
The Verdict
Best for: Running an assistant or agent that works all day on tool calls, where cost per task matters more than personality.
Pros
- Cheap enough to leave running: a personal assistant costs a few francs a month
- Takes text, images, video, and files in, which covers most of what an assistant is handed
- Open weights under MIT, so nothing stops you self-hosting it later
- Z.ai's own tables put it ahead of Claude Opus 4.8 on several agentic and automation benchmarks
Cons
- Thinking is always on and defaults to maximum effort, so routine jobs cost more than they need to until you turn it down
- Less personality than the frontier chat models
- Benchmark wins are vendor-reported, and it still trails on the harder coding and reasoning lines
- Running the full 320B-parameter model yourself needs serious hardware
Specs
- Pricing $0.15/M input, $0.50/M output · cached input $0.03/M
- Cost Tier budget
- ⚡Speed Tier fast
- License MIT