GLM-5.3-Flash is the first natively multimodal model in Z.ai's GLM-5 series, a mixture-of-experts transformer with roughly 320 billion total parameters and 18 billion activated per token. It routes each token through 8 of 288 experts across 45 language layers that interleave KDA linear attention with sparse multi-head latent attention, and pairs them with a 24-layer vision encoder that handles image and video input. The checkpoint declares a maximum context length of 1,048,576 tokens, ships in native FP8, and includes a multi-token prediction draft layer for speculative decoding. Z.ai reports that the hybrid attention design reduces attention computation by 3.01x and KV cache size by 4.44x relative to GLM-5.3.
The model starts from a newly trained base built on a 30 trillion token multimodal pre-training corpus and adopts Manifold-Constrained Hyper-Connections to improve scaling efficiency. Vision is integrated into the coding and agent loop, so the model can inspect interfaces, rendered output, and images while operating across code, browsers, and graphical user interfaces. Z.ai reports scores of 84.3 on Terminal-Bench 2.1, 63.4 on DeepSWE 1.1, 55.3 on Humanity's Last Exam with tools, and 48.8 on AutomationBench, and the model exposes low, high, and max thinking modes.
Drag and drop an image here, or click to browse
Model settings
Thinking level
Max output tokens
Default 65,536 · max 65,536
Sign in to adjust thinking and output length per run.
Results appear here. Add an image or pick an example to run GLM 5.3 Flash.
Usage
Past 30 DaysVision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.
Evals updated September 5, 2026Pricing updated September 13, 2026
GLM 5.3 Flash averages 66.3% across the six Vision Evals tasks, ranking #33 of 53 models overall.
Its weakest relative showing is Object Detection, ranking #46 of 53 at 33.1%.
At $0.0005 per sample it is the 7th cheapest of the 53 benchmarked models, and its average inference time of 6.8s per sample makes it the 14th fastest.
Field medians: Object Detection 53.9%, Counting 56.8%, Identification 84.4%, OCR 88.7%, Data Extraction 84.5%, Reasoning 54.1%.
| Task | Score | Field (0 to 100) | Rank | Cost / sample | Speed |
|---|---|---|---|---|---|
| Object Detection | 33.1% | #46 of 53 | $0.0008 | 9.10s | |
| Counting | 55.4% | #30 of 53 | $0.0002 | 6.20s | |
| Identification | 84.4% | #23 of 53 | $0.0002 | 6.45s | |
| OCR | 90.6% | #21 of 53 | $0.0004 | 7.41s | |
| Data Extraction | 83.5% | #32 of 53 | $0.0002 | 4.78s | |
| Reasoning (low) | 51.0% | #29 of 53 | $0.0002 | 4.44s | |
| Reasoning (high) | 59.6% | #31 of 39 | $0.0003 | 5.38s |
Overall benchmark score against estimated cost per sample, on a log scale. Upper-left is the sweet spot: high quality at low cost.
52 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort · GLM 5.3 Flash highlighted
GLM 5.3 Flash scores are from a single run per task; a three-run re-run under the current protocol is pending · Methodology
View all Vision Evals →GLM 5.3 Flash costs $0.150 per 1M input tokens and $0.500 per 1M output tokens.
Pricing updated Sep 13, 2026
Other models worth comparing for similar use cases.
GLM 5.3 Flash is released under MIT, a permissive license. The GLM 5.3 Flash license lets you use, modify, and sell work built on the model, with the copyright notice as the only real obligation and no requirement to open-source related code changes.
MIT grants no explicit patent license and disclaims all warranties. If patent exposure is a concern for your deployment, review it with counsel before launch.
Read the full MIT license ↗No commercial license is needed for GLM 5.3 Flash: permissive terms let you keep related code private while deploying commercially.
Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.
Talk to salesThis model is released under the MIT License, a short and permissive open-source license that allows commercial use, modification, and redistribution.
Yes. Under the terms of the MIT license, you can freely use this model for commercial purposes. You must retain the copyright notice and license text when redistributing.
License information is provided as a guide and is not legal advice.
Yes. GLM 5.3 Flash accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is OCR at 90.6% (#21 of 53 at low effort). You can test it on your own image in the demo above.
Yes. its transcriptions match the ground truth 90.6% on average (#21 of 53 at low effort) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 83.5%.
Not its strength. On Vision Evals, GLM 5.3 Flash scores 33.1% mAP@50 on object detection (#46 of 53 at low effort) and 55.4% judge-graded accuracy on object counting. For production counting or precise localization, pairing it with a specialized detector like RF-DETR or your own trained model in a Roboflow Workflow is usually more reliable: detect the objects, then count the detections.
On our benchmark's task mix, GLM 5.3 Flash averages $0.0005 per sample at $0.15 per 1M input and $0.50 per 1M output tokens (#7 of 53 on cost), with an average speed of 6.8s per sample across the benchmark. Actual cost depends on your images and prompts.
On the overall Vision Evals ranking, GLM 5.3 Flash sits #33 of 53 at 66.3%, just behind Claude Sonnet 5 (66.4%) and just ahead of Gemini 2.5 Pro (66%). See the full side-by-side: GLM 5.3 Flash vs Claude Sonnet 5.