Roboflow
Z.ai

Z.ai: GLM 5.3 Flash

GLM 5.3 Flash Overview

GLM-5.3-Flash is the first natively multimodal model in Z.ai's GLM-5 series, a mixture-of-experts transformer with roughly 320 billion total parameters and 18 billion activated per token. It routes each token through 8 of 288 experts across 45 language layers that interleave KDA linear attention with sparse multi-head latent attention, and pairs them with a 24-layer vision encoder that handles image and video input. The checkpoint declares a maximum context length of 1,048,576 tokens, ships in native FP8, and includes a multi-token prediction draft layer for speculative decoding. Z.ai reports that the hybrid attention design reduces attention computation by 3.01x and KV cache size by 4.44x relative to GLM-5.3.

The model starts from a newly trained base built on a 30 trillion token multimodal pre-training corpus and adopts Manifold-Constrained Hyper-Connections to improve scaling efficiency. Vision is integrated into the coding and agent loop, so the model can inspect interfaces, rendered output, and images while operating across code, browsers, and graphical user interfaces. Z.ai reports scores of 84.3 on Terminal-Bench 2.1, 63.4 on DeepSWE 1.1, 55.3 on Humanity's Last Exam with tools, and 48.8 on AutomationBench, and the model exposes low, high, and max thinking modes.

GLM 5.3 Flash Interactive Demo

Model settings

Max output tokens

Default 65,536 · max 65,536

Sign in to adjust thinking and output length per run.

Results appear here. Add an image or pick an example to run GLM 5.3 Flash.

GLM 5.3 Flash Details & Performance

Details

Resources

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRObject DetectionVision LanguageVisual Question Answering

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Performance

Avg. Latency

Arena Rankings

GLM 5.3 Flash Vision Evals

Vision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.

Evals updated August 26, 2026Pricing updated August 26, 2026

Overall score#22 of 33
66.3%
Avg cost / sample#2 of 33
$0.0002
Avg speed / sample#13 of 33
6.78s
Avg tokens / sample
2.4K

Strengths and weaknesses

GLM 5.3 Flash averages 66.3% across the six Vision Evals tasks, ranking #22 of 33 models overall.

Its weakest relative showing is Object Detection, ranking #30 of 33 at 33.1%.

At $0.0002 per sample it is the 2nd cheapest of the 33 benchmarked models, and its average inference time of 6.8s per sample makes it the 13th fastest.

Performance profile

Field medianGLM 5.3 Flash

Field medians: Object Detection 54.4%, Counting 60.8%, Identification 84.4%, OCR 89.3%, Data Extraction 85.6%, Reasoning 55.6%.

Results by task

TaskScoreField (0 to 100)RankCost / sampleSpeed
Object Detection
33.1%
#30 of 33$0.00049.10s
Counting
55.4%
#19 of 33$0.00016.20s
Identification
84.4%
#13 of 33$0.00016.45s
OCR
90.6%
#16 of 33$0.00027.41s
Data Extraction
83.5%
#21 of 33$0.00014.78s
Reasoning (low)
51.0%
#20 of 33$0.00014.44s
Reasoning (high)
59.6%
#25 of 33$0.00015.38s
  • Thinking longer helps: 8.6 points higher on reasoning at high effort for 1.2x the cost and 1.2x the latency.

Price vs. performance

Score vs. cost

Overall benchmark score against estimated cost per sample. Upper-left is the sweet spot: high quality at low cost.

33 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort · GLM 5.3 Flash highlighted

GLM 5.3 Flash scores from a single evaluation run · Methodology

View all Vision Evals →

GLM 5.3 Flash License

MIT · Permissive license

GLM 5.3 Flash is released under MIT, a permissive license. The GLM 5.3 Flash license lets you use, modify, and sell work built on the model, with the copyright notice as the only real obligation and no requirement to open-source related code changes.

Commercial use
Permitted with no separate commercial license. No usage caps, revenue thresholds, or field-of-use limits apply to GLM 5.3 Flash.
Modification
Permitted. You can fine-tune or rewrite GLM 5.3 Flash and keep the result closed-source.
Redistribution
Permitted. Include the original copyright and permission notice in copies or substantial portions of the work.

MIT grants no explicit patent license and disclaims all warranties. If patent exposure is a concern for your deployment, review it with counsel before launch.

Read the full MIT license ↗

Do I need a commercial license for GLM 5.3 Flash?

No commercial license is needed for GLM 5.3 Flash: permissive terms let you keep related code private while deploying commercially.

Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.

Talk to sales

This model is released under the MIT License, a short and permissive open-source license that allows commercial use, modification, and redistribution.

Yes. Under the terms of the MIT license, you can freely use this model for commercial purposes. You must retain the copyright notice and license text when redistributing.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About GLM 5.3 Flash Vision

Yes. GLM 5.3 Flash accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is Identification at 84.4% (#13 of 33). You can test it on your own image in the demo above.

Yes. its transcriptions match the ground truth 90.6% on average (#16 of 33) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 83.5%.

Not its strength. On Vision Evals, GLM 5.3 Flash scores 33.1% mAP@50 on object detection (#30 of 33) and 55.4% exact-match accuracy on object counting. For production counting or precise localization, pairing it with a specialized detector like RF-DETR or your own trained model in a Roboflow Workflow is usually more reliable: detect the objects, then count the detections.

On our benchmark's task mix, GLM 5.3 Flash averages $0.0002 per sample (#2 of 33 on cost), with an average speed of 6.8s per sample across the benchmark. Actual cost depends on your images and prompts.

On the overall Vision Evals ranking, GLM 5.3 Flash sits #22 of 33 at 66.3%, just behind Claude Sonnet 5 (66.4%) and just ahead of Gemini 2.5 Pro (66%). See the full side-by-side: GLM 5.3 Flash vs Claude Sonnet 5.