Qwen3 VL 235B A22B Instruct is a flagship multimodal vision-language model developed by Qwen (Alibaba Cloud), designed for instruction-following tasks that combine advanced text generation with visual understanding. It serves as a high-end open-weight model for developers and researchers building multimodal AI systems that require strong reasoning, perception, and long-context capabilities.
The model supports interleaved text and image inputs, very long context windows (up to roughly 256K tokens), and efficient inference through a mixture-of-experts architecture with about 22B active parameters out of 235B total. In today’s landscape, it competes with top-tier proprietary vision-language models while offering the advantages of open weights and flexible deployment. Typical applications include multimodal assistants, document and image analysis, visual reasoning, and large-context instruction-based workflows.
Drag and drop an image here, or click to browse
Captioning will run automatically
—
Usage
Past 30 DaysVision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.
Evals updated July 10, 2026Pricing updated July 21, 2026
Qwen3 VL 235B A22B Instruct averages 66.4% across the six Vision Evals tasks, ranking #10 of 16 models overall.
Its weakest relative showing is Counting, ranking #16 of 16 at 47.3%.
At $0.0007 per sample it is the 2nd cheapest of the 16 benchmarked models, and its average inference time of 8.2s per sample makes it the 12th fastest.
Field medians: Object Detection 41.5%, Counting 62.2%, Identification 84.4%, OCR 89.1%, Data Extraction 85.6%, Reasoning 76.1%.
| Task | Score | Field (0 to 100) | Rank | Cost / sample | Speed |
|---|---|---|---|---|---|
| Object Detection | 42.3% | #8 of 16 | $0.0011 | 12.3s | |
| Counting | 47.3% | #16 of 16 | $0.0002 | 3.9s | |
| Identification | 90.6% | #6 of 16 | $0.0002 | 3.2s | |
| OCR | 88.1% | #13 of 16 | $0.0010 | 10.5s | |
| Data Extraction | 86.6% | #8 of 16 | $0.0002 | 2.8s | |
| Reasoning | 43.5% | #13 of 16 | $0.0002 | 2.9s |
Overall benchmark score against estimated cost per sample. Upper-left is the sweet spot: high quality at low cost.
16 models on the current benchmark · scores and efficiency pooled across all six tasks · Qwen3-VL 235B highlighted
Qwen3 VL 235B A22B Instruct scores from a single evaluation run · Methodology
View all Vision Evals →Qwen3 VL 235B A22B Instruct costs $0.210 per 1M input tokens and $1.90 per 1M output tokens.
Pricing updated Jul 21, 2026
Other models worth comparing for similar use cases.
License terms and commercial-use guidance for Qwen3 VL 235B A22B Instruct.
This model is released under the Apache License 2.0, a permissive open-source license that allows commercial use, modification, distribution, and patent use.
Yes. Under the terms of the Apache 2.0 license, you can freely use this model for commercial purposes, including in proprietary products. You must retain the copyright notice and disclaimers when redistributing.
License information is provided as a guide and is not legal advice.
Yes. Qwen3 VL 235B A22B Instruct accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is Identification at 90.6% (#6 of 16). You can test it on your own image in the demo above.
Yes. its transcriptions match the ground truth 88.1% on average (#13 of 16) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 86.6%.
Not its strength. On Vision Evals, Qwen3 VL 235B A22B Instruct scores 42.3% mAP@50 on object detection (#8 of 16) and 47.3% exact-match accuracy on object counting. For production counting or precise localization, pairing it with a specialized detector like RF-DETR or your own trained model in a Roboflow Workflow is usually more reliable: detect the objects, then count the detections.
On our benchmark's task mix, Qwen3 VL 235B A22B Instruct averages $0.0007 per sample at $0.21 per 1M input and $1.90 per 1M output tokens (#2 of 16 on cost), with an average speed of 8.2s per sample across the benchmark. Actual cost depends on your images and prompts.
On the overall Vision Evals ranking, Qwen3 VL 235B A22B Instruct sits #10 of 16 at 66.4%, just behind Gemini 2.5 Pro (67.9%) and just ahead of Qwen 3.7 Plus (66.2%). See the full side-by-side: Qwen3 VL 235B A22B Instruct vs Gemini 2.5 Pro.