Roboflow
Qwen

Qwen: Qwen3 VL 235B A22B Instruct

Qwen3 VL 235B A22B Instruct Overview

Qwen3 VL 235B A22B Instruct is a flagship multimodal vision-language model developed by Qwen (Alibaba Cloud), designed for instruction-following tasks that combine advanced text generation with visual understanding. It serves as a high-end open-weight model for developers and researchers building multimodal AI systems that require strong reasoning, perception, and long-context capabilities.

The model supports interleaved text and image inputs, very long context windows (up to roughly 256K tokens), and efficient inference through a mixture-of-experts architecture with about 22B active parameters out of 235B total. In today’s landscape, it competes with top-tier proprietary vision-language models while offering the advantages of open weights and flexible deployment. Typical applications include multimodal assistants, document and image analysis, visual reasoning, and large-context instruction-based workflows.

Qwen3 VL 235B A22B Instruct Interactive Demo

Qwen3 VL 235B A22B Instruct Details & Performance

Details

Resources

Vision Tasks

Vision LanguageObject DetectionOCRVisual Question AnsweringCaptioning

Features

LLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Performance

Avg. Latency

Arena Rankings

Qwen3 VL 235B A22B Instruct Vision Evals

Vision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.

Evals updated July 10, 2026Pricing updated July 21, 2026

Overall score#10 of 16
66.4%
Avg cost / sample#2 of 16
$0.0007
Avg speed / sample#12 of 16
8.2s
Avg tokens / sample
1.6K

Strengths and weaknesses

Qwen3 VL 235B A22B Instruct averages 66.4% across the six Vision Evals tasks, ranking #10 of 16 models overall.

Its weakest relative showing is Counting, ranking #16 of 16 at 47.3%.

At $0.0007 per sample it is the 2nd cheapest of the 16 benchmarked models, and its average inference time of 8.2s per sample makes it the 12th fastest.

Performance profile

Field medianQwen3 VL 235B A22B Instruct

Field medians: Object Detection 41.5%, Counting 62.2%, Identification 84.4%, OCR 89.1%, Data Extraction 85.6%, Reasoning 76.1%.

Results by task

TaskScoreField (0 to 100)RankCost / sampleSpeed
Object Detection
42.3%
#8 of 16$0.001112.3s
Counting
47.3%
#16 of 16$0.00023.9s
Identification
90.6%
#6 of 16$0.00023.2s
OCR
88.1%
#13 of 16$0.001010.5s
Data Extraction
86.6%
#8 of 16$0.00022.8s
Reasoning
43.5%
#13 of 16$0.00022.9s

Price vs. performance

Score vs. cost

Overall benchmark score against estimated cost per sample. Upper-left is the sweet spot: high quality at low cost.

16 models on the current benchmark · scores and efficiency pooled across all six tasks · Qwen3-VL 235B highlighted

Qwen3 VL 235B A22B Instruct scores from a single evaluation run · Methodology

View all Vision Evals →

Qwen3 VL 235B A22B Instruct Pricing

Qwen3 VL 235B A22B Instruct costs $0.210 per 1M input tokens and $1.90 per 1M output tokens.

Input$0.210 / 1M tokens
Output$1.90 / 1M tokens
Cached input$0.100 / 1M tokens

Pricing updated Jul 21, 2026

Alternatives to Qwen3 VL 235B A22B Instruct

Other models worth comparing for similar use cases.

Qwen
Qwen3.5 397B A17B
Qwen3.5-397B-A17B is a 397B-parameter (17B active) open-weight multimodal model developed by Alibaba’s Qwen team, released on 2026-02-16 under Apache-2.0. It supports text and image inputs with text outputs, combining a sparse Mixture-of-Experts architecture with Gated Delta Networks for efficient scaling. The model provides native vision-language reasoning and a large ~262K token context window, extendable to ~1M tokens.As the first open-weight release in the Qwen3.5 family, it positions itself as a high-capacity, long-context alternative in the large vision-language space, balancing scale and efficiency via sparse activation. It is designed for advanced reasoning, coding, agent workflows, and multimodal understanding tasks.
Qwen
Qwen3.5 122B A10B
Qwen3.5-122B-A10B is a high-capacity multimodal Mixture-of-Experts (MoE) model developed by Alibaba’s Qwen team as part of the Qwen3.5 model family. The architecture contains 122 billion total parameters while activating roughly 10 billion per token through sparse expert routing, allowing the model to balance large-scale reasoning ability with relatively efficient inference compared to dense models of similar size.The model is designed to process both text and visual inputs within a unified multimodal framework, enabling tasks that require reasoning across images, documents, charts, and natural language. This makes it suitable for applications such as document understanding, diagram interpretation, and complex visual question answering.Qwen3.5-122B-A10B supports a native context window of approximately 256,000 tokens, which can be extended further through techniques such as YaRN scaling to support very long-context workloads. Released under the Apache 2.0 license, it builds on earlier Qwen multimodal systems and provides developers with an open-weight model capable of handling demanding multimodal reasoning and analysis tasks.
Meta
Llama 4 Maverick
Llama 4 Maverick, introduced on April 5, 2025, is one of the first models in Meta’s Llama 4 family, designed as a natively multimodal model supporting text + image inputs with text outputs. It employs a Mixture-of-Experts (MoE) architecture with 128 experts, activating ~17B parameters per token out of a pool of ~400B total parameters. This design improves scalability, efficiency, and reasoning capacity. Maverick has a 1M-token context window, enabling it to handle large documents, extended conversations, and multimodal reasoning. Its knowledge cutoff is August 2024.The model is released under the Llama 4 Community License and comes in both base and instruction-tuned (“Instruct”) versions. Maverick is widely deployed via Hugging Face, Google Vertex AI, Amazon Bedrock, and Oracle Cloud, making it one of the most accessible large open-weight models. However, it outputs text only (no image/audio generation) and, while input capacity is huge, output limits are typically much smaller. The MoE design also raises hardware demands, as maintaining 128 experts requires significant compute resources, and Meta’s license introduces restrictions around commercial-scale use.
Google
Gemini 2.5 Pro
Gemini 2.5 Pro, released on June 17, 2025, is Google DeepMind’s most capable model in the Gemini 2.5 family, optimized for deep reasoning, coding, and complex multimodal tasks. It accepts text, images, audio, video, and PDFs as input and outputs text. The model supports 1 million input tokens with an output capacity of up to 65K tokens, enabling large-scale comprehension of datasets, codebases, and technical documents. Its training knowledge extends to January 2025.Pro outperforms earlier Gemini 2.0 models across benchmarks, including agentic coding tasks where it achieved ~63.8% on SWE-Bench Verified. It supports structured outputs, function calling, code execution, search grounding, and URL context, making it well-suited for enterprise, STEM, and developer workflows. However, it does not currently support image or audio generation in its stable release, and its higher computational cost and latency make it less efficient than Flash or Flash-Lite. It is available via the Gemini API, Google AI Studio, and Vertex AI.
Anthropic
Claude Opus 4.1
Claude 4.1 Opus, released by Anthropic in August 2025, is the upgraded flagship of the Claude 4 family, building on Opus 4 with stronger reasoning and agentic capabilities. Like its predecessor, it is multimodal and optimized for text, code, and tool use, with support for large context windows suited to multi-file codebases, technical workflows, and long-horizon problem solving. On benchmarks, Opus 4.1 improves coding performance, reaching ~74.5% on SWE-Bench Verified compared to Opus 4’s ~72.5%. It demonstrates more precise debugging, refactoring, and orchestration of agentic tasks while maintaining similar safety and alignment safeguards. It is best suited for enterprise-scale software development, research automation, and advanced reasoning workflows where reliability and depth of analysis are critical.
MoonshotAI
Kimi K2.5
Kimi K2.5 is a frontier-scale multimodal AI model developed by Moonshot AI and released on January 27, 2026. As a significant advancement within the Kimi K2 family, it utilizes a sparse Mixture-of-Experts (MoE) architecture with 1 trillion total parameters (32 billion active per inference) and a massive 256K-token context window. The model features native multimodal integration via a 400M-parameter MoonViT encoder, allowing it to process text, images, and video frames simultaneously. Built for both speed and depth, it offers "Instant" and "Thinking" modes, the latter of which excels at expert-level reasoning, scoring 50.2% on the Humanity’s Last Exam (HLE) benchmark when equipped with tools.The model is released under a Modified MIT License, which remains open-weight but requires attribution for high-revenue commercial entities. It introduces an "Agent Swarm" paradigm capable of coordinating up to 100 specialized sub-agents for parallel workflows, significantly reducing latency in complex research tasks. For vision tasks, Kimi K2.5 demonstrates strong autonomous visual debugging capabilities, where it can inspect its own generated UI outputs against visual specifications to iteratively refine frontend code. This makes it a powerful choice for developers testing automated UI reconstruction, high-fidelity OCR document processing, and multi-step agentic research grounded in complex visual data.

Qwen3 VL 235B A22B Instruct License

Apache 2.0

License terms and commercial-use guidance for Qwen3 VL 235B A22B Instruct.

This model is released under the Apache License 2.0, a permissive open-source license that allows commercial use, modification, distribution, and patent use.

Read the full Apache 2.0 license ↗

Yes. Under the terms of the Apache 2.0 license, you can freely use this model for commercial purposes, including in proprietary products. You must retain the copyright notice and disclaimers when redistributing.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About Qwen3 VL 235B A22B Instruct Vision

Yes. Qwen3 VL 235B A22B Instruct accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is Identification at 90.6% (#6 of 16). You can test it on your own image in the demo above.

Yes. its transcriptions match the ground truth 88.1% on average (#13 of 16) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 86.6%.

Not its strength. On Vision Evals, Qwen3 VL 235B A22B Instruct scores 42.3% mAP@50 on object detection (#8 of 16) and 47.3% exact-match accuracy on object counting. For production counting or precise localization, pairing it with a specialized detector like RF-DETR or your own trained model in a Roboflow Workflow is usually more reliable: detect the objects, then count the detections.

On our benchmark's task mix, Qwen3 VL 235B A22B Instruct averages $0.0007 per sample at $0.21 per 1M input and $1.90 per 1M output tokens (#2 of 16 on cost), with an average speed of 8.2s per sample across the benchmark. Actual cost depends on your images and prompts.

On the overall Vision Evals ranking, Qwen3 VL 235B A22B Instruct sits #10 of 16 at 66.4%, just behind Gemini 2.5 Pro (67.9%) and just ahead of Qwen 3.7 Plus (66.2%). See the full side-by-side: Qwen3 VL 235B A22B Instruct vs Gemini 2.5 Pro.