Roboflow
Qwen

Qwen: Qwen3.5 122B A10B

Qwen3.5 122B A10B Overview

Qwen3.5-122B-A10B is a high-capacity multimodal Mixture-of-Experts (MoE) model developed by Alibaba’s Qwen team as part of the Qwen3.5 model family. The architecture contains 122 billion total parameters while activating roughly 10 billion per token through sparse expert routing, allowing the model to balance large-scale reasoning ability with relatively efficient inference compared to dense models of similar size.

The model is designed to process both text and visual inputs within a unified multimodal framework, enabling tasks that require reasoning across images, documents, charts, and natural language. This makes it suitable for applications such as document understanding, diagram interpretation, and complex visual question answering.

Qwen3.5-122B-A10B supports a native context window of approximately 256,000 tokens, which can be extended further through techniques such as YaRN scaling to support very long-context workloads. Released under the Apache 2.0 license, it builds on earlier Qwen multimodal systems and provides developers with an open-weight model capable of handling demanding multimodal reasoning and analysis tasks.

Qwen3.5 122B A10B Interactive Demo

Results appear here. Add an image or pick an example to run Qwen3.5 122B A10B.

Qwen3.5 122B A10B Details & Performance

Details

Resources

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRObject DetectionVision LanguageVisual Question Answering

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Performance

Avg. Latency

Qwen3.5 122B A10B Vision Evals

Qwen3.5 122B A10B has not yet been evaluated on the current benchmark. The results below are from the legacy version of Vision Evals, our previous benchmark. See the current Vision Evals

Visual Understanding

77 models · 67 tasks
HighestLowest
This model#9 of 7776.12% pass rate · better than 86%
Score76.12%pass rate across 67 tasks
Speed1.77savg response per task
Cost$0.0003 / task$0.260 in · $2.08 out / 1M
Tokens1.2K / task1.2K in · 7 out
Score key:≥75%40–74%<40%
CategoryPassedScore
Object Understanding13 / 14
92.9%
Defect Detection13 / 15
86.7%
Document Understanding7 / 9
77.8%
Spatial Understanding14 / 19
73.7%
Object Counting4 / 10
40%

Scores based on a single evaluation run · Methodology

View all legacy Vision Evals results →

Qwen3.5 122B A10B Pricing

Qwen3.5 122B A10B costs $0.260 per 1M input tokens and $2.08 per 1M output tokens.

Input$0.260 / 1M tokens
Output$2.08 / 1M tokens

Pricing updated Sep 22, 2026

Price vs. performance

Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.

11 of 11 models plotted

ModelScoreMedian tokensEst. cost / taskCompare
GoogleGemini 3.5 Flash79.1%1.4K$0.0043Compare
AnthropicClaude Fable 579.1%2.9K$0.041Compare
OpenAIGPT-5.4 Mini77.6%1.9K$0.0015Compare
OpenAIGPT-5.477.6%1.7K$0.0052Compare
OpenAIGPT-5.577.6%1.7K$0.011Compare
QwenQwen3.5 122B A10B(this model)76.1%1.2K$0.0003
OpenAIGPT-5.6 Terra76.1%1.5K$0.0033Compare
OpenAIGPT-5.6 Sol76.1%1.5K$0.0029Compare
GoogleGemini 3.1 Pro75.8%1.1K$0.0024Compare
GoogleGemini 3 Flash74.6%1.4K$0.0014Compare
OpenAIGPT-5 Mini73.1%1.8K$0.0006Compare

Alternatives to Qwen3.5 122B A10B

Other models worth comparing for similar use cases.

Qwen
Qwen3 VL 235B A22B Instruct
Qwen3 VL 235B A22B Instruct is a flagship multimodal vision-language model developed by Qwen (Alibaba Cloud), designed for instruction-following tasks that combine advanced text generation with visual understanding. It serves as a high-end open-weight model for developers and researchers building multimodal AI systems that require strong reasoning, perception, and long-context capabilities.The model supports interleaved text and image inputs, very long context windows (up to roughly 256K tokens), and efficient inference through a mixture-of-experts architecture with about 22B active parameters out of 235B total. In today’s landscape, it competes with top-tier proprietary vision-language models while offering the advantages of open weights and flexible deployment. Typical applications include multimodal assistants, document and image analysis, visual reasoning, and large-context instruction-based workflows.
Google
Gemini 3.1 Pro
Gemini 3.1 Pro is a proprietary multimodal model from Google’s Gemini 3 series, released in early 2026 and designed for advanced reasoning across large multimodal datasets. It accepts text, images, audio, video, and documents, supporting up to a 1-million-token input context with up to 64k output tokens. Compared with Gemini 3 Pro, it improves long-context synthesis and multi-step reasoning, enabling more reliable analysis of large documents, datasets, and software codebases.The model also advances visual understanding and grounding, allowing it to interpret UI screenshots, diagrams, and real-world scenes while referencing specific regions within images or video. These capabilities make Gemini 3.1 Pro well suited for multimodal workflows involving document processing, interface analysis, robotics research, and complex visual reasoning.
Anthropic
Claude Opus 4.7
Claude Opus 4.7 is a proprietary multimodal language model developed by Anthropic, released on April 16, 2026. It is designed for agentic coding, long-horizon task execution, and enterprise knowledge work. The model supports text and vision inputs and operates with a context window of up to 1,000,000 tokens. It introduces adaptive thinking, which dynamically allocates reasoning based on task complexity, along with configurable effort controls including a new xhigh setting that sits between the existing high and max levels. It achieves 87.6% on SWE-bench Verified and 78.0% on OSWorld-Verified, reflecting strong performance on autonomous software engineering and computer use tasks respectively.Compared to Claude Opus 4.6, version 4.7 shows improved instruction following and higher reliability in extended agentic tasks. Vision capabilities now support high-resolution inputs up to 2,576px on the long edge (~3.75 megapixels), more than three times the resolution of prior Claude models, enabling finer interpretation of dense diagrams, UI screenshots, and document layouts. These improvements, combined with self-verification on long-running tasks and a new task budget system for controlling agentic loops, make it well-suited for complex software engineering, technical analysis, and multimodal vision workflows.
OpenAI
GPT-5.6 Sol
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 family, which also includes Terra (a balanced everyday-work tier) and Luna (a fast, cost-efficient tier). Sol is designed for demanding reasoning, long-horizon agentic workflows, software engineering, computer use, scientific research, and cybersecurity tasks. It introduces two new capability modes: a "max" reasoning effort setting that allocates additional compute time for difficult problems, and an "ultra" mode that coordinates multiple subagents in parallel to accelerate complex, multi-step work. The model supports native multimodal input, allowing it to process screenshots, diagrams, charts, documents, and photographs alongside text. A reported context window of approximately 1.5 million tokens enables processing of large codebases, lengthy research documents, and extended agentic sessions.GPT-5.6 Sol was announced on June 26, 2026, initially in a limited preview for trusted partners, and reached general availability on July 9, 2026. On the Agents' Last Exam benchmark, which evaluates long-running professional workflows across 55 fields, Sol scores 53.6. On Terminal-Bench 2.1, which tests command-line agentic coding workflows, Sol Ultra achieves 91.9%. The model also demonstrates gains in life sciences evaluations, including long-horizon genomics and quantitative biology analyses. OpenAI paired the release with its most extensive safety evaluation to date, combining human red teaming with large-scale automated testing, and classified Sol as High capability in both cybersecurity and biological risk under its Preparedness Framework, though it does not cross the Critical threshold in either category.
Qwen
Qwen3.6 Plus
Qwen3.6 Plus is a flagship model in Alibaba’s Qwen Plus series, designed for agentic workflows, coding, and multi-step reasoning. It supports a 1 million token context window and up to 65,536 output tokens, with built-in reasoning capabilities. The model is available as a hosted, proprietary API through Alibaba Cloud. Compared to Qwen3.5, it improves reliability in multi-step execution and frontend code generation, with stronger performance on agentic coding tasks. It also supports document and image understanding, though its vision capabilities are more limited than dedicated Qwen-VL models. Qwen3.6 Plus is part of a broader Qwen ecosystem that includes both closed-source APIs and open-weight models.
Grok
Grok 4.5
Grok 4.5 is a proprietary reasoning model from SpaceXAI (xAI) that accepts interleaved text and image input and returns text, with a 500,000 token context window. xAI positions it as a model for coding, agentic software work, and knowledge tasks, and states it was trained in the company's Memphis data centers on datasets spanning science, engineering, and mathematics. Its reinforcement learning stage covers hundreds of thousands of multi step software engineering tasks scored by automated checks and model based grading, and training is reported to have run on tens of thousands of NVIDIA GB300 GPUs using an asynchronous scheme in which multi hour agentic rollouts continue while learning proceeds in parallel, targeting long horizon autonomous operation rather than single turn inference.For vision, the model consumes JPEG and PNG images in any order relative to text prompts, covering visual question answering, description of chart and document imagery, and reading text rendered inside a scene. Reasoning effort is configurable, and the model supports function calling and structured outputs, so image inputs can be interleaved with tool calls inside agent loops. xAI has not published a technical report, architecture details, or parameter count, and reported mixture of experts sizing figures come from secondary coverage rather than official documentation.

Qwen3.5 122B A10B License

Apache-2.0 · Permissive license

Qwen3.5 122B A10B is released under Apache-2.0, a permissive license. The Qwen3.5 122B A10B license lets you run, fine-tune, and redistribute the model in commercial products with no obligation to open-source related code changes, so no separate commercial license is required.

Commercial use
Permitted. Because Apache-2.0 is permissive, Qwen3.5 122B A10B can ship inside paid products and internal systems with no commercial license and no revenue threshold.
Modification
Permitted. Fine-tuning, quantizing, and distilling are all allowed, and your code changes can stay closed. Files you change must be marked as changed.
Redistribution
Permitted with attribution. Ship the Apache-2.0 license text and any NOTICE file alongside the weights or derived code.

Apache-2.0 grants an express patent license that terminates if you bring a patent claim over the work, and it disclaims warranties. Validate Qwen3.5 122B A10B on your own data before you depend on it in production.

Read the full Apache 2.0 license ↗

Do I need a commercial license for Qwen3.5 122B A10B?

This is the straightforward case: a permissive license is the best technical solution and you are free to deploy Qwen3.5 122B A10B commercially without open-sourcing your own code.

Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.

Talk to sales

This model is released under the Apache License 2.0, a permissive open-source license that allows commercial use, modification, distribution, and patent use.

Yes. Under the terms of the Apache 2.0 license, you can freely use this model for commercial purposes, including in proprietary products. You must retain the copyright notice and disclaimers when redistributing.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About Qwen3.5 122B A10B Vision

Yes. Qwen3.5 122B A10B accepts image input, and on Roboflow's previous vision benchmark it passed 76.1% of visual understanding tasks (#9 of 77). You can test it on your own image in the demo above.

Qwen3.5 122B A10B has not yet been evaluated on Roboflow's current Vision Evals. The results on this page are from the previous benchmark.

Yes. The demo on this page runs Qwen3.5 122B A10B in the free Roboflow Playground: upload an image and see results in seconds. A free account unlocks unlimited runs.