Qwen3.5-27B vs Qwen3.6 35B A3B
Compare Qwen3.5-27B and Qwen3.6 35B A3B side-by-side. See how these vision models stack up in Image Captioning, Open Prompt, Classification, Object Detection, and OCR.
Compare Qwen3.5-27B vs Qwen3.6 35B A3B live
Run the same image across every model that supports a task and compare their outputs side-by-side.
Detect and compare bounding boxes across models on the same image.
Upload an image
Drag and drop an image here, or click to browse
Models in this comparison
Qwen3.5-27B vs Qwen3.6 35B A3B on Vision Evals
Qwen3.6 35B A3B scores higher on 4 of the six Vision Evals tasks.
The widest gap is Object Detection, where Qwen3.6 35B A3B leads 57.0% to 50.5%.
Overall, Qwen3.5-27B averages 70.8% (#25 of 59) against 71.9% (#23 of 59) for Qwen3.6 35B A3B.
Qwen3.6 35B A3B is both cheaper ($0.0012 vs $0.0043 per sample) and faster (27.1s vs 80.4s per sample).
Qwen3.5-27B vs Qwen3.6 35B A3B Comparison Table
Evals updated September 22, 2026Pricing updated September 25, 2026
| Property | Qwen3.5-27B | Qwen3.6 35B A3B |
|---|---|---|
| Organization | Qwen | Qwen |
| Category | open | open |
| Modality | multimodal | multimodal |
| Release Date | Feb 2026 | Apr 2026 |
| Context Window | 262K | 262K |
| Parameters | 27B | 35B total, 3B active |
| License | Apache 2.0 | Apache 2.0 |
| Pricing per 1M tokens | ||
| Input $/1M | $0.195 | $0.150 |
| Output $/1M | $1.56 | $1.00 |
| Vision Tasks | ||
| Captioning | Demo | Demo |
| Chart Question Answering | ||
| Classification | Demo | Demo |
| Document Question Answering | ||
| Image Tagging | ||
| Multi-Label Classification | ||
| Object Detection | Demo | Demo |
| OCR | Demo | Demo |
| Vision Language | ||
| Visual Question Answering | Demo | Demo |
| Phrase Grounding | ||
| Video Classification | ||
| Model Features | ||
| Foundation Vision | ||
| LLMs with Vision Capabilities | ||
| Multimodal Vision | ||
Vision Evalsground-truth scores across 6 vision tasks | ||
| Overall | 70.8% | 71.9% |
| Quantizationsself-hosted | ||
| Avg cost / sample | $0.0043 | $0.0012 |
| Avg speed / sample | 80.37s | 27.10s |
| By task | ||
| Object Detection | 50.5% ±3.5, Mean of 3 runs, range 46.1 to 53.0 | 57.0% ±1.3, Mean of 3 runs, range 56.1 to 58.7 |
| Counting | 67.6% ±1.4, Mean of 3 runs, range 66.2 to 68.9 | 65.3% ±2.7, Mean of 3 runs, range 62.2 to 67.6 |
| Identification | 80.2% ±4.7, Mean of 3 runs, range 75.0 to 84.4 | 82.3% ±6.3, Mean of 3 runs, range 75.0 to 87.5 |
| OCR | 84.7% ±3.3, Mean of 3 runs, range 80.8 to 87.3 | 87.7% ±0.0, Mean of 3 runs, range 87.6 to 87.7 |
| Data Extraction | 83.8% ±1.5, Mean of 3 runs, range 82.5 to 85.6 | 84.5% ±1.0, Mean of 3 runs, range 83.5 to 85.6 |
| Reasoning | 58.1% ±2.3, Mean of 3 runs, range 55.6 to 60.3 | 54.8% ±1.3, Mean of 3 runs, range 53.0 to 55.6 |
Qwen3.5-27B vs Qwen3.6 35B A3B: Overview
Qwen3.5-27B is a multimodal dense hybrid model developed by Alibaba Cloud’s Qwen team and released in February 2026 as a high-precision entry in the Qwen3.5 "Medium" series. Unlike its Mixture-of-Experts (MoE) siblings, the 27B model utilizes a dense architecture combining Gated Delta Networks with a feed-forward structure, activating its full parameter suite for every inference to maximize reliability. This design provides the highest instruction-following and coding accuracy in its class, with a notable IFEval score of 95.0. The model features a native 262K-token context window, extensible to 1M tokens via YaRN (RoPE scaling), and is released under the Apache-2.0 license.
Optimized for agentic workflows, Qwen3.5-27B employs an early-fusion architecture that treats visual and textual data as a unified stream for deep cross-modal reasoning. This unified approach allows the model to excel in technical analysis and software engineering, matching GPT-5-mini with a 72.4% score on SWE-bench Verified. While the larger MoE variants in the family lead in raw knowledge benchmarks, the 27B model offers a stable and high-density alternative for structured data extraction and spatial perception, contributing to the Qwen3.5 family’s generational leap in OCR accuracy over the previous Qwen3-VL series.
Qwen3.6-35B-A3B is a sparse Mixture-of-Experts (MoE) multimodal language model developed by the Qwen team at Alibaba Group. It carries 35 billion total parameters but activates only approximately 3 billion per forward pass via a learned routing mechanism, giving it the representational capacity of a large dense model at a fraction of the inference compute. The model is natively multimodal, processing images, documents, and video alongside text as a core architectural capability rather than an add-on. It supports a native context window of 262,144 tokens, extensible up to 1,010,000 tokens via YaRN. A key design feature is the unified thinking/non-thinking mode framework: users can switch between deliberate chain-of-thought reasoning and fast direct responses within a single model, and a "thinking preservation" option retains reasoning context across multi-turn agentic workflows to reduce redundant computation.
The model is specifically optimized for agentic coding tasks, including repository-level reasoning, frontend workflow generation, multi-step tool use, and MCP (Model Context Protocol) integration. On SWE-bench Verified it scores 73.4%, on Terminal-Bench 2.0 it scores 51.5%, and on MCPMark it scores 37.0%. For vision-language tasks it achieves 92.0 on RefCOCO, 89.9 on OmniDocBench 1.5, and 83.7 on VideoMMMU. The model also supports Multi-Token Prediction (MTP) for speculative decoding. All Qwen3.6 open-weight models are released under the Apache 2.0 license.