Roboflow

Qwen3.5-27B vs Qwen3.6 35B A3B

Compare Qwen3.5-27B and Qwen3.6 35B A3B side-by-side. See how these vision models stack up in Image Captioning, Open Prompt, Classification, Object Detection, and OCR.

Compare Qwen3.5-27B vs Qwen3.6 35B A3B live

Run the same image across every model that supports a task and compare their outputs side-by-side.

Detect and compare bounding boxes across models on the same image.

Open Object Detection in the full playground
QwenQwen3.5-27B
Run to compare this model.
QwenQwen3.6 35B A3B
Run to compare this model.

Models in this comparison

Qwen3.5-27B vs Qwen3.6 35B A3B on Vision Evals

Qwen3.6 35B A3B scores higher on 4 of the six Vision Evals tasks.

The widest gap is Object Detection, where Qwen3.6 35B A3B leads 57.0% to 50.5%.

Overall, Qwen3.5-27B averages 70.8% (#25 of 59) against 71.9% (#23 of 59) for Qwen3.6 35B A3B.

Qwen3.6 35B A3B is both cheaper ($0.0012 vs $0.0043 per sample) and faster (27.1s vs 80.4s per sample).

Qwen3.5-27BQwen3.6 35B A3B

Qwen3.5-27B vs Qwen3.6 35B A3B Comparison Table

Evals updated September 22, 2026Pricing updated September 25, 2026

PropertyQwen3.5-27BQwen3.6 35B A3B
OrganizationQwenQwen
Categoryopenopen
Modalitymultimodalmultimodal
Release DateFeb 2026Apr 2026
Context Window262K262K
Parameters27B35B total, 3B active
LicenseApache 2.0Apache 2.0
Pricing per 1M tokens
Input $/1M$0.195$0.150
Output $/1M$1.56$1.00
Vision Tasks
CaptioningDemoDemo
Chart Question Answering
ClassificationDemoDemo
Document Question Answering
Image Tagging
Multi-Label Classification
Object DetectionDemoDemo
OCRDemoDemo
Vision Language
Visual Question AnsweringDemoDemo
Phrase Grounding
Video Classification
Model Features
Foundation Vision
LLMs with Vision Capabilities
Multimodal Vision
Vision Evalsground-truth scores across 6 vision tasks
Overall
70.8%
71.9%
Quantizationsself-hosted
BF1670.8%FP868.2%AWQ-INT469.3%hardware →
FP871.9%AWQ-INT468.7%hardware →
Avg cost / sample$0.0043$0.0012
Avg speed / sample80.37s27.10s
By task
Object Detection
50.5%
±3.5, Mean of 3 runs, range 46.1 to 53.0
$0
57.0%
±1.3, Mean of 3 runs, range 56.1 to 58.7
$0
Counting
67.6%
±1.4, Mean of 3 runs, range 66.2 to 68.9
$0
65.3%
±2.7, Mean of 3 runs, range 62.2 to 67.6
$0
Identification
80.2%
±4.7, Mean of 3 runs, range 75.0 to 84.4
$0
82.3%
±6.3, Mean of 3 runs, range 75.0 to 87.5
$0
OCR
84.7%
±3.3, Mean of 3 runs, range 80.8 to 87.3
$0
87.7%
±0.0, Mean of 3 runs, range 87.6 to 87.7
$0
Data Extraction
83.8%
±1.5, Mean of 3 runs, range 82.5 to 85.6
$0
84.5%
±1.0, Mean of 3 runs, range 83.5 to 85.6
$0
Reasoning
58.1%
±2.3, Mean of 3 runs, range 55.6 to 60.3
$0
54.8%
±1.3, Mean of 3 runs, range 53.0 to 55.6
$0

Qwen3.5-27B vs Qwen3.6 35B A3B: Overview

Qwen3.5-27B

Qwen3.5-27B is a multimodal dense hybrid model developed by Alibaba Cloud’s Qwen team and released in February 2026 as a high-precision entry in the Qwen3.5 "Medium" series. Unlike its Mixture-of-Experts (MoE) siblings, the 27B model utilizes a dense architecture combining Gated Delta Networks with a feed-forward structure, activating its full parameter suite for every inference to maximize reliability. This design provides the highest instruction-following and coding accuracy in its class, with a notable IFEval score of 95.0. The model features a native 262K-token context window, extensible to 1M tokens via YaRN (RoPE scaling), and is released under the Apache-2.0 license.

Optimized for agentic workflows, Qwen3.5-27B employs an early-fusion architecture that treats visual and textual data as a unified stream for deep cross-modal reasoning. This unified approach allows the model to excel in technical analysis and software engineering, matching GPT-5-mini with a 72.4% score on SWE-bench Verified. While the larger MoE variants in the family lead in raw knowledge benchmarks, the 27B model offers a stable and high-density alternative for structured data extraction and spatial perception, contributing to the Qwen3.5 family’s generational leap in OCR accuracy over the previous Qwen3-VL series.

Qwen3.6 35B A3B

Qwen3.6-35B-A3B is a sparse Mixture-of-Experts (MoE) multimodal language model developed by the Qwen team at Alibaba Group. It carries 35 billion total parameters but activates only approximately 3 billion per forward pass via a learned routing mechanism, giving it the representational capacity of a large dense model at a fraction of the inference compute. The model is natively multimodal, processing images, documents, and video alongside text as a core architectural capability rather than an add-on. It supports a native context window of 262,144 tokens, extensible up to 1,010,000 tokens via YaRN. A key design feature is the unified thinking/non-thinking mode framework: users can switch between deliberate chain-of-thought reasoning and fast direct responses within a single model, and a "thinking preservation" option retains reasoning context across multi-turn agentic workflows to reduce redundant computation.

The model is specifically optimized for agentic coding tasks, including repository-level reasoning, frontend workflow generation, multi-step tool use, and MCP (Model Context Protocol) integration. On SWE-bench Verified it scores 73.4%, on Terminal-Bench 2.0 it scores 51.5%, and on MCPMark it scores 37.0%. For vision-language tasks it achieves 92.0 on RefCOCO, 89.9 on OmniDocBench 1.5, and 83.7 on VideoMMMU. The model also supports Multi-Token Prediction (MTP) for speculative decoding. All Qwen3.6 open-weight models are released under the Apache 2.0 license.