Roboflow

LLaVA-1.5 vs Qwen3.5 9b

Compare LLaVA-1.5 and Qwen3.5 9b side-by-side.

Compare LLaVA-1.5 vs Qwen3.5 9b live

Run the same image across every model that supports a task and compare their outputs side-by-side.

These models don't share enough common tasks for a side-by-side demo. See the comparison table below for their capabilities.

Models in this comparison

LLaVA-1.5 vs Qwen3.5 9b Comparison Table

Evals updated September 5, 2026Pricing updated September 20, 2026

PropertyLLaVA-1.5Qwen3.5 9b
OrganizationMicrosoftQwen
Categoryopenopen
Modalitymultimodalmultimodal
Release DateOct 2023Mar 2026
Context Window262K
Parameters7B, 13B9B
LicenseCustomApache 2.0
Pricing per 1M tokens
Input $/1M$0.100
Output $/1M$0.150
Vision Tasks
Vision Language
Visual Question AnsweringDemo
CaptioningDemo
Chart Question Answering
Classification
Document Question Answering
Image Tagging
Multi-Label Classification
Object Detection
OCRDemo
Model Features
LLMs with Vision Capabilities
Multimodal Vision
Foundation Vision
Vision Evalsground-truth scores across 6 vision tasks
OverallNot evaluated
64.3%
Quantizationsself-hosted
BF1664.8%FP864.4%AWQ-INT464.3%hardware →
Avg cost / sample$0.0017
Avg speed / sample33.65s
By task
Object Detection
46.6%
±0.6, Mean of 3 runs, range 45.8 to 47.0
$0
Counting
51.8%
±2.7, Mean of 3 runs, range 48.6 to 54.0
$0
Identification
83.3%
±1.6, Mean of 3 runs, range 81.3 to 84.4
$0
OCR
77.5%
±7.3, Mean of 3 runs, range 68.4 to 83.0
$0
Data Extraction
78.7%
±2.1, Mean of 3 runs, range 76.3 to 80.4
$0
Reasoning
47.7%
±2.3, Mean of 3 runs, range 45.7 to 50.3
$0

LLaVA-1.5 vs Qwen3.5 9b: Overview

LLaVA-1.5

LLaVA-1.5 is an open-source large multimodal model released in October 2023 by researchers at the University of Wisconsin-Madison and Microsoft Research. It builds on the original LLaVA architecture by introducing targeted refinements: switching the vision encoder to CLIP-ViT-L at 336-pixel resolution, replacing the projection layer with a two-layer MLP, and adding academic-task-oriented visual question answering data with response formatting prompts during training. These modifications achieve state-of-the-art performance across 11 benchmarks at release, with training completing in approximately one day on a single 8-A100 node.

The model accepts an image paired with a text prompt and generates natural language responses, supporting visual question answering, image captioning, and open-ended visual conversation. LLaVA-1.5 is available in 7B and 13B parameter variants built on the Vicuna language model, and is distributed under the Llama 2 Community License due to its Llama-2-based foundation. The original LLaVA paper was presented as an oral at NeurIPS 2023. Subsequent releases in the series (LLaVA-NeXT (LLaVA-1.6), LLaVA-NeXT-Video, and LLaVA-OneVision) are separate models with their own release pages and build on this foundation with expanded OCR, video, and multi-image capabilities.

Qwen3.5 9b

Qwen3.5-9B is a 9-billion-parameter multimodal foundation model developed by Alibaba Cloud's Qwen team, released on March 2, 2026 as part of the Qwen3.5 model family. Designed for efficient multimodal reasoning and long-context language tasks, it notably outperforms the older Qwen3-30B, a model more than three times its size, on key benchmarks including GPQA Diamond, IFEval, and LongBench.

The model supports vision-language inputs through an early-fusion multimodal architecture built on a dense hybrid foundation of Gated Delta Networks and Gated Attention. It can also operate in a text-only mode by skipping the vision encoder during inference. It provides a 262,144-token context window (extensible to ~1M tokens via YaRN) and is released under the Apache License 2.0. Within the current AI landscape, Qwen3.5-9B offers a strong balance of capability and efficiency, making it well-suited for multimodal assistants, document analysis, long-context reasoning, and developer-deployed agentic systems.