Roboflow

Mistral Vision Models

Compare all 3 Mistral vision models we track, 2 of them open-weight. Try them on your own images, free in the Roboflow Playground.

3 models · 2 open-weight · 3 free to try · Updated Aug 2025

About Mistral Vision Models

Mistral ships vision models on two tiers, and our catalog tracks 3 of them. 2 are open-weight: the weights are downloadable, so you can self-host under their own licenses. The other 1 is proprietary, reached through an API and billed by the provider. The open-weight side runs on Apache 2.0 terms. The newest additions are Mistral Medium 3.1 (Aug 2025) and Mistral Small 3.1 24B (Mar 2025).

Between them, the Mistral models we list cover 9 distinct vision tasks. The widest coverage is captioning (3 models), chart question answering (3), and classification (3). For captioning, start with Mistral Medium 3.1, Mistral Small 3.1 24B, or Pixtral 12B.

On the open-weight tier, Mistral Small 3.1 24B is the largest at 24B parameters and Pixtral 12B the smallest at 12B. Both carry a permissive license, so commercial use is straightforward.

On the API tier, the one option is Mistral Medium 3.1 (Aug 2025), billed per token by the provider rather than run on your own hardware. Which tier to start on is a constraints question, not a quality one: reach for the open-weight side when you need offline inference, predictable per-image cost, or a checkpoint you can fine-tune on your own data, and for the API side when you want broad general reasoning without managing GPUs. All 3 Mistral vision models run live in the Roboflow Playground, so you can run the same image through several of them and compare the answers before committing to one.

Which Mistral Model Should You Use?

What each of the 3 Mistral vision models in our catalog is built for, and how you run it.

Mistral Medium 3.1
Best for: OCR and Document Question Answering. Proprietary: no downloadable weights, billed per token through Mistral's API. Runnable in the Playground.
Mistral Small 3.1 24B
Best for: OCR and Document Question Answering. Open weights: Apache 2.0 license, 24B parameters. Self-host it or run it here. Runnable in the Playground.
Pixtral 12B
Best for: OCR and Document Question Answering. Open weights: Apache 2.0 license, 12B parameters. Self-host it or run it here. Runnable in the Playground.

Open-Source Mistral Models

2 models with downloadable weights you can self-host under their licenses (Apache 2.0). All run live in the Playground through hosted APIs, so self-hosting is optional.

Mistral
Mistral Small 3.1 24B
Mistral Small 3.1 24B, released on March 17, 2025, is an open-weight multimodal model from Mistral AI, distributed under the Apache-2.0 license. With around 24B parameters and a 128K token context window, it is available in both base and instruction-tuned (“Instruct”) variants. The model introduces vision support alongside text, enabling tasks like multimodal reasoning, captioning, and image-based Q&A.It is multilingual, supporting many languages, and is optimized for fast responses, function calling, structured dialogue, and long-context reasoning. Despite its size, the model can be run locally in quantized formats, fitting on machines with ~32GB RAM, making it accessible to developers outside large cloud setups. However, the output length is smaller than the 128K input window, meaning long generations may require chaining. In addition, using full vision features or the maximum context window significantly increases compute costs, and performance on highly complex reasoning or enterprise-scale tasks still trails larger proprietary frontier models.
Mistral
Pixtral 12B
Pixtral-12B is a vision-language model introduced by Mistral AI in September 2024 under the Apache 2.0 license, designed to process both text and images in a unified context. With ~12 billion parameters in its decoder and an additional ~400 million in a custom-trained vision encoder, it supports long-context reasoning up to 128k tokens and accepts multiple images per input. Its architecture is optimized for handling variable image sizes and aspect ratios, making it flexible for diverse multimodal tasks.As Mistral’s first VLM, Pixtral-12B delivers strong performance not only on image-text reasoning benchmarks but also in text-only applications, positioning it as a versatile alternative to models like GPT-4V and LLaVA. Its open availability via Hugging Face and major cloud providers such as Amazon Bedrock and SageMaker makes it accessible for research and production. Typical use cases include document analysis, visual QA, data extraction, and multimodal assistants requiring both textual and visual understanding.

Mistral Models via API

1 proprietary model where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.

Frequently Asked Questions About Mistral Vision Models

Which Mistral models can do captioning?

3 of the 3 Mistral vision models we track handle captioning: Mistral Medium 3.1, Mistral Small 3.1 24B, and Pixtral 12B. Each model page lists its full task coverage, license, and specs.

Which Mistral models can do chart question answering?

3 of the 3 Mistral vision models we track handle chart question answering: Mistral Medium 3.1, Mistral Small 3.1 24B, and Pixtral 12B. Each model page lists its full task coverage, license, and specs.

Are Mistral vision models open source?

Partly. 2 of the 3 Mistral vision models we track are open-weight (for example Mistral Small 3.1 24B and Pixtral 12B), downloadable and self-hostable under their licenses (Apache 2.0). The other 1 are proprietary and reached through Mistral's API.

What is the best Mistral model for captioning?

We do not publish a Mistral-only ranking, so pick on constraints rather than a label. 3 Mistral models handle captioning; the most recent is Mistral Medium 3.1 (Aug 2025), and Mistral Small 3.1 24B is the newest open-weight option if you need to self-host. For fixed categories in production, a model fine-tuned on your own data typically beats any general-purpose model.

How many Mistral vision models are on Roboflow Playground?

We track 3 live Mistral vision models, 2 open-weight and 1 available through an API. The most recent addition is Mistral Medium 3.1, released Aug 2025.

Can I try Mistral vision models for free?

Yes. All 3 Mistral models run live in the Roboflow Playground. Upload your own image, run several models on it at once, and compare the outputs side by side. No setup and no account required.

This page lists all 3 Mistral vision models in the Roboflow Playground catalog: 2 open-weight models you can self-host and 1 proprietary model accessed through an API. They cover captioning, chart question answering, and classification, among other tasks. All of them run live in the Roboflow Playground on your own images. Compare licenses, parameters, prices, and release dates side by side, or open any model page for full details.