Roboflow

Xiaomi Vision Models

Compare both Xiaomi vision models we track, all of them open-weight. Try them on your own images, free in the Roboflow Playground.

2 models · 2 open-weight · 2 free to try · Updated Sep 2026

About Xiaomi Vision Models

Every Xiaomi vision model in our catalog is open-weight. Both publish downloadable weights, so you can self-host under their own licenses instead of paying per request. The open-weight side runs on MIT terms. The newest additions are MiMo V2.6 Flash (Sep 2026) and MiMo V2.6 Pro (Sep 2026).

Between them, the Xiaomi models we list cover 10 distinct vision tasks. The widest coverage is captioning (2 models), chart question answering (2), and classification (2). For captioning, start with MiMo V2.6 Flash or MiMo V2.6 Pro.

On the open-weight tier, MiMo V2.6 Flash is the largest at 309B total, 15B active parameters and MiMo V2.6 Pro the smallest at 1.02T total, 42B active. Both carry a permissive license, so commercial use is straightforward.

With every model here open-weight, the real trade is size against generality: the smaller checkpoints fine-tune and deploy cheaply on your own hardware, while the larger ones cover more ground out of the box. Both Xiaomi vision models run live in the Roboflow Playground, so you can run the same image through both and compare the answers before committing to one.

Which Xiaomi Model Should You Use?

What each of the 2 Xiaomi vision models in our catalog is built for, and how you run it.

MiMo V2.6 Flash
Best for: Object Detection and OCR. Open weights: MIT license, 309B total, 15B active parameters. Self-host it or run it here. Runnable in the Playground.
MiMo V2.6 Pro
Best for: Object Detection and OCR. Open weights: MIT license, 1.02T total, 42B active parameters. Self-host it or run it here. Runnable in the Playground.

Open-Source Xiaomi Models

2 models with downloadable weights you can self-host under their licenses (MIT). All run live in the Playground through hosted APIs, so self-hosting is optional.

MiMo V2.6 Flash
NEW
MiMo-V2.6-Flash is the efficiency-oriented checkpoint of Xiaomi's MiMo-V2.6 series, a natively omnimodal foundation model that accepts text, image, video, and audio in a single model and supports a one million token context window. The language backbone is a sparse mixture-of-experts transformer with roughly 309 billion total parameters and 15 billion activated per token, organized as 48 layers with 256 routed experts and top-8 routing. It uses a hybrid attention scheme that interleaves sliding window attention with global attention layers to cut key-value cache cost on long sequences, and pairs the backbone with a vision encoder, an audio encoder, and an audio tokenizer, plus a multi-token prediction module and a draft model for faster decoding.Training emphasizes large scale reinforcement learning on verifiable, long-horizon tasks, with RL compute, environment diversity, and grader compute scaled together in a single mixed run. Xiaomi reports gains during RL on SWE-bench Verified, Terminal Bench, a cybersecurity benchmark, and an internal visual coding benchmark, reflecting a focus on agentic coding, computer use, and multimodal document and screen understanding rather than single turn chat.
MiMo V2.6 Pro
NEW
MiMo V2.6 Pro is the flagship omni-modal foundation model in Xiaomi's MiMo V2.6 series, released as open weights alongside a Flash variant and a 9B distillation of Qwen3.5. It uses a sparse mixture-of-experts transformer with 1.02 trillion total parameters and roughly 42 billion activated per token, paired with a hybrid attention design that interleaves sliding-window and global attention layers to support a context window of about one million tokens. Dedicated encoders handle non-text inputs, including a vision encoder of roughly 681 million parameters and an audio tokenizer stack, so the model accepts text, images, video, and audio and returns text.Post-training centers on large-scale reinforcement learning across thousands of interactive environments, combined with agentic grading, self-correction cold start, and a multi-prefix multi-teacher on-policy distillation stage that extends behavior to tasks that are hard to verify automatically. The resulting model targets long-horizon agentic work such as software engineering, terminal and computer-use operation, tool calling, cybersecurity analysis, and visual coding, and it reports gains over the prior MiMo generation on SWE-bench Verified, Terminal Bench, and internal visual coding and cyber benchmarks.

Frequently Asked Questions About Xiaomi Vision Models

Which Xiaomi models can do captioning?

2 of the 2 Xiaomi vision models we track handle captioning: MiMo V2.6 Flash and MiMo V2.6 Pro. Each model page lists its full task coverage, license, and specs.

Which Xiaomi models can do chart question answering?

2 of the 2 Xiaomi vision models we track handle chart question answering: MiMo V2.6 Flash and MiMo V2.6 Pro. Each model page lists its full task coverage, license, and specs.

Are Xiaomi vision models open source?

Yes. Both Xiaomi vision models we track publish downloadable weights you can self-host under their licenses (MIT).

What is the best Xiaomi model for captioning?

We do not publish a Xiaomi-only ranking, so pick on constraints rather than a label. 2 Xiaomi models handle captioning; the most recent is MiMo V2.6 Flash (Sep 2026). For fixed categories in production, a model fine-tuned on your own data typically beats any general-purpose model.

How many Xiaomi vision models are on Roboflow Playground?

We track 2 live Xiaomi vision models. The most recent addition is MiMo V2.6 Flash, released Sep 2026.

Can I try Xiaomi vision models for free?

Yes. Both Xiaomi models run live in the Roboflow Playground. Upload your own image, run several models on it at once, and compare the outputs side by side. No setup and no account required.

This page lists both Xiaomi vision models in the Roboflow Playground catalog, all of them open-weight and free to self-host. They cover captioning, chart question answering, and classification, among other tasks. All of them run live in the Roboflow Playground on your own images. Compare licenses, parameters, prices, and release dates side by side, or open any model page for full details.