MiMo V2.6 Pro is the flagship omni-modal foundation model in Xiaomi's MiMo V2.6 series, released as open weights alongside a Flash variant and a 9B distillation of Qwen3.5. It uses a sparse mixture-of-experts transformer with 1.02 trillion total parameters and roughly 42 billion activated per token, paired with a hybrid attention design that interleaves sliding-window and global attention layers to support a context window of about one million tokens. Dedicated encoders handle non-text inputs, including a vision encoder of roughly 681 million parameters and an audio tokenizer stack, so the model accepts text, images, video, and audio and returns text.
Post-training centers on large-scale reinforcement learning across thousands of interactive environments, combined with agentic grading, self-correction cold start, and a multi-prefix multi-teacher on-policy distillation stage that extends behavior to tasks that are hard to verify automatically. The resulting model targets long-horizon agentic work such as software engineering, terminal and computer-use operation, tool calling, cybersecurity analysis, and visual coding, and it reports gains over the prior MiMo generation on SWE-bench Verified, Terminal Bench, and internal visual coding and cyber benchmarks.
Drag and drop an image here, or click to browse
Model settings
Thinking level
Max output tokens
Default 65,536 · max 65,536
Sign in to adjust thinking and output length per run.
Results appear here. Add an image or pick an example to run MiMo V2.6 Pro.
—
Usage
Past 30 DaysVision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.
Evals updated September 22, 2026Pricing updated September 23, 2026
MiMo V2.6 Pro averages 62.5% across the six Vision Evals tasks, ranking #49 of 59 models overall.
Its weakest relative showing is Identification, ranking #52 of 59 at 76.0%.
At $0.0008 per sample it is the 15th cheapest of the 59 benchmarked models, and its average inference time of 8.5s per sample makes it the 26th fastest.
Field medians: Object Detection 53.9%, Counting 59.5%, Identification 84.4%, OCR 88.7%, Data Extraction 84.5%, Reasoning 54.1%.
| Task | Score | Field (0 to 100) | Rank | Cost / sample | Speed |
|---|---|---|---|---|---|
| Object Detection (low) | 42.0% ±1.1, Mean of 3 runs, range 40.9 to 43.1 | #40 of 59 | $0.0014 | 15.50s | |
| Object Detection (high) | 46.7% ±0.8, Mean of 3 runs, range 45.7 to 47.3 | #19 of 25 | $0.0030 | 45.14s | |
| Counting (low) | 50.0% ±2.0, Mean of 3 runs, range 48.6 to 52.7 | #43 of 59 | $0.0005 | 3.47s | |
| Counting (high) | 59.0% ±5.4, Mean of 3 runs, range 52.7 to 63.5 | #23 of 25 | $0.0016 | 42.05s | |
| Identification (low) | 76.0% ±1.6, Mean of 3 runs, range 75.0 to 78.1 | #52 of 59 | $0.0004 | 3.33s | |
| Identification (high) | 78.1% ±4.7, Mean of 3 runs, range 71.9 to 81.3 | #25 of 25 | $0.0012 | 28.01s | |
| OCR (low) | 90.7% ±1.7, Mean of 3 runs, range 88.5 to 91.9 | #21 of 59 | $0.0008 | 8.11s | |
| OCR (high) | 87.5% ±2.7, Mean of 3 runs, range 85.3 to 90.6 | #22 of 25 | $0.0048 | 101.54s | |
| Data Extraction (low) | 81.1% ±0.5, Mean of 3 runs, range 80.4 to 81.4 | #39 of 59 | $0.0005 | 3.26s | |
| Data Extraction (high) | 80.4% ±1.5, Mean of 3 runs, range 79.4 to 82.5 | #23 of 25 | $0.0013 | 25.71s | |
| Reasoning (low) | 35.1% ±2.6, Mean of 3 runs, range 32.5 to 37.8 | #45 of 59 | $0.0005 | 3.82s | |
| Reasoning (high) | 55.9% ±2.3, Mean of 3 runs, range 54.3 to 58.9 | #39 of 45 | $0.0026 | 68.29s |
Overall benchmark score against estimated cost per sample, on a log scale. Upper-left is the sweet spot: high quality at low cost.
58 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort · MiMo V2.6 Pro highlighted
MiMo V2.6 Pro scores are the mean of 3 runs per task at both low and high effort · Methodology
View all Vision Evals →MiMo V2.6 Pro costs $0.435 per 1M input tokens and $0.870 per 1M output tokens.
Pricing updated Sep 23, 2026
MiMo V2.6 Pro is released under MIT, a permissive license. The MiMo V2.6 Pro license lets you use, modify, and sell work built on the model, with the copyright notice as the only real obligation and no requirement to open-source related code changes.
MIT grants no explicit patent license and disclaims all warranties. If patent exposure is a concern for your deployment, review it with counsel before launch.
Read the full MIT license ↗No commercial license is needed for MiMo V2.6 Pro: permissive terms let you keep related code private while deploying commercially.
Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.
Talk to salesThis model is released under the MIT License, a short and permissive open-source license that allows commercial use, modification, and redistribution.
Yes. Under the terms of the MIT license, you can freely use this model for commercial purposes. You must retain the copyright notice and license text when redistributing.
License information is provided as a guide and is not legal advice.
Yes. MiMo V2.6 Pro accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is OCR at 90.7% (#21 of 59 at low effort). You can test it on your own image in the demo above.
Yes. its transcriptions match the ground truth 90.7% on average (#21 of 59 at low effort) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 81.1%.
Not its strength. On Vision Evals, MiMo V2.6 Pro scores 42% mAP@50 on object detection (#40 of 59 at low effort) and 50% judge-graded accuracy on object counting. For production counting or precise localization, pairing it with a specialized detector like RF-DETR or your own trained model in a Roboflow Workflow is usually more reliable: detect the objects, then count the detections.
On our benchmark's task mix, MiMo V2.6 Pro averages $0.0008 per sample at $0.43 per 1M input and $0.87 per 1M output tokens (#15 of 59 on cost), with an average speed of 8.5s per sample across the benchmark. Actual cost depends on your images and prompts.
On the overall Vision Evals ranking, MiMo V2.6 Pro sits #49 of 59 at 62.5%, just behind Qwen3.6 27B (62.7%) and just ahead of Qwen3.7 Flash (61.5%). See the full side-by-side: MiMo V2.6 Pro vs Qwen3.6 27B.