MiMo-V2.6-Flash is the efficiency-oriented checkpoint of Xiaomi's MiMo-V2.6 series, a natively omnimodal foundation model that accepts text, image, video, and audio in a single model and supports a one million token context window. The language backbone is a sparse mixture-of-experts transformer with roughly 309 billion total parameters and 15 billion activated per token, organized as 48 layers with 256 routed experts and top-8 routing. It uses a hybrid attention scheme that interleaves sliding window attention with global attention layers to cut key-value cache cost on long sequences, and pairs the backbone with a vision encoder, an audio encoder, and an audio tokenizer, plus a multi-token prediction module and a draft model for faster decoding.
Training emphasizes large scale reinforcement learning on verifiable, long-horizon tasks, with RL compute, environment diversity, and grader compute scaled together in a single mixed run. Xiaomi reports gains during RL on SWE-bench Verified, Terminal Bench, a cybersecurity benchmark, and an internal visual coding benchmark, reflecting a focus on agentic coding, computer use, and multimodal document and screen understanding rather than single turn chat.
Drag and drop an image here, or click to browse
Model settings
Thinking level
Max output tokens
Default 65,536 · max 65,536
Sign in to adjust thinking and output length per run.
Results appear here. Add an image or pick an example to run MiMo V2.6 Flash.
—
Usage
Past 30 DaysVision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.
Evals updated September 22, 2026Pricing updated September 23, 2026
MiMo V2.6 Flash averages 60.6% across the six Vision Evals tasks, ranking #52 of 59 models overall.
Its weakest relative showing is Identification, ranking #52 of 59 at 76.0%.
At $0.0003 per sample it is the 5th cheapest of the 59 benchmarked models, and its average inference time of 9.2s per sample makes it the 33rd fastest.
Field medians: Object Detection 53.9%, Counting 59.5%, Identification 84.4%, OCR 88.7%, Data Extraction 84.5%, Reasoning 54.1%.
| Task | Score | Field (0 to 100) | Rank | Cost / sample | Speed |
|---|---|---|---|---|---|
| Object Detection (low) | 37.8% ±1.1, Mean of 3 runs, range 36.4 to 38.7 | #47 of 59 | $0.0004 | 13.02s | |
| Object Detection (high) | 45.0% ±2.5, Mean of 3 runs, range 42.2 to 47.1 | #20 of 25 | $0.0009 | 28.18s | |
| Counting (low) | 49.5% ±8.1, Mean of 3 runs, range 41.9 to 58.1 | #45 of 59 | $0.0002 | 5.76s | |
| Counting (high) | 64.9% ±1.4, Mean of 3 runs, range 63.5 to 66.2 | #18 of 25 | $0.0003 | 15.63s | |
| Identification (low) | 76.0% ±3.1, Mean of 3 runs, range 71.9 to 78.1 | #52 of 59 | $0.0001 | 6.00s | |
| Identification (high) | 82.3% ±4.7, Mean of 3 runs, range 78.1 to 87.5 | #21 of 25 | $0.0003 | 16.68s | |
| OCR (low) | 87.0% ±0.5, Mean of 3 runs, range 86.6 to 87.7 | #43 of 59 | $0.0003 | 15.58s | |
| OCR (high) | 87.0% ±2.1, Mean of 3 runs, range 84.3 to 88.5 | #24 of 25 | $0.0016 | 73.66s | |
| Data Extraction (low) | 80.1% ±1.0, Mean of 3 runs, range 79.4 to 81.4 | #43 of 59 | $0.0002 | 5.71s | |
| Data Extraction (high) | 82.5% ±1.0, Mean of 3 runs, range 81.4 to 83.5 | #17 of 25 | $0.0004 | 15.30s | |
| Reasoning (low) | 33.1% ±2.0, Mean of 3 runs, range 31.1 to 35.1 | #49 of 59 | $0.0002 | 5.90s | |
| Reasoning (high) | 58.5% ±1.3, Mean of 3 runs, range 57.0 to 59.6 | #38 of 45 | $0.0008 | 36.78s |
Overall benchmark score against estimated cost per sample, on a log scale. Upper-left is the sweet spot: high quality at low cost.
58 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort · MiMo V2.6 Flash highlighted
MiMo V2.6 Flash scores are the mean of 3 runs per task at both low and high effort · Methodology
View all Vision Evals →MiMo V2.6 Flash costs $0.140 per 1M input tokens and $0.280 per 1M output tokens.
Pricing updated Sep 23, 2026
MiMo V2.6 Flash is released under MIT, a permissive license. The MiMo V2.6 Flash license lets you use, modify, and sell work built on the model, with the copyright notice as the only real obligation and no requirement to open-source related code changes.
MIT grants no explicit patent license and disclaims all warranties. If patent exposure is a concern for your deployment, review it with counsel before launch.
Read the full MIT license ↗No commercial license is needed for MiMo V2.6 Flash: permissive terms let you keep related code private while deploying commercially.
Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.
Talk to salesThis model is released under the MIT License, a short and permissive open-source license that allows commercial use, modification, and redistribution.
Yes. Under the terms of the MIT license, you can freely use this model for commercial purposes. You must retain the copyright notice and license text when redistributing.
License information is provided as a guide and is not legal advice.
Yes. MiMo V2.6 Flash accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is OCR at 87.1% (#43 of 59 at low effort). You can test it on your own image in the demo above.
Yes. its transcriptions match the ground truth 87.1% on average (#43 of 59 at low effort) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 80.1%.
Not its strength. On Vision Evals, MiMo V2.6 Flash scores 37.8% mAP@50 on object detection (#47 of 59 at low effort) and 49.6% judge-graded accuracy on object counting. For production counting or precise localization, pairing it with a specialized detector like RF-DETR or your own trained model in a Roboflow Workflow is usually more reliable: detect the objects, then count the detections.
On our benchmark's task mix, MiMo V2.6 Flash averages $0.0003 per sample at $0.14 per 1M input and $0.28 per 1M output tokens (#5 of 59 on cost), with an average speed of 9.2s per sample across the benchmark. Actual cost depends on your images and prompts.
On the overall Vision Evals ranking, MiMo V2.6 Flash sits #52 of 59 at 60.6%, just behind Qwen3.8 27B (61.2%) and just ahead of GLM-4.6V Flash (59.7%). See the full side-by-side: MiMo V2.6 Flash vs Qwen3.8 27B.