Roboflow

Xiaomi: MiMo V2.6 Flash

MiMo V2.6 Flash Overview

MiMo-V2.6-Flash is the efficiency-oriented checkpoint of Xiaomi's MiMo-V2.6 series, a natively omnimodal foundation model that accepts text, image, video, and audio in a single model and supports a one million token context window. The language backbone is a sparse mixture-of-experts transformer with roughly 309 billion total parameters and 15 billion activated per token, organized as 48 layers with 256 routed experts and top-8 routing. It uses a hybrid attention scheme that interleaves sliding window attention with global attention layers to cut key-value cache cost on long sequences, and pairs the backbone with a vision encoder, an audio encoder, and an audio tokenizer, plus a multi-token prediction module and a draft model for faster decoding.

Training emphasizes large scale reinforcement learning on verifiable, long-horizon tasks, with RL compute, environment diversity, and grader compute scaled together in a single mixed run. Xiaomi reports gains during RL on SWE-bench Verified, Terminal Bench, a cybersecurity benchmark, and an internal visual coding benchmark, reflecting a focus on agentic coding, computer use, and multimodal document and screen understanding rather than single turn chat.

MiMo V2.6 Flash Interactive Demo

Model settings

Thinking level

Max output tokens

Default 65,536 · max 65,536

Sign in to adjust thinking and output length per run.

Results appear here. Add an image or pick an example to run MiMo V2.6 Flash.

MiMo V2.6 Flash Details & Performance

Details

Resources

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRVision LanguageVisual Question Answering

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Performance

Avg. Latency

MiMo V2.6 Flash Vision Evals

Vision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.

Evals updated September 22, 2026Pricing updated September 23, 2026

Overall score#52 of 59
60.6%
Avg cost / sample#5 of 59
$0.0003
Avg speed / sample#33 of 59
9.20s
Avg tokens / sample
1.7K

Strengths and weaknesses

MiMo V2.6 Flash averages 60.6% across the six Vision Evals tasks, ranking #52 of 59 models overall.

Its weakest relative showing is Identification, ranking #52 of 59 at 76.0%.

At $0.0003 per sample it is the 5th cheapest of the 59 benchmarked models, and its average inference time of 9.2s per sample makes it the 33rd fastest.

Performance profile

Field medianMiMo V2.6 Flash

Field medians: Object Detection 53.9%, Counting 59.5%, Identification 84.4%, OCR 88.7%, Data Extraction 84.5%, Reasoning 54.1%.

Results by task

TaskScoreField (0 to 100)RankCost / sampleSpeed
Object Detection (low)
37.8%
±1.1, Mean of 3 runs, range 36.4 to 38.7
#47 of 59$0.000413.02s
Object Detection (high)
45.0%
±2.5, Mean of 3 runs, range 42.2 to 47.1
#20 of 25$0.000928.18s
Counting (low)
49.5%
±8.1, Mean of 3 runs, range 41.9 to 58.1
#45 of 59$0.00025.76s
Counting (high)
64.9%
±1.4, Mean of 3 runs, range 63.5 to 66.2
#18 of 25$0.000315.63s
Identification (low)
76.0%
±3.1, Mean of 3 runs, range 71.9 to 78.1
#52 of 59$0.00016.00s
Identification (high)
82.3%
±4.7, Mean of 3 runs, range 78.1 to 87.5
#21 of 25$0.000316.68s
OCR (low)
87.0%
±0.5, Mean of 3 runs, range 86.6 to 87.7
#43 of 59$0.000315.58s
OCR (high)
87.0%
±2.1, Mean of 3 runs, range 84.3 to 88.5
#24 of 25$0.001673.66s
Data Extraction (low)
80.1%
±1.0, Mean of 3 runs, range 79.4 to 81.4
#43 of 59$0.00025.71s
Data Extraction (high)
82.5%
±1.0, Mean of 3 runs, range 81.4 to 83.5
#17 of 25$0.000415.30s
Reasoning (low)
33.1%
±2.0, Mean of 3 runs, range 31.1 to 35.1
#49 of 59$0.00025.90s
Reasoning (high)
58.5%
±1.3, Mean of 3 runs, range 57.0 to 59.6
#38 of 45$0.000836.78s
  • Thinking longer helps: 7.3 points higher on object detection at high effort for 2.3x the cost and 2.2x the latency.
  • Thinking longer helps: 15.3 points higher on counting at high effort for 2.2x the cost and 2.7x the latency.
  • Thinking longer helps: 6.3 points higher on identification at high effort for 2.5x the cost and 2.8x the latency.
  • Thinking longer does not help: 0.1 points lower on ocr at high effort for 5.6x the cost and 4.7x the latency.
  • Thinking longer helps: 2.4 points higher on data extraction at high effort for 2.2x the cost and 2.7x the latency.
  • Thinking longer helps: 25.4 points higher on reasoning at high effort for 5x the cost and 6.2x the latency.

Price vs. performance

Score vs. cost

Overall benchmark score against estimated cost per sample, on a log scale. Upper-left is the sweet spot: high quality at low cost.

58 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort · MiMo V2.6 Flash highlighted

MiMo V2.6 Flash scores are the mean of 3 runs per task at both low and high effort · Methodology

View all Vision Evals →

MiMo V2.6 Flash Pricing

MiMo V2.6 Flash costs $0.140 per 1M input tokens and $0.280 per 1M output tokens.

Input$0.140 / 1M tokens
Output$0.280 / 1M tokens
Cached input$0.003 / 1M tokens

Pricing updated Sep 23, 2026

MiMo V2.6 Flash License

MIT · Permissive license

MiMo V2.6 Flash is released under MIT, a permissive license. The MiMo V2.6 Flash license lets you use, modify, and sell work built on the model, with the copyright notice as the only real obligation and no requirement to open-source related code changes.

Commercial use
Permitted with no separate commercial license. No usage caps, revenue thresholds, or field-of-use limits apply to MiMo V2.6 Flash.
Modification
Permitted. You can fine-tune or rewrite MiMo V2.6 Flash and keep the result closed-source.
Redistribution
Permitted. Include the original copyright and permission notice in copies or substantial portions of the work.

MIT grants no explicit patent license and disclaims all warranties. If patent exposure is a concern for your deployment, review it with counsel before launch.

Read the full MIT license ↗

Do I need a commercial license for MiMo V2.6 Flash?

No commercial license is needed for MiMo V2.6 Flash: permissive terms let you keep related code private while deploying commercially.

Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.

Talk to sales

This model is released under the MIT License, a short and permissive open-source license that allows commercial use, modification, and redistribution.

Yes. Under the terms of the MIT license, you can freely use this model for commercial purposes. You must retain the copyright notice and license text when redistributing.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About MiMo V2.6 Flash Vision

Yes. MiMo V2.6 Flash accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is OCR at 87.1% (#43 of 59 at low effort). You can test it on your own image in the demo above.

Yes. its transcriptions match the ground truth 87.1% on average (#43 of 59 at low effort) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 80.1%.

Not its strength. On Vision Evals, MiMo V2.6 Flash scores 37.8% mAP@50 on object detection (#47 of 59 at low effort) and 49.6% judge-graded accuracy on object counting. For production counting or precise localization, pairing it with a specialized detector like RF-DETR or your own trained model in a Roboflow Workflow is usually more reliable: detect the objects, then count the detections.

On our benchmark's task mix, MiMo V2.6 Flash averages $0.0003 per sample at $0.14 per 1M input and $0.28 per 1M output tokens (#5 of 59 on cost), with an average speed of 9.2s per sample across the benchmark. Actual cost depends on your images and prompts.

On the overall Vision Evals ranking, MiMo V2.6 Flash sits #52 of 59 at 60.6%, just behind Qwen3.8 27B (61.2%) and just ahead of GLM-4.6V Flash (59.7%). See the full side-by-side: MiMo V2.6 Flash vs Qwen3.8 27B.