Kimi K2.5 is a frontier-scale multimodal AI model developed by Moonshot AI and released on January 27, 2026. As a significant advancement within the Kimi K2 family, it utilizes a sparse Mixture-of-Experts (MoE) architecture with 1 trillion total parameters (32 billion active per inference) and a massive 256K-token context window. The model features native multimodal integration via a 400M-parameter MoonViT encoder, allowing it to process text, images, and video frames simultaneously. Built for both speed and depth, it offers "Instant" and "Thinking" modes, the latter of which excels at expert-level reasoning, scoring 50.2% on the Humanity’s Last Exam (HLE) benchmark when equipped with tools.
The model is released under a Modified MIT License, which remains open-weight but requires attribution for high-revenue commercial entities. It introduces an "Agent Swarm" paradigm capable of coordinating up to 100 specialized sub-agents for parallel workflows, significantly reducing latency in complex research tasks. For vision tasks, Kimi K2.5 demonstrates strong autonomous visual debugging capabilities, where it can inspect its own generated UI outputs against visual specifications to iteratively refine frontend code. This makes it a powerful choice for developers testing automated UI reconstruction, high-fidelity OCR document processing, and multi-step agentic research grounded in complex visual data.
Drag and drop an image here, or click to browse
Captioning will run automatically
—
Usage
Past 30 DaysKimi K2.5 has not yet been evaluated on the current benchmark. The results below are from the legacy version of Vision Evals, our previous benchmark. See the current Vision Evals
| Category | Passed | Score |
|---|---|---|
| Document Understanding | 5 / 9 | 55.6% |
| Defect Detection | 7 / 15 | 46.7% |
| Object Understanding | 6 / 14 | 42.9% |
| Spatial Understanding | 5 / 19 | 26.3% |
| Object Counting | 1 / 10 | 10% |
| Category | Passed | Score |
|---|---|---|
| Handwritten Math | 5 / 10 | 50% |
| VQA & Extraction | 20 / 60 | 33.3% |
| Text Recognition | 8 / 30 | 26.7% |
| Focused Scene OCR | 10 / 99 | 10.1% |
| License Plate Recognition | 2 / 30 | 6.7% |
Scores based on a single evaluation run · Methodology
View all legacy Vision Evals results →Kimi K2.5 costs $0.570 per 1M input tokens and $2.85 per 1M output tokens.
Pricing updated Jul 24, 2026
Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.
6 of 6 models plotted
| Model | Score | Median tokens | Est. cost / task | Compare |
|---|---|---|---|---|
| Claude Haiku 4.5 | 58.2% | 2.3K | $0.0030 | Compare |
| GPT-5 Nano | 58.2% | 2.7K | $0.0003 | Compare |
| Qwen3.5 397B A17B | 58.2% | 1.5K | $0.0006 | Compare |
| Gemini 2.5 Flash | 55.2% | 476 | $0.0005 | Compare |
| Gemini 2.5 Flash-Lite | 53.7% | 301 | <$0.0001 | Compare |
| Kimi K2.5(this model) | 35.8% | 2.7K | $0.0031 | — |
Other models worth comparing for similar use cases.
License terms and commercial-use guidance for Kimi K2.5.
This model is released under a modified version of the MIT License. The base permissions of MIT apply, but additional terms or restrictions have been added by the model authors.
Commercial use is generally permitted, but the modified terms may add restrictions specific to this model. Review the full license text before deploying commercially.
License information is provided as a guide and is not legal advice.
Yes. Kimi K2.5 accepts image input, and on Roboflow's previous vision benchmark it passed 35.8% of visual understanding tasks (#74 of 77) and scored 19.7% on OCR. You can test it on your own image in the demo above.
Kimi K2.5 has not yet been evaluated on Roboflow's current Vision Evals. The results on this page are from the previous benchmark.
Yes. The demo on this page runs Kimi K2.5 in the free Roboflow Playground: upload an image and see results in seconds. A free account unlocks unlimited runs.