Roboflow

Kimi K2.5 Overview

Kimi K2.5 is a frontier-scale multimodal AI model developed by Moonshot AI and released on January 27, 2026. As a significant advancement within the Kimi K2 family, it utilizes a sparse Mixture-of-Experts (MoE) architecture with 1 trillion total parameters (32 billion active per inference) and a massive 256K-token context window. The model features native multimodal integration via a 400M-parameter MoonViT encoder, allowing it to process text, images, and video frames simultaneously. Built for both speed and depth, it offers "Instant" and "Thinking" modes, the latter of which excels at expert-level reasoning, scoring 50.2% on the Humanity’s Last Exam (HLE) benchmark when equipped with tools.

The model is released under a Modified MIT License, which remains open-weight but requires attribution for high-revenue commercial entities. It introduces an "Agent Swarm" paradigm capable of coordinating up to 100 specialized sub-agents for parallel workflows, significantly reducing latency in complex research tasks. For vision tasks, Kimi K2.5 demonstrates strong autonomous visual debugging capabilities, where it can inspect its own generated UI outputs against visual specifications to iteratively refine frontend code. This makes it a powerful choice for developers testing automated UI reconstruction, high-fidelity OCR document processing, and multi-step agentic research grounded in complex visual data.

Kimi K2.5 Interactive Demo

Results appear here. Add an image or pick an example to run Kimi K2.5.

Kimi K2.5 Details & Performance

Details

Resources

—

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRVision LanguageVisual Question Answering

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Performance

Avg. Latency

Kimi K2.5 Vision Evals

Kimi K2.5 has not yet been evaluated on the current benchmark. The results below are from the legacy version of Vision Evals, our previous benchmark. See the current Vision Evals

Visual Understanding

77 models · 67 tasks
HighestLowest
This model#74 of 7735.82% pass rate · better than 1%
Score35.82%pass rate across 67 tasks
Speed14.81savg response per task
Cost$0.0024 / task$0.450 in · $2.25 out / 1M
Tokens2.7K / task1.6K in · 766 out
Score key:≥75%40–74%<40%
CategoryPassedScore
Document Understanding5 / 9
55.6%
Defect Detection7 / 15
46.7%
Object Understanding6 / 14
42.9%
Spatial Understanding5 / 19
26.3%
Object Counting1 / 10
10%
HighestLowest
This model#58 of 5819.65% pass rate · better than 0%
Score19.65%pass rate across 229 tasks
Speed13.09savg response per task
Cost$0.0006 / task$0.450 in · $2.25 out / 1M
Tokens706 / task119 in · 258 out
Score key:≥75%40–74%<40%
CategoryPassedScore
Handwritten Math5 / 10
50%
VQA & Extraction20 / 60
33.3%
Text Recognition8 / 30
26.7%
Focused Scene OCR10 / 99
10.1%
License Plate Recognition2 / 30
6.7%

Scores based on a single evaluation run · Methodology

View all legacy Vision Evals results →

Kimi K2.5 Pricing

Kimi K2.5 costs $0.450 per 1M input tokens and $2.25 per 1M output tokens.

Input$0.450 / 1M tokens
Output$2.25 / 1M tokens
Cached input$0.070 / 1M tokens

Pricing updated Oct 3, 2026

Price vs. performance

Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.

6 of 6 models plotted

ModelScoreMedian tokensEst. cost / taskCompare
AnthropicClaude Haiku 4.558.2%2.3K$0.0030Compare
OpenAIGPT-5 Nano58.2%2.7K$0.0003Compare
QwenQwen3.5 397B A17B58.2%1.5K$0.0008Compare
GoogleGemini 2.5 Flash55.2%476$0.0005Compare
GoogleGemini 2.5 Flash-Lite53.7%301<$0.0001Compare
MoonshotAIKimi K2.5(this model)35.8%2.7K$0.0024—

Alternatives to Kimi K2.5

Other models worth comparing for similar use cases.

MiMo V2.6 Pro
MiMo V2.6 Pro is the flagship omni-modal foundation model in Xiaomi's MiMo V2.6 series, released as open weights alongside a Flash variant and a 9B distillation of Qwen3.5. It uses a sparse mixture-of-experts transformer with 1.02 trillion total parameters and roughly 42 billion activated per token, paired with a hybrid attention design that interleaves sliding-window and global attention layers to support a context window of about one million tokens. Dedicated encoders handle non-text inputs, including a vision encoder of roughly 681 million parameters and an audio tokenizer stack, so the model accepts text, images, video, and audio and returns text.Post-training centers on large-scale reinforcement learning across thousands of interactive environments, combined with agentic grading, self-correction cold start, and a multi-prefix multi-teacher on-policy distillation stage that extends behavior to tasks that are hard to verify automatically. The resulting model targets long-horizon agentic work such as software engineering, terminal and computer-use operation, tool calling, cybersecurity analysis, and visual coding, and it reports gains over the prior MiMo generation on SWE-bench Verified, Terminal Bench, and internal visual coding and cyber benchmarks.
Google
Gemini 3.1 Pro
Gemini 3.1 Pro is a proprietary multimodal model from Google’s Gemini 3 series, released in early 2026 and designed for advanced reasoning across large multimodal datasets. It accepts text, images, audio, video, and documents, supporting up to a 1-million-token input context with up to 64k output tokens. Compared with Gemini 3 Pro, it improves long-context synthesis and multi-step reasoning, enabling more reliable analysis of large documents, datasets, and software codebases.The model also advances visual understanding and grounding, allowing it to interpret UI screenshots, diagrams, and real-world scenes while referencing specific regions within images or video. These capabilities make Gemini 3.1 Pro well suited for multimodal workflows involving document processing, interface analysis, robotics research, and complex visual reasoning.
OpenAI
GPT-6 Sol
GPT-6 Sol is a proprietary multimodal reasoning model from OpenAI, released on September 22, 2026 alongside GPT-6 Luna as an efficiency-oriented tier of the GPT-6 family that began with GPT-6 Astra. OpenAI states that Sol and Luna are trained with methods similar to those used for Astra, carrying the same work on professional tasks, factuality, coding, computer use, and alignment into models that run faster. Sol accepts text and image input and returns text output, and OpenAI documents a context window of roughly one million tokens together with a knowledge cutoff of April 20, 2026.The model targets complex coding and agentic workflows and exposes a configurable reasoning effort setting with levels of none, low, medium, high, xhigh, and max, which trades latency and token consumption against answer quality. OpenAI reports results including 33.2% on AutomationBench at xhigh effort and 56.4% on Agents' Last Exam at max effort, while its reported DeepSWE and OSWorld 2.0 figures of 68.8% and 64.4% fall below those of the earlier GPT-5.6 Sol. Its vision behavior covers image understanding tasks such as visual question answering, captioning, document and chart interpretation, and text recognition.
Anthropic
Claude Opus 4.7
Claude Opus 4.7 is a proprietary multimodal language model developed by Anthropic, released on April 16, 2026. It is designed for agentic coding, long-horizon task execution, and enterprise knowledge work. The model supports text and vision inputs and operates with a context window of up to 1,000,000 tokens. It introduces adaptive thinking, which dynamically allocates reasoning based on task complexity, along with configurable effort controls including a new xhigh setting that sits between the existing high and max levels. It achieves 87.6% on SWE-bench Verified and 78.0% on OSWorld-Verified, reflecting strong performance on autonomous software engineering and computer use tasks respectively.Compared to Claude Opus 4.6, version 4.7 shows improved instruction following and higher reliability in extended agentic tasks. Vision capabilities now support high-resolution inputs up to 2,576px on the long edge (~3.75 megapixels), more than three times the resolution of prior Claude models, enabling finer interpretation of dense diagrams, UI screenshots, and document layouts. These improvements, combined with self-verification on long-running tasks and a new task budget system for controlling agentic loops, make it well-suited for complex software engineering, technical analysis, and multimodal vision workflows.
Grok
Grok 4.6
Grok 4.6 is a proprietary reasoning model from xAI aimed at long-running agentic workflows, coding, and knowledge work. It accepts text and image input and returns text, with a 500,000 token context window and a knowledge cutoff of February 1, 2026. The model exposes an adjustable reasoning budget with low, medium, high, and xhigh settings, where high is the default, and it supports function calling, structured outputs, web and X search, and code execution as documented tool behaviors. Its visual capability covers interpreting images supplied alongside text prompts, which places it in the visual question answering and document understanding family, and it can also return object detection boxes as text coordinates when prompted.xAI characterizes Grok 4.6 as the result of an extended post-training run over the Grok 4.5 lineage rather than a new pretrained base. The described recipe combines curated model-generated reasoning and technical data, engineering data, a revised optimizer, regenerated supervised fine-tuning trajectories, and reinforcement learning across agent environments spanning knowledge work, coding, kernel optimization, web development, and computer-aided design. Parameter count and architecture specifics are not disclosed. Independent measurement from Artificial Analysis places the model at 61 on its Intelligence Index, five points above Grok 4.5.

Kimi K2.5 License

Modified MIT · Permissive license, with added terms

Kimi K2.5 ships under a modified MIT license: MIT's permissive core with vendor-specific clauses layered on top. The Kimi K2.5 license is permissive in outline, but the added clauses are where commercial restrictions hide.

Commercial use
Usually permitted without a separate commercial license, though the added clauses can restrict it — named-user caps, industry carve-outs, or scale thresholds are common. Confirm against the Kimi K2.5 license text.
Modification
Permitted under the MIT core, subject to whatever the modified terms add. Derivative works may inherit the same additions.
Redistribution
Permitted with attribution, and the modified terms must travel with every copy you distribute.

Uncertainty around licensing can delay or stop a project. Diff the Kimi K2.5 terms against stock MIT, and read any acceptable-use policy attached to them, before you build on the model.

Read the full Modified MIT license ↗

Do I need a commercial license for Kimi K2.5?

If the added clauses rule out your use case, you need a commercial license from the rights holder. Roboflow's licensing page lists the models whose commercial license is included in a Roboflow plan, and which deployment methods it covers — Kimi K2.5 is worth checking against that list before you commit.

Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.

Talk to sales

This model is released under a modified version of the MIT License. The base permissions of MIT apply, but additional terms or restrictions have been added by the model authors.

Commercial use is generally permitted, but the modified terms may add restrictions specific to this model. Review the full license text before deploying commercially.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About Kimi K2.5 Vision

Yes. Kimi K2.5 accepts image input, and on Roboflow's previous vision benchmark it passed 35.8% of visual understanding tasks (#74 of 77) and scored 19.7% on OCR. You can test it on your own image in the demo above.

Kimi K2.5 has not yet been evaluated on Roboflow's current Vision Evals. The results on this page are from the previous benchmark.

Yes. The demo on this page runs Kimi K2.5 in the free Roboflow Playground: upload an image and see results in seconds. A free account unlocks unlimited runs.