Roboflow

Kimi K2.5 Overview

Kimi K2.5 is a frontier-scale multimodal AI model developed by Moonshot AI and released on January 27, 2026. As a significant advancement within the Kimi K2 family, it utilizes a sparse Mixture-of-Experts (MoE) architecture with 1 trillion total parameters (32 billion active per inference) and a massive 256K-token context window. The model features native multimodal integration via a 400M-parameter MoonViT encoder, allowing it to process text, images, and video frames simultaneously. Built for both speed and depth, it offers "Instant" and "Thinking" modes, the latter of which excels at expert-level reasoning, scoring 50.2% on the Humanity’s Last Exam (HLE) benchmark when equipped with tools.

The model is released under a Modified MIT License, which remains open-weight but requires attribution for high-revenue commercial entities. It introduces an "Agent Swarm" paradigm capable of coordinating up to 100 specialized sub-agents for parallel workflows, significantly reducing latency in complex research tasks. For vision tasks, Kimi K2.5 demonstrates strong autonomous visual debugging capabilities, where it can inspect its own generated UI outputs against visual specifications to iteratively refine frontend code. This makes it a powerful choice for developers testing automated UI reconstruction, high-fidelity OCR document processing, and multi-step agentic research grounded in complex visual data.

Kimi K2.5 Interactive Demo

Results appear here. Add an image or pick an example to run Kimi K2.5.

Kimi K2.5 Details & Performance

Details

Resources

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRVision LanguageVisual Question Answering

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Performance

Avg. Latency

Kimi K2.5 Vision Evals

Kimi K2.5 has not yet been evaluated on the current benchmark. The results below are from the legacy version of Vision Evals, our previous benchmark. See the current Vision Evals

Visual Understanding

77 models · 67 tasks
HighestLowest
This model#74 of 7735.82% pass rate · better than 1%
Score35.82%pass rate across 67 tasks
Speed14.81savg response per task
Cost$0.0024 / task$0.450 in · $2.25 out / 1M
Tokens2.7K / task1.6K in · 766 out
Score key:≥75%40–74%<40%
CategoryPassedScore
Document Understanding5 / 9
55.6%
Defect Detection7 / 15
46.7%
Object Understanding6 / 14
42.9%
Spatial Understanding5 / 19
26.3%
Object Counting1 / 10
10%
HighestLowest
This model#58 of 5819.65% pass rate · better than 0%
Score19.65%pass rate across 229 tasks
Speed13.09savg response per task
Cost$0.0006 / task$0.450 in · $2.25 out / 1M
Tokens706 / task119 in · 258 out
Score key:≥75%40–74%<40%
CategoryPassedScore
Handwritten Math5 / 10
50%
VQA & Extraction20 / 60
33.3%
Text Recognition8 / 30
26.7%
Focused Scene OCR10 / 99
10.1%
License Plate Recognition2 / 30
6.7%

Scores based on a single evaluation run · Methodology

View all legacy Vision Evals results →

Kimi K2.5 Pricing

Kimi K2.5 costs $0.450 per 1M input tokens and $2.25 per 1M output tokens.

Input$0.450 / 1M tokens
Output$2.25 / 1M tokens
Cached input$0.070 / 1M tokens

Pricing updated Sep 7, 2026

Price vs. performance

Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.

6 of 6 models plotted

ModelScoreMedian tokensEst. cost / taskCompare
AnthropicClaude Haiku 4.558.2%2.3K$0.0030Compare
OpenAIGPT-5 Nano58.2%2.7K$0.0003Compare
QwenQwen3.5 397B A17B58.2%1.5K$0.0006Compare
GoogleGemini 2.5 Flash55.2%476$0.0005Compare
GoogleGemini 2.5 Flash-Lite53.7%301<$0.0001Compare
MoonshotAIKimi K2.5(this model)35.8%2.7K$0.0024

Alternatives to Kimi K2.5

Other models worth comparing for similar use cases.

Google
Gemini 3.1 Pro
Gemini 3.1 Pro is a proprietary multimodal model from Google’s Gemini 3 series, released in early 2026 and designed for advanced reasoning across large multimodal datasets. It accepts text, images, audio, video, and documents, supporting up to a 1-million-token input context with up to 64k output tokens. Compared with Gemini 3 Pro, it improves long-context synthesis and multi-step reasoning, enabling more reliable analysis of large documents, datasets, and software codebases.The model also advances visual understanding and grounding, allowing it to interpret UI screenshots, diagrams, and real-world scenes while referencing specific regions within images or video. These capabilities make Gemini 3.1 Pro well suited for multimodal workflows involving document processing, interface analysis, robotics research, and complex visual reasoning.
Anthropic
Claude Opus 4.6
Claude Opus 4.6 is the flagship large language model from Anthropic, released on 2026-02-05 for advanced reasoning, complex coding, and enterprise agent workflows. It supports text and image inputs via API, offers a 200K-token standard context window with a 1M-token beta option, and enables outputs up to 128K tokens, with adaptive reasoning and context compaction for sustained tasks.As of 2026-02-17, Anthropic also released Claude Sonnet 4.6, extending the 1M-token context window to a broader tier. Opus remains positioned for maximum depth and benchmark performance, while Sonnet 4.6 brings long-context capability to more cost- and latency-sensitive production use cases.
Qwen
Qwen3.6 Plus
Qwen3.6 Plus is a flagship model in Alibaba’s Qwen Plus series, designed for agentic workflows, coding, and multi-step reasoning. It supports a 1 million token context window and up to 65,536 output tokens, with built-in reasoning capabilities. The model is available as a hosted, proprietary API through Alibaba Cloud. Compared to Qwen3.5, it improves reliability in multi-step execution and frontend code generation, with stronger performance on agentic coding tasks. It also supports document and image understanding, though its vision capabilities are more limited than dedicated Qwen-VL models. Qwen3.6 Plus is part of a broader Qwen ecosystem that includes both closed-source APIs and open-weight models.
OpenAI
GPT-5.4
GPT-5.4 is a proprietary multimodal large language model developed by OpenAI and released on March 5, 2026. It is designed for professional workloads such as advanced software development, research, and agentic automation. The model combines the general reasoning capabilities of the GPT-5 series with software engineering improvements derived from GPT-5.3-Codex. In the API and Codex environments it supports context windows of up to 1 million tokens, enabling long-context reasoning and large-scale code or document workflows.Compared with GPT-5.2, GPT-5.4 reduces false individual claims by 33% and lowers overall response errors by 18%, improving factual reliability across complex tasks. It is also the first general-purpose OpenAI release with native computer-use capabilities, allowing agents to interact with desktops, browsers, and external applications to complete multi-step workflows. The model family includes three variants: GPT-5.4 (standard), GPT-5.4 Pro for higher-performance workloads, and GPT-5.4 Thinking, a reasoning-oriented version in ChatGPT that presents an upfront plan before generating its response. The API also introduces a Tool Search system that allows models to retrieve tool definitions dynamically, reducing token usage in tool-heavy integrations.
Z.ai
GLM 5V Turbo
GLM-5V-Turbo is a native multimodal model from Z.ai that extends the GLM family with joint image, video, and text input aimed at vision-centered coding and agent workflows. The model reads screenshots, design drafts, document layouts, and interface captures and generates runnable code from them, covering tasks such as turning a visual design into a working front end, diagnosing rendering and layout defects from screen captures, and operating graphical user interfaces during long-horizon agent runs. It accepts roughly 200,000 input tokens and can emit up to 131,072 output tokens in a single response, which supports sessions that hold specifications, source files, logs, and visual references at the same time.Training includes a joint reinforcement learning stage spanning more than 30 tasks simultaneously, an approach Z.ai describes as a way to counter the trade-off in which improving visual recognition degrades programming ability and the reverse. Reported evaluations cover pure-text coding on the backend, frontend, and repository exploration tracks of CC-Bench-V2, together with agent execution suites such as PinchBench, ClawEval, and ZClawBench, indicating that text coding behavior is retained after visual input is added.
Grok
Grok 4.5
Grok 4.5 is a proprietary reasoning model from SpaceXAI (xAI) that accepts interleaved text and image input and returns text, with a 500,000 token context window. xAI positions it as a model for coding, agentic software work, and knowledge tasks, and states it was trained in the company's Memphis data centers on datasets spanning science, engineering, and mathematics. Its reinforcement learning stage covers hundreds of thousands of multi step software engineering tasks scored by automated checks and model based grading, and training is reported to have run on tens of thousands of NVIDIA GB300 GPUs using an asynchronous scheme in which multi hour agentic rollouts continue while learning proceeds in parallel, targeting long horizon autonomous operation rather than single turn inference.For vision, the model consumes JPEG and PNG images in any order relative to text prompts, covering visual question answering, description of chart and document imagery, and reading text rendered inside a scene. Reasoning effort is configurable, and the model supports function calling and structured outputs, so image inputs can be interleaved with tool calls inside agent loops. xAI has not published a technical report, architecture details, or parameter count, and reported mixture of experts sizing figures come from secondary coverage rather than official documentation.

Kimi K2.5 License

Modified MIT · Permissive license, with added terms

Kimi K2.5 ships under a modified MIT license: MIT's permissive core with vendor-specific clauses layered on top. The Kimi K2.5 license is permissive in outline, but the added clauses are where commercial restrictions hide.

Commercial use
Usually permitted without a separate commercial license, though the added clauses can restrict it — named-user caps, industry carve-outs, or scale thresholds are common. Confirm against the Kimi K2.5 license text.
Modification
Permitted under the MIT core, subject to whatever the modified terms add. Derivative works may inherit the same additions.
Redistribution
Permitted with attribution, and the modified terms must travel with every copy you distribute.

Uncertainty around licensing can delay or stop a project. Diff the Kimi K2.5 terms against stock MIT, and read any acceptable-use policy attached to them, before you build on the model.

Read the full Modified MIT license ↗

Do I need a commercial license for Kimi K2.5?

If the added clauses rule out your use case, you need a commercial license from the rights holder. Roboflow's licensing page lists the models whose commercial license is included in a Roboflow plan, and which deployment methods it covers — Kimi K2.5 is worth checking against that list before you commit.

Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.

Talk to sales

This model is released under a modified version of the MIT License. The base permissions of MIT apply, but additional terms or restrictions have been added by the model authors.

Commercial use is generally permitted, but the modified terms may add restrictions specific to this model. Review the full license text before deploying commercially.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About Kimi K2.5 Vision

Yes. Kimi K2.5 accepts image input, and on Roboflow's previous vision benchmark it passed 35.8% of visual understanding tasks (#74 of 77) and scored 19.7% on OCR. You can test it on your own image in the demo above.

Kimi K2.5 has not yet been evaluated on Roboflow's current Vision Evals. The results on this page are from the previous benchmark.

Yes. The demo on this page runs Kimi K2.5 in the free Roboflow Playground: upload an image and see results in seconds. A free account unlocks unlimited runs.