Roboflow

Kimi K2.5 Overview

Kimi K2.5 is a frontier-scale multimodal AI model developed by Moonshot AI and released on January 27, 2026. As a significant advancement within the Kimi K2 family, it utilizes a sparse Mixture-of-Experts (MoE) architecture with 1 trillion total parameters (32 billion active per inference) and a massive 256K-token context window. The model features native multimodal integration via a 400M-parameter MoonViT encoder, allowing it to process text, images, and video frames simultaneously. Built for both speed and depth, it offers "Instant" and "Thinking" modes, the latter of which excels at expert-level reasoning, scoring 50.2% on the Humanity’s Last Exam (HLE) benchmark when equipped with tools.

The model is released under a Modified MIT License, which remains open-weight but requires attribution for high-revenue commercial entities. It introduces an "Agent Swarm" paradigm capable of coordinating up to 100 specialized sub-agents for parallel workflows, significantly reducing latency in complex research tasks. For vision tasks, Kimi K2.5 demonstrates strong autonomous visual debugging capabilities, where it can inspect its own generated UI outputs against visual specifications to iteratively refine frontend code. This makes it a powerful choice for developers testing automated UI reconstruction, high-fidelity OCR document processing, and multi-step agentic research grounded in complex visual data.

Kimi K2.5 Interactive Demo

Kimi K2.5 Details & Performance

Details

Resources

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRVision LanguageVisual Question Answering

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Performance

Avg. Latency

Arena Rankings

Kimi K2.5 Vision Evals

Kimi K2.5 has not yet been evaluated on the current benchmark. The results below are from the legacy version of Vision Evals, our previous benchmark. See the current Vision Evals

Visual Understanding

77 models · 67 tasks
HighestLowest
This model#74 of 7735.82% pass rate · better than 1%
Score35.82%pass rate across 67 tasks
Speed14.81savg response per task
Cost$0.0024 / task$0.450 in · $2.25 out / 1M
Tokens2.7K / task1.6K in · 766 out
Score key:≥75%40–74%<40%
CategoryPassedScore
Document Understanding5 / 9
55.6%
Defect Detection7 / 15
46.7%
Object Understanding6 / 14
42.9%
Spatial Understanding5 / 19
26.3%
Object Counting1 / 10
10%
HighestLowest
This model#58 of 5819.65% pass rate · better than 0%
Score19.65%pass rate across 229 tasks
Speed13.09savg response per task
Cost$0.0006 / task$0.450 in · $2.25 out / 1M
Tokens706 / task119 in · 258 out
Score key:≥75%40–74%<40%
CategoryPassedScore
Handwritten Math5 / 10
50%
VQA & Extraction20 / 60
33.3%
Text Recognition8 / 30
26.7%
Focused Scene OCR10 / 99
10.1%
License Plate Recognition2 / 30
6.7%

Scores based on a single evaluation run · Methodology

View all legacy Vision Evals results →

Kimi K2.5 Pricing

Kimi K2.5 costs $0.450 per 1M input tokens and $2.25 per 1M output tokens.

Input$0.450 / 1M tokens
Output$2.25 / 1M tokens
Cached input$0.070 / 1M tokens

Pricing updated Aug 19, 2026

Price vs. performance

Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.

6 of 6 models plotted

ModelScoreMedian tokensEst. cost / taskCompare
AnthropicClaude Haiku 4.558.2%2.3K$0.0030Compare
OpenAIGPT-5 Nano58.2%2.7K$0.0003Compare
QwenQwen3.5 397B A17B58.2%1.5K$0.0006Compare
GoogleGemini 2.5 Flash55.2%476$0.0005Compare
GoogleGemini 2.5 Flash-Lite53.7%301<$0.0001Compare
MoonshotAIKimi K2.5(this model)35.8%2.7K$0.0024

Alternatives to Kimi K2.5

Other models worth comparing for similar use cases.

MoonshotAI
Kimi K3
Kimi K3 is a sparse Mixture-of-Experts large language model developed by Moonshot AI, with 2.8 trillion total parameters and a 1-million-token context window. The model activates 16 out of 896 experts per token using the Stable LatentMoE framework, and is built on two architectural innovations: Kimi Delta Attention (KDA), a hybrid linear attention mechanism that enables up to 6.3x faster decoding in long-context settings, and Attention Residuals (AttnRes), which selectively retrieves representations across model depth and delivers roughly 25% higher training efficiency. Together with refined training and data recipes, these structural advances yield approximately 2.5x better overall scaling efficiency compared to its predecessor Kimi K2. The model applies quantization-aware training from the supervised fine-tuning stage onward, using MXFP4 weights with MXFP8 activations for hardware compatibility. Thinking mode is always enabled at launch, with reasoning effort configurable via the reasoning_effort field.Kimi K3 supports native visual understanding alongside text, accepting image inputs for tasks that combine software engineering and visual reasoning. It targets long-horizon coding, knowledge work, and agentic workflows, and ships in two variants: K3 Max for general chat and agent tasks, and K3 Swarm Max for large-scale parallel processing across many coordinated sub-agents. The model is compatible with the OpenAI SDK via an OpenAI-compatible API. Full model weights are scheduled for release by July 27, 2026 under a Modified MIT license, following the open-weight pattern established by the Kimi K2 model family. A technical report with full architecture, training, and evaluation details is expected to accompany the weights release.
Qwen
Qwen3.5 397B A17B
Qwen3.5-397B-A17B is a 397B-parameter (17B active) open-weight multimodal model developed by Alibaba’s Qwen team, released on 2026-02-16 under Apache-2.0. It supports text and image inputs with text outputs, combining a sparse Mixture-of-Experts architecture with Gated Delta Networks for efficient scaling. The model provides native vision-language reasoning and a large ~262K token context window, extendable to ~1M tokens.As the first open-weight release in the Qwen3.5 family, it positions itself as a high-capacity, long-context alternative in the large vision-language space, balancing scale and efficiency via sparse activation. It is designed for advanced reasoning, coding, agent workflows, and multimodal understanding tasks.
Qwen
Qwen3 VL 235B A22B Instruct
Qwen3 VL 235B A22B Instruct is a flagship multimodal vision-language model developed by Qwen (Alibaba Cloud), designed for instruction-following tasks that combine advanced text generation with visual understanding. It serves as a high-end open-weight model for developers and researchers building multimodal AI systems that require strong reasoning, perception, and long-context capabilities.The model supports interleaved text and image inputs, very long context windows (up to roughly 256K tokens), and efficient inference through a mixture-of-experts architecture with about 22B active parameters out of 235B total. In today’s landscape, it competes with top-tier proprietary vision-language models while offering the advantages of open weights and flexible deployment. Typical applications include multimodal assistants, document and image analysis, visual reasoning, and large-context instruction-based workflows.
Google
Gemini 3.1 Pro
Gemini 3.1 Pro is a proprietary multimodal model from Google’s Gemini 3 series, released in early 2026 and designed for advanced reasoning across large multimodal datasets. It accepts text, images, audio, video, and documents, supporting up to a 1-million-token input context with up to 64k output tokens. Compared with Gemini 3 Pro, it improves long-context synthesis and multi-step reasoning, enabling more reliable analysis of large documents, datasets, and software codebases.The model also advances visual understanding and grounding, allowing it to interpret UI screenshots, diagrams, and real-world scenes while referencing specific regions within images or video. These capabilities make Gemini 3.1 Pro well suited for multimodal workflows involving document processing, interface analysis, robotics research, and complex visual reasoning.
Anthropic
Claude Opus 4.6
Claude Opus 4.6 is the flagship large language model from Anthropic, released on 2026-02-05 for advanced reasoning, complex coding, and enterprise agent workflows. It supports text and image inputs via API, offers a 200K-token standard context window with a 1M-token beta option, and enables outputs up to 128K tokens, with adaptive reasoning and context compaction for sustained tasks.As of 2026-02-17, Anthropic also released Claude Sonnet 4.6, extending the 1M-token context window to a broader tier. Opus remains positioned for maximum depth and benchmark performance, while Sonnet 4.6 brings long-context capability to more cost- and latency-sensitive production use cases.
Meta
Llama 4 Maverick
Llama 4 Maverick, introduced on April 5, 2025, is one of the first models in Meta’s Llama 4 family, designed as a natively multimodal model supporting text + image inputs with text outputs. It employs a Mixture-of-Experts (MoE) architecture with 128 experts, activating ~17B parameters per token out of a pool of ~400B total parameters. This design improves scalability, efficiency, and reasoning capacity. Maverick has a 1M-token context window, enabling it to handle large documents, extended conversations, and multimodal reasoning. Its knowledge cutoff is August 2024.The model is released under the Llama 4 Community License and comes in both base and instruction-tuned (“Instruct”) versions. Maverick is widely deployed via Hugging Face, Google Vertex AI, Amazon Bedrock, and Oracle Cloud, making it one of the most accessible large open-weight models. However, it outputs text only (no image/audio generation) and, while input capacity is huge, output limits are typically much smaller. The MoE design also raises hardware demands, as maintaining 128 experts requires significant compute resources, and Meta’s license introduces restrictions around commercial-scale use.

Kimi K2.5 License

Modified MIT · Permissive license, with added terms

Kimi K2.5 ships under a modified MIT license: MIT's permissive core with vendor-specific clauses layered on top. The Kimi K2.5 license is permissive in outline, but the added clauses are where commercial restrictions hide.

Commercial use
Usually permitted without a separate commercial license, though the added clauses can restrict it — named-user caps, industry carve-outs, or scale thresholds are common. Confirm against the Kimi K2.5 license text.
Modification
Permitted under the MIT core, subject to whatever the modified terms add. Derivative works may inherit the same additions.
Redistribution
Permitted with attribution, and the modified terms must travel with every copy you distribute.

Uncertainty around licensing can delay or stop a project. Diff the Kimi K2.5 terms against stock MIT, and read any acceptable-use policy attached to them, before you build on the model.

Read the full Modified MIT license ↗

Do I need a commercial license for Kimi K2.5?

If the added clauses rule out your use case, you need a commercial license from the rights holder. Roboflow's licensing page lists the models whose commercial license is included in a Roboflow plan, and which deployment methods it covers — Kimi K2.5 is worth checking against that list before you commit.

Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.

Talk to sales

This model is released under a modified version of the MIT License. The base permissions of MIT apply, but additional terms or restrictions have been added by the model authors.

Commercial use is generally permitted, but the modified terms may add restrictions specific to this model. Review the full license text before deploying commercially.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About Kimi K2.5 Vision

Yes. Kimi K2.5 accepts image input, and on Roboflow's previous vision benchmark it passed 35.8% of visual understanding tasks (#74 of 77) and scored 19.7% on OCR. You can test it on your own image in the demo above.

Kimi K2.5 has not yet been evaluated on Roboflow's current Vision Evals. The results on this page are from the previous benchmark.

Yes. The demo on this page runs Kimi K2.5 in the free Roboflow Playground: upload an image and see results in seconds. A free account unlocks unlimited runs.