Roboflow

Kimi K3 Overview

Kimi K3 is a sparse Mixture-of-Experts large language model developed by Moonshot AI, with 2.8 trillion total parameters and a 1-million-token context window. The model activates 16 out of 896 experts per token using the Stable LatentMoE framework, and is built on two architectural innovations: Kimi Delta Attention (KDA), a hybrid linear attention mechanism that enables up to 6.3x faster decoding in long-context settings, and Attention Residuals (AttnRes), which selectively retrieves representations across model depth and delivers roughly 25% higher training efficiency. Together with refined training and data recipes, these structural advances yield approximately 2.5x better overall scaling efficiency compared to its predecessor Kimi K2. The model applies quantization-aware training from the supervised fine-tuning stage onward, using MXFP4 weights with MXFP8 activations for hardware compatibility. Thinking mode is always enabled at launch, with reasoning effort configurable via the reasoning_effort field.

Kimi K3 supports native visual understanding alongside text, accepting image inputs for tasks that combine software engineering and visual reasoning. It targets long-horizon coding, knowledge work, and agentic workflows, and ships in two variants: K3 Max for general chat and agent tasks, and K3 Swarm Max for large-scale parallel processing across many coordinated sub-agents. The model is compatible with the OpenAI SDK via an OpenAI-compatible API. Full model weights are scheduled for release by July 27, 2026 under a Modified MIT license, following the open-weight pattern established by the Kimi K2 model family. A technical report with full architecture, training, and evaluation details is expected to accompany the weights release.

Kimi K3 Interactive Demo

Kimi K3 Details & Performance

Details

Resources

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRObject DetectionVision LanguageVisual Question Answering

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Performance

Avg. Latency

Arena Rankings

Kimi K3 Vision Evals

Vision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.

Evals updated August 14, 2026Pricing updated August 18, 2026

Overall score#20 of 30
66.5%
Avg cost / sample#25 of 30
$0.011
Avg speed / sample#27 of 30
12.71s
Avg tokens / sample
2.6K

Strengths and weaknesses

Kimi K3 averages 66.5% across the six Vision Evals tasks, ranking #20 of 30 models overall.

Its weakest relative showing is Counting, ranking #28 of 30 at 46.0%.

At $0.011 per sample it is the 25th cheapest of the 30 benchmarked models, and its average inference time of 12.7s per sample makes it the 27th fastest.

Performance profile

Field medianKimi K3

Field medians: Object Detection 55.2%, Counting 64.2%, Identification 84.4%, OCR 90.0%, Data Extraction 86.6%, Reasoning 58.0%.

Results by task

TaskScoreField (0 to 100)RankCost / sampleSpeed
Object Detection
51.9%
#19 of 30$0.02023.23s
Counting
46.0%
#28 of 30$0.00466.12s
Identification
81.3%
#17 of 30$0.00415.21s
OCR
93.0%
#5 of 30$0.009413.18s
Data Extraction
84.5%
#17 of 30$0.00465.65s
Reasoning (low)
42.4%
#22 of 30$0.00444.51s
Reasoning (high)
74.2%
#8 of 30$0.03779.92s
  • Thinking longer helps: 31.8 points higher on reasoning at high effort for 8.4x the cost and 17.7x the latency.

Price vs. performance

Score vs. cost

Overall benchmark score against estimated cost per sample. Upper-left is the sweet spot: high quality at low cost.

30 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort · Kimi K3 highlighted

Kimi K3 scores from a single evaluation run · Methodology

View all Vision Evals →

Kimi K3 Pricing

Kimi K3 costs $3.00 per 1M input tokens and $15.00 per 1M output tokens.

Input$3.00 / 1M tokens
Output$15.00 / 1M tokens
Cached input$0.300 / 1M tokens

Pricing updated Aug 18, 2026

Alternatives to Kimi K3

Other models worth comparing for similar use cases.

Qwen
Qwen3.8 Max
Qwen3.8 Max is the flagship tier of Alibaba's Qwen3.8 family, a sparse mixture-of-experts multimodal model with roughly 2.4 trillion total parameters of which about 95 billion activate per token, which keeps serving cost and latency well below what the total parameter count would imply. It builds on the architectural foundation established by Qwen3.5 and accepts text, images, video, and documents as input while producing text output. Reported context handling reaches close to one million tokens, with a maximum generation length of 131,072 tokens, so the model is aimed at long-horizon agentic work such as repository-scale coding, multi-step research, data analysis, and office document workflows.For vision work the model performs image and video understanding, document and chart interpretation, text recognition inside images, and grounded visual question answering, and Alibaba reports gains concentrated in multimodal and agentic evaluation categories rather than general reasoning. Published figures include 86.6 on Terminal-Bench 2.1, 67.7 on SWE-bench Pro, 93.0 on PaperBench, 82.8 on IFBench, and 92.6 on GPQA Diamond. It is the first model in the Max tier of the Qwen line for which the team states weights will be released publicly, alongside a smaller Qwen3.8 27B checkpoint. No training or safety model card has been published.
Qwen
Qwen3.5 397B A17B
Qwen3.5-397B-A17B is a 397B-parameter (17B active) open-weight multimodal model developed by Alibaba’s Qwen team, released on 2026-02-16 under Apache-2.0. It supports text and image inputs with text outputs, combining a sparse Mixture-of-Experts architecture with Gated Delta Networks for efficient scaling. The model provides native vision-language reasoning and a large ~262K token context window, extendable to ~1M tokens.As the first open-weight release in the Qwen3.5 family, it positions itself as a high-capacity, long-context alternative in the large vision-language space, balancing scale and efficiency via sparse activation. It is designed for advanced reasoning, coding, agent workflows, and multimodal understanding tasks.
Anthropic
Claude Opus 5
Claude Opus 5 is a large language model with multimodal vision capabilities developed by Anthropic, released on July 24, 2026 as the fourth model in the Claude 5 family. It sits in the Opus tier of Anthropic's lineup, positioned below the Mythos-class Fable 5 and Mythos 5 models, and is framed by Anthropic as the go-to model for most knowledge work and automation tasks. The model approaches Fable 5's capabilities at roughly half the cost, priced at $5 per million input tokens and $25 per million output tokens. It becomes the default model on Claude Max and the strongest model available on Claude Pro. The model ships with a 1 million token context window and an adjustable "effort" parameter that allows users to trade reasoning depth for speed and token savings. Early enterprise customers reported that Opus 5 achieved comparable performance to Opus 4.8's maximum-reasoning mode while generating significantly fewer tokens on average, and demonstrated higher accuracy on financial modeling tasks with fewer tool calls and less time.Claude Opus 5 supports multimodal inputs including images and text, and is designed for agentic workflows, coding, scientific research, and complex enterprise tasks. Anthropic reports the model scores 10.2 percentage points higher than Opus 4.8 on an internal chemistry benchmark, making it the most capable generally available model for scientific research in the Claude lineup. Cyber classifiers on Opus 5 are designed to intervene approximately 85 percent less often than those on Fable 5, with fallback to Opus 4.8 when a classifier triggers. The model does not retain user data for 30 days, unlike Fable 5. It is available across Anthropic's platforms including Claude Code and Claude Cowork, as well as cloud partners.
OpenAI
GPT-5.6 Sol
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 family, which also includes Terra (a balanced everyday-work tier) and Luna (a fast, cost-efficient tier). Sol is designed for demanding reasoning, long-horizon agentic workflows, software engineering, computer use, scientific research, and cybersecurity tasks. It introduces two new capability modes: a "max" reasoning effort setting that allocates additional compute time for difficult problems, and an "ultra" mode that coordinates multiple subagents in parallel to accelerate complex, multi-step work. The model supports native multimodal input, allowing it to process screenshots, diagrams, charts, documents, and photographs alongside text. A reported context window of approximately 1.5 million tokens enables processing of large codebases, lengthy research documents, and extended agentic sessions.GPT-5.6 Sol was announced on June 26, 2026, initially in a limited preview for trusted partners, and reached general availability on July 9, 2026. On the Agents' Last Exam benchmark, which evaluates long-running professional workflows across 55 fields, Sol scores 53.6. On Terminal-Bench 2.1, which tests command-line agentic coding workflows, Sol Ultra achieves 91.9%. The model also demonstrates gains in life sciences evaluations, including long-horizon genomics and quantitative biology analyses. OpenAI paired the release with its most extensive safety evaluation to date, combining human red teaming with large-scale automated testing, and classified Sol as High capability in both cybersecurity and biological risk under its Preparedness Framework, though it does not cross the Critical threshold in either category.
Google
Gemini 3.1 Pro
Gemini 3.1 Pro is a proprietary multimodal model from Google’s Gemini 3 series, released in early 2026 and designed for advanced reasoning across large multimodal datasets. It accepts text, images, audio, video, and documents, supporting up to a 1-million-token input context with up to 64k output tokens. Compared with Gemini 3 Pro, it improves long-context synthesis and multi-step reasoning, enabling more reliable analysis of large documents, datasets, and software codebases.The model also advances visual understanding and grounding, allowing it to interpret UI screenshots, diagrams, and real-world scenes while referencing specific regions within images or video. These capabilities make Gemini 3.1 Pro well suited for multimodal workflows involving document processing, interface analysis, robotics research, and complex visual reasoning.
MoonshotAI
Kimi K2.5
Kimi K2.5 is a frontier-scale multimodal AI model developed by Moonshot AI and released on January 27, 2026. As a significant advancement within the Kimi K2 family, it utilizes a sparse Mixture-of-Experts (MoE) architecture with 1 trillion total parameters (32 billion active per inference) and a massive 256K-token context window. The model features native multimodal integration via a 400M-parameter MoonViT encoder, allowing it to process text, images, and video frames simultaneously. Built for both speed and depth, it offers "Instant" and "Thinking" modes, the latter of which excels at expert-level reasoning, scoring 50.2% on the Humanity’s Last Exam (HLE) benchmark when equipped with tools.The model is released under a Modified MIT License, which remains open-weight but requires attribution for high-revenue commercial entities. It introduces an "Agent Swarm" paradigm capable of coordinating up to 100 specialized sub-agents for parallel workflows, significantly reducing latency in complex research tasks. For vision tasks, Kimi K2.5 demonstrates strong autonomous visual debugging capabilities, where it can inspect its own generated UI outputs against visual specifications to iteratively refine frontend code. This makes it a powerful choice for developers testing automated UI reconstruction, high-fidelity OCR document processing, and multi-step agentic research grounded in complex visual data.

Deploy Kimi K3 with an API

Kimi K3 runs as a hosted REST endpoint through Roboflow Workflows. Pick a task, then hand the prompt to your coding agent or copy the code. Deploying the workflow into a free Roboflow workspace replaces the your-workspace and YOUR_API_KEY placeholders with your own.

Connect your agent to Roboflow (once)

Add the Roboflow MCP server

claude mcp add --transport http roboflow https://mcp.roboflow.com/mcp

Run /mcp and authorize Roboflow in your browser when the OAuth flow opens.

Start a new Claude Code session so the MCP loads, then paste the prompt below (it works the same in any agent).

Deploy this workflow to your Roboflow workspace to use it.

Integrate the Roboflow "Kimi K3" workflow into my app.

- Endpoint: POST https://serverless.roboflow.com/<your-workspace>/workflows/kimi-k3-object-detection
- Auth: send my Roboflow API key as `api_key` in the request body, read from the ROBOFLOW_API_KEY env var (never hardcode).
- Body: { "api_key": ..., "inputs": { `image`: { type: "url" | "base64", value }, `classes`: string array } }.
- Billing: this workflow needs no provider API key — inference runs on my Roboflow credits. A BYO provider key can be added to the model step in the Roboflow workflow editor later.

With the Roboflow MCP connected, call `workflows_get` on "kimi-k3-object-detection" to read the exact input schema (the source of truth), then `workflows_run` on a sample image to confirm the output shape before writing code (the MCP is authenticated, so this needs no key). Without the MCP, use the contract above.

Before running the app, set up these keys so it does not error at runtime:
- `ROBOFLOW_API_KEY` (sent as `api_key`) from https://app.roboflow.com/settings/api
Create a .gitignore'd .env with these variables, using placeholder values for any I haven't given you. Then pause and tell me directly, in your reply: the full path to the .env file, exactly which keys I need to paste in, and the link to get each one. Wait for me to confirm I've added them before you run anything. Do not run the app until I confirm.

Then add the integration to my codebase: match my project's language, framework, and conventions; read every key from environment variables (never hardcode); add basic error handling; and include a small runnable example. If you can't tell what language my project uses, ask me.
Installpip install inference-sdk

Deploy this workflow to your Roboflow workspace to use it.

# Inference runs on your Roboflow credits — no provider API key needed. To bill your own provider account instead, add an api_key to the model step in the Roboflow workflow editor.
# 1. Import the library
from inference_sdk import InferenceHTTPClient

# 2. Connect to your workflow
client = InferenceHTTPClient(
  api_url="https://serverless.roboflow.com",
  api_key="YOUR_API_KEY"
)

# 3. Run your workflow on an image
result = client.run_workflow(
  workspace_name="your-workspace",
  workflow_id="kimi-k3-object-detection",
  images={
    "image": "YOUR_IMAGE.jpg"  # Path to your image file
  },
  parameters={
    "classes": ["class1", "class2", "class3"]
  },
  use_cache=True  # cache workflow definition for 15 minutes
)

# 4. Get your results
print(result)

Deploy this workflow to your Roboflow workspace to use it.

// Inference runs on your Roboflow credits — no provider API key needed. To bill your own provider account instead, add an api_key to the model step in the Roboflow workflow editor.
const response = await fetch('https://serverless.roboflow.com/your-workspace/workflows/kimi-k3-object-detection', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json'
  },
  body: JSON.stringify({
    api_key: 'YOUR_API_KEY',
    inputs: {
      "image": {"type": "url", "value": "IMAGE_URL"},
      "classes": ["class1", "class2", "class3"]
    }
  })
});

const result = await response.json();
console.log(result);

Deploy this workflow to your Roboflow workspace to use it.

# Inference runs on your Roboflow credits — no provider API key needed. To bill your own provider account instead, add an api_key to the model step in the Roboflow workflow editor.
curl --location 'https://serverless.roboflow.com/your-workspace/workflows/kimi-k3-object-detection' \
--header 'Content-Type: application/json' \
--data '{
  "api_key": "YOUR_API_KEY",
  "inputs": {
    "image": {"type": "url", "value": "IMAGE_URL"},
    "classes": ["class1", "class2", "class3"]
  }
}'

Kimi K3 License

Modified MIT · Permissive license, with added terms

Kimi K3 ships under a modified MIT license: MIT's permissive core with vendor-specific clauses layered on top. The Kimi K3 license is permissive in outline, but the added clauses are where commercial restrictions hide.

Commercial use
Usually permitted without a separate commercial license, though the added clauses can restrict it — named-user caps, industry carve-outs, or scale thresholds are common. Confirm against the Kimi K3 license text.
Modification
Permitted under the MIT core, subject to whatever the modified terms add. Derivative works may inherit the same additions.
Redistribution
Permitted with attribution, and the modified terms must travel with every copy you distribute.

Uncertainty around licensing can delay or stop a project. Diff the Kimi K3 terms against stock MIT, and read any acceptable-use policy attached to them, before you build on the model.

Read the full Modified MIT license ↗

Do I need a commercial license for Kimi K3?

If the added clauses rule out your use case, you need a commercial license from the rights holder. Roboflow's licensing page lists the models whose commercial license is included in a Roboflow plan, and which deployment methods it covers — Kimi K3 is worth checking against that list before you commit.

Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.

Talk to sales

This model is released under a modified version of the MIT License. The base permissions of MIT apply, but additional terms or restrictions have been added by the model authors.

Commercial use is generally permitted, but the modified terms may add restrictions specific to this model. Review the full license text before deploying commercially.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About Kimi K3 Vision

Yes. Kimi K3 accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is OCR at 93% (#5 of 30). You can test it on your own image in the demo above.

Yes, and it is one of the model's strongest vision skills: its transcriptions match the ground truth 93% on average (#5 of 30) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 84.5%.

Not its strength. On Vision Evals, Kimi K3 scores 51.9% mAP@50 on object detection (#19 of 30) and 46% exact-match accuracy on object counting. For production counting or precise localization, pairing it with a specialized detector like RF-DETR or your own trained model in a Roboflow Workflow is usually more reliable: detect the objects, then count the detections.

On our benchmark's task mix, Kimi K3 averages $0.01 per sample at $3.00 per 1M input and $15.00 per 1M output tokens (#25 of 30 on cost), with an average speed of 12.7s per sample across the benchmark. Actual cost depends on your images and prompts.

On the overall Vision Evals ranking, Kimi K3 sits #20 of 30 at 66.5%, just behind Claude Opus 4.8 (66.8%) and just ahead of Claude Sonnet 5 (66.4%). See the full side-by-side: Kimi K3 vs Claude Opus 4.8.