Roboflow
Meta

Meta: Muse Spark 1.2

Muse Spark 1.2 Overview

Muse Spark 1.2 is a proprietary multimodal reasoning model from Meta Superintelligence Labs, released as a coding-focused update to Muse Spark 1.1. It accepts text, images, video, audio, and PDF documents and returns text, with a context window of roughly one million tokens that allows whole repositories, long documents, and extended agent trajectories to be held in a single request. The model thinks before answering, and the amount of reasoning effort it spends is configurable per request. Alongside its visual and document understanding, it supports structured output and parallel function calling, and it is designed to operate either as a planning agent that delegates work or as a subagent executing tasks in parallel.

Training for version 1.2 scaled up compute on coding tasks and widened the diversity of training environments, concentrating on long-horizon work such as whole-repository generation, large end-to-end projects, and automated research. Part of the training data was self-generated, with Muse Spark 1.1 producing coding environments and instruction-following templates and grading candidate solutions against them. The model was co-trained with the Muse Code terminal agent, incorporating rejection-sampled harness trajectories and that toolset. Meta reports 82.9 percent on Terminal-Bench 2.1, an improvement of 6.7 points over Muse Spark 1.1. Multimodal use cases documented for the family include visual-to-code generation and detailed image and video captioning.

Muse Spark 1.2 Interactive Demo

Muse Spark 1.2 Details & Performance

Details

Resources

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRVision LanguageVisual Question Answering

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Performance

Avg. Latency

Arena Rankings

Muse Spark 1.2 Vision Evals

Vision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.

Evals updated August 6, 2026Pricing updated August 7, 2026

Overall score#5 of 25
80.4%
Avg cost / sample#16 of 25
$0.0071
Avg speed / sample#16 of 25
7.78s
Avg tokens / sample
2.9K

Strengths and weaknesses

Muse Spark 1.2 averages 80.4% across the six Vision Evals tasks, ranking #5 of 25 models overall.

It places in the top three for OCR and Reasoning.

Its weakest relative showing is Data Extraction, ranking #8 of 25 at 88.7%.

At $0.0071 per sample it is the 16th cheapest of the 25 benchmarked models, and its average inference time of 7.8s per sample makes it the 16th fastest.

Performance profile

Field medianMuse Spark 1.2

Field medians: Object Detection 56.0%, Counting 63.5%, Identification 84.4%, OCR 89.3%, Data Extraction 86.6%, Reasoning 55.6%.

Results by task

TaskScoreField (0 to 100)RankCost / sampleSpeed
Object Detection
60.1%
#7 of 25$0.00948.52s
Counting
74.3%
#5 of 25$0.00496.40s
Identification
90.6%
#7 of 25$0.00385.34s
OCR
93.8%
#2 of 25$0.00796.88s
Data Extraction
88.7%
#8 of 25$0.00334.18s
Reasoning (low)
74.8%
#3 of 25$0.007410.27s
Reasoning (high)
76.2%
#4 of 25$0.01216.31s
  • Thinking longer helps: 1.3 points higher on reasoning at high effort for 1.6x the cost and 1.6x the latency.

Price vs. performance

Score vs. cost

Overall benchmark score against estimated cost per sample. Upper-left is the sweet spot: high quality at low cost.

25 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort · Muse Spark 1.2 highlighted

Muse Spark 1.2 scores from a single evaluation run · Methodology

View all Vision Evals →

Muse Spark 1.2 Pricing

Muse Spark 1.2 costs $1.25 per 1M input tokens and $4.25 per 1M output tokens.

Input$1.25 / 1M tokens
Output$4.25 / 1M tokens
Cached input$0.150 / 1M tokens

Pricing updated Aug 7, 2026

Alternatives to Muse Spark 1.2

Other models worth comparing for similar use cases.

OpenAI
GPT-5.6 Sol
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 family, which also includes Terra (a balanced everyday-work tier) and Luna (a fast, cost-efficient tier). Sol is designed for demanding reasoning, long-horizon agentic workflows, software engineering, computer use, scientific research, and cybersecurity tasks. It introduces two new capability modes: a "max" reasoning effort setting that allocates additional compute time for difficult problems, and an "ultra" mode that coordinates multiple subagents in parallel to accelerate complex, multi-step work. The model supports native multimodal input, allowing it to process screenshots, diagrams, charts, documents, and photographs alongside text. A reported context window of approximately 1.5 million tokens enables processing of large codebases, lengthy research documents, and extended agentic sessions.GPT-5.6 Sol was announced on June 26, 2026, initially in a limited preview for trusted partners, and reached general availability on July 9, 2026. On the Agents' Last Exam benchmark, which evaluates long-running professional workflows across 55 fields, Sol scores 53.6. On Terminal-Bench 2.1, which tests command-line agentic coding workflows, Sol Ultra achieves 91.9%. The model also demonstrates gains in life sciences evaluations, including long-horizon genomics and quantitative biology analyses. OpenAI paired the release with its most extensive safety evaluation to date, combining human red teaming with large-scale automated testing, and classified Sol as High capability in both cybersecurity and biological risk under its Preparedness Framework, though it does not cross the Critical threshold in either category.
Anthropic
Claude Opus 5
Claude Opus 5 is a large language model with multimodal vision capabilities developed by Anthropic, released on July 24, 2026 as the fourth model in the Claude 5 family. It sits in the Opus tier of Anthropic's lineup, positioned below the Mythos-class Fable 5 and Mythos 5 models, and is framed by Anthropic as the go-to model for most knowledge work and automation tasks. The model approaches Fable 5's capabilities at roughly half the cost, priced at $5 per million input tokens and $25 per million output tokens. It becomes the default model on Claude Max and the strongest model available on Claude Pro. The model ships with a 1 million token context window and an adjustable "effort" parameter that allows users to trade reasoning depth for speed and token savings. Early enterprise customers reported that Opus 5 achieved comparable performance to Opus 4.8's maximum-reasoning mode while generating significantly fewer tokens on average, and demonstrated higher accuracy on financial modeling tasks with fewer tool calls and less time.Claude Opus 5 supports multimodal inputs including images and text, and is designed for agentic workflows, coding, scientific research, and complex enterprise tasks. Anthropic reports the model scores 10.2 percentage points higher than Opus 4.8 on an internal chemistry benchmark, making it the most capable generally available model for scientific research in the Claude lineup. Cyber classifiers on Opus 5 are designed to intervene approximately 85 percent less often than those on Fable 5, with fallback to Opus 4.8 when a classifier triggers. The model does not retain user data for 30 days, unlike Fable 5. It is available across Anthropic's platforms including Claude Code and Claude Cowork, as well as cloud partners.
Google
Gemini 3.1 Pro
Gemini 3.1 Pro is a proprietary multimodal model from Google’s Gemini 3 series, released in early 2026 and designed for advanced reasoning across large multimodal datasets. It accepts text, images, audio, video, and documents, supporting up to a 1-million-token input context with up to 64k output tokens. Compared with Gemini 3 Pro, it improves long-context synthesis and multi-step reasoning, enabling more reliable analysis of large documents, datasets, and software codebases.The model also advances visual understanding and grounding, allowing it to interpret UI screenshots, diagrams, and real-world scenes while referencing specific regions within images or video. These capabilities make Gemini 3.1 Pro well suited for multimodal workflows involving document processing, interface analysis, robotics research, and complex visual reasoning.
Qwen
Qwen3.8 Max
Qwen3.8 Max is the flagship tier of Alibaba's Qwen3.8 family, a sparse mixture-of-experts multimodal model with roughly 2.4 trillion total parameters of which about 95 billion activate per token, which keeps serving cost and latency well below what the total parameter count would imply. It builds on the architectural foundation established by Qwen3.5 and accepts text, images, video, and documents as input while producing text output. Reported context handling reaches close to one million tokens, with a maximum generation length of 131,072 tokens, so the model is aimed at long-horizon agentic work such as repository-scale coding, multi-step research, data analysis, and office document workflows.For vision work the model performs image and video understanding, document and chart interpretation, text recognition inside images, and grounded visual question answering, and Alibaba reports gains concentrated in multimodal and agentic evaluation categories rather than general reasoning. Published figures include 86.6 on Terminal-Bench 2.1, 67.7 on SWE-bench Pro, 93.0 on PaperBench, 82.8 on IFBench, and 92.6 on GPQA Diamond. It is the first model in the Max tier of the Qwen line for which the team states weights will be released publicly, alongside a smaller Qwen3.8 27B checkpoint. No training or safety model card has been published.
MoonshotAI
Kimi K3
Kimi K3 is a sparse Mixture-of-Experts large language model developed by Moonshot AI, with 2.8 trillion total parameters and a 1-million-token context window. The model activates 16 out of 896 experts per token using the Stable LatentMoE framework, and is built on two architectural innovations: Kimi Delta Attention (KDA), a hybrid linear attention mechanism that enables up to 6.3x faster decoding in long-context settings, and Attention Residuals (AttnRes), which selectively retrieves representations across model depth and delivers roughly 25% higher training efficiency. Together with refined training and data recipes, these structural advances yield approximately 2.5x better overall scaling efficiency compared to its predecessor Kimi K2. The model applies quantization-aware training from the supervised fine-tuning stage onward, using MXFP4 weights with MXFP8 activations for hardware compatibility. Thinking mode is always enabled at launch, with reasoning effort configurable via the reasoning_effort field.Kimi K3 supports native visual understanding alongside text, accepting image inputs for tasks that combine software engineering and visual reasoning. It targets long-horizon coding, knowledge work, and agentic workflows, and ships in two variants: K3 Max for general chat and agent tasks, and K3 Swarm Max for large-scale parallel processing across many coordinated sub-agents. The model is compatible with the OpenAI SDK via an OpenAI-compatible API. Full model weights are scheduled for release by July 27, 2026 under a Modified MIT license, following the open-weight pattern established by the Kimi K2 model family. A technical report with full architecture, training, and evaluation details is expected to accompany the weights release.
Meta
Muse Spark 1.1
Muse Spark 1.1 is a natively multimodal reasoning model from Meta Superintelligence Labs, released on July 9, 2026, as a significant upgrade to the original Muse Spark. The model accepts text, image, video, PDF, and audio as input and produces text output. It operates with a 1-million-token context window (1,048,576 tokens per the Meta Model API documentation) and is designed specifically for agentic tasks that require planning, tool use, computer use, and multi-agent orchestration. The model runs in a "Thinking" mode, where adjustable reasoning effort is applied before generating a response. It can function both as a main agent gathering context, forming plans, and delegating to parallel subagents and as a subagent that adheres to assigned tasks and escalates when needed. It is trained to decide autonomously when to write automation scripts versus interact directly with a user interface.Muse Spark 1.1 supports a range of multimodal capabilities including visual perception, image and video captioning, visual-to-code generation, and document analysis. The model was evaluated under Meta's Advanced AI Scaling Framework across frontier risk categories including chemical and biological threats, cybersecurity, and loss-of-control scenarios. Parameter count, architecture details, and training data composition are not publicly disclosed. The model is proprietary and closed-weight, accessible to consumers through the Meta AI app and to developers via the Meta Model API, which launched in public preview alongside this release.

Deploy Muse Spark 1.2 with an API

Muse Spark 1.2 runs as a hosted REST endpoint through Roboflow Workflows. Pick a task, then hand the prompt to your coding agent or copy the code. Forking the workflow into a free Roboflow workspace replaces the your-workspace and YOUR_API_KEY placeholders with your own.

Connect your agent to Roboflow (once)

Add the Roboflow MCP server

claude mcp add --transport http roboflow https://mcp.roboflow.com/mcp

Run /mcp and authorize Roboflow in your browser when the OAuth flow opens.

Start a new Claude Code session so the MCP loads, then paste the prompt below (it works the same in any agent).

Fork this workflow to your Roboflow workspace to use it.

Integrate the Roboflow "Muse Spark 1.2" workflow into my app.

- Endpoint: POST https://serverless.roboflow.com/<your-workspace>/workflows/playground-muse-spark-1-2-c
- Auth: send my Roboflow API key as `api_key` in the request body, read from the ROBOFLOW_API_KEY env var (never hardcode).
- Body: { "api_key": ..., "inputs": { `image`: { type: "url" | "base64", value }, `model_api_key`: my provider key } }.

With the Roboflow MCP connected, call `workflows_get` on "playground-muse-spark-1-2-c" to read the exact input schema and treat it as the source of truth. A live `workflows_run` for this workflow also needs my OpenRouter key (`model_api_key`) passed as a runtime parameter; if you don't have it yet, skip the test run — it will fail with a server error without the provider key, which is expected and not a problem with your code — and rely on the schema. Validate the real run via the REST call once the keys below are set. Without the MCP, use the contract above.

Before running the app, set up these keys so it does not error at runtime:
- `ROBOFLOW_API_KEY` (sent as `api_key`) from https://app.roboflow.com/settings/api
- `OPENROUTER_API_KEY` (sent as `model_api_key`) from https://openrouter.ai/keys — my OpenRouter key
Create a .gitignore'd .env with these variables, using placeholder values for any I haven't given you. Then pause and tell me directly, in your reply: the full path to the .env file, exactly which keys I need to paste in, and the link to get each one. Wait for me to confirm I've added them before you run anything. Do not run the app until I confirm.

Then add the integration to my codebase: match my project's language, framework, and conventions; read every key from environment variables (never hardcode); add basic error handling; and include a small runnable example. If you can't tell what language my project uses, ask me.
Installpip install inference-sdk

Fork this workflow to your Roboflow workspace to use it.

# 1. Import the library
from inference_sdk import InferenceHTTPClient

# 2. Connect to your workflow
client = InferenceHTTPClient(
  api_url="https://serverless.roboflow.com",
  api_key="YOUR_API_KEY"
)

# 3. Run your workflow on an image
result = client.run_workflow(
  workspace_name="your-workspace",
  workflow_id="playground-muse-spark-1-2-c",
  images={
    "image": "YOUR_IMAGE.jpg"  # Path to your image file
  },
  parameters={
    "model_api_key": "YOUR_OPENROUTER_API_KEY"
  },
  use_cache=True  # cache workflow definition for 15 minutes
)

# 4. Get your results
print(result)

Fork this workflow to your Roboflow workspace to use it.

const response = await fetch('https://serverless.roboflow.com/your-workspace/workflows/playground-muse-spark-1-2-c', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json'
  },
  body: JSON.stringify({
    api_key: 'YOUR_API_KEY',
    inputs: {
      "image": {"type": "url", "value": "IMAGE_URL"},
      "model_api_key": "YOUR_OPENROUTER_API_KEY"
    }
  })
});

const result = await response.json();
console.log(result);

Fork this workflow to your Roboflow workspace to use it.

curl --location 'https://serverless.roboflow.com/your-workspace/workflows/playground-muse-spark-1-2-c' \
--header 'Content-Type: application/json' \
--data '{
  "api_key": "YOUR_API_KEY",
  "inputs": {
    "image": {"type": "url", "value": "IMAGE_URL"},
    "model_api_key": "YOUR_OPENROUTER_API_KEY"
  }
}'

Muse Spark 1.2 License

Proprietary

License terms and commercial-use guidance for Muse Spark 1.2.

This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.

Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About Muse Spark 1.2 Vision

Yes. Muse Spark 1.2 accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is OCR at 93.8% (#2 of 25). You can test it on your own image in the demo above.

Yes, and it is one of the model's strongest vision skills: its transcriptions match the ground truth 93.8% on average (#2 of 25) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 88.7%.

It's serviceable. On Vision Evals, Muse Spark 1.2 scores 60.1% mAP@50 on object detection (#7 of 25) and 74.3% exact-match accuracy on object counting.

On our benchmark's task mix, Muse Spark 1.2 averages $0.0071 per sample at $1.25 per 1M input and $4.25 per 1M output tokens (#16 of 25 on cost), with an average speed of 7.8s per sample across the benchmark. Actual cost depends on your images and prompts.

On the overall Vision Evals ranking, Muse Spark 1.2 sits #5 of 25 at 80.4%, just behind Gemini 3.6 Flash (83.1%) and just ahead of Muse Spark 1.1 (79.2%). See the full side-by-side: Muse Spark 1.2 vs Gemini 3.6 Flash.