Roboflow

Not available to run yet

Grok 4.7 is on the Vision Evals leaderboard. Running it in the Playground is not available yet. View evals

Grok 4.7 Overview

Grok 4.7 is a proprietary model from SpaceXAI, released on September 21, 2026. It accepts text and images as input and returns text. It extends Grok 4.6 and is listed at the same API price.

Its Vision Evals scores are on the leaderboard. Running it in the Playground is not available yet, because the inference workflow is not ready.

Grok 4.7 Details & Performance

Details

Resources

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRVision LanguageVisual Question AnsweringObject Detection

Features

Usage

Past 30 Days

Not available

Not in Playground

Performance

Avg. Latency

Grok 4.7 Vision Evals

Vision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.

Evals updated September 22, 2026Pricing updated September 22, 2026

Overall score#22 of 54
71.9%
Avg cost / sample#48 of 54
$0.012
Avg speed / sample#43 of 54
23.55s
Avg tokens / sample
4.2K

Strengths and weaknesses

Grok 4.7 averages 71.9% across the six Vision Evals tasks, ranking #22 of 54 models overall.

Its weakest relative showing is Object Detection, ranking #39 of 54 at 40.4%.

At $0.012 per sample it is the 48th cheapest of the 54 benchmarked models, and its average inference time of 23.6s per sample makes it the 43rd fastest.

Performance profile

Field medianGrok 4.7

Field medians: Object Detection 53.4%, Counting 59.0%, Identification 84.4%, OCR 88.7%, Data Extraction 84.5%, Reasoning 54.4%.

Results by task

TaskScoreField (0 to 100)RankCost / sampleSpeed
Object Detection (low)
40.4%
±0.6, Mean of 3 runs, range 39.8 to 41.0
#39 of 54$0.01732.56s
Object Detection (high)
41.2%
±1.6, Mean of 3 runs, range 39.6 to 42.8
#17 of 20$0.02349.09s
Counting (low)
61.7%
±1.3, Mean of 3 runs, range 60.8 to 63.5
#26 of 54$0.008616.76s
Counting (high)
60.8%
±1.3, Mean of 3 runs, range 59.5 to 62.2
#18 of 20$0.01332.45s
Identification (low)
87.5%
±3.1, Mean of 3 runs, range 84.4 to 90.6
#20 of 54$0.00506.18s
Identification (high)
80.2%
±1.6, Mean of 3 runs, range 78.1 to 81.3
#20 of 20$0.007413.61s
OCR (low)
92.6%
±0.7, Mean of 3 runs, range 92.1 to 93.4
#9 of 54$0.01430.02s
OCR (high)
93.5%
±0.3, Mean of 3 runs, range 93.1 to 93.8
#2 of 20$0.03480.11s
Data Extraction (low)
84.9%
±2.6, Mean of 3 runs, range 82.5 to 87.6
#23 of 54$0.00484.90s
Data Extraction (high)
87.6%
±1.5, Mean of 3 runs, range 86.6 to 89.7
#10 of 20$0.00546.13s
Reasoning (low)
64.2%
±2.3, Mean of 3 runs, range 62.3 to 66.9
#17 of 54$0.01226.03s
Reasoning (high)
66.9%
±1.3, Mean of 3 runs, range 65.6 to 68.2
#20 of 40$0.01951.09s
  • Thinking longer helps: 0.7 points higher on object detection at high effort for 1.4x the cost and 1.5x the latency.
  • Thinking longer does not help: 0.9 points lower on counting at high effort for 1.6x the cost and 1.9x the latency.
  • Thinking longer does not help: 7.3 points lower on identification at high effort for 1.5x the cost and 2.2x the latency.
  • Thinking longer helps: 0.9 points higher on ocr at high effort for 2.3x the cost and 2.7x the latency.
  • Thinking longer helps: 2.8 points higher on data extraction at high effort for 1.1x the cost and 1.3x the latency.
  • Thinking longer helps: 2.7 points higher on reasoning at high effort for 1.7x the cost and 2x the latency.

Price vs. performance

Score vs. cost

Overall benchmark score against estimated cost per sample, on a log scale. Upper-left is the sweet spot: high quality at low cost.

53 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort · Grok 4.7 highlighted

Grok 4.7 scores are the mean of 3 runs per task at both low and high effort · Methodology

View all Vision Evals →

Grok 4.7 Pricing

Grok 4.7 costs $1.60 per 1M input tokens and $4.80 per 1M output tokens.

Input$1.60 / 1M tokens
Output$4.80 / 1M tokens
Cached input$0.400 / 1M tokens

Pricing updated Sep 22, 2026

Grok 4.7 License

Proprietary

Grok 4.7 is proprietary: the weights are not distributed, and the Grok 4.7 license is the vendor's commercial terms of service that you accept when you call the API.

Commercial use
Permitted under the vendor terms, typically metered per token or per request, with the vendor usage policy applying to your inputs and outputs.
Modification
Not available. Grok 4.7 weights are closed, so you can configure prompts and use vendor-hosted fine-tuning where it is offered, but you cannot modify the model itself.
Redistribution
Not permitted. You cannot self-host or resell the model; you build on the hosted API instead.

Vendor terms govern data retention, whether your inputs can be trained on, rate limits, and regional availability, and they can change with notice. Review them if you handle regulated or customer data.

Do I need a commercial license for Grok 4.7?

Proprietary terms are set by the vendor rather than negotiated per project, and no open-source obligation attaches to your code. If you would rather deploy a model whose commercial license is included in your plan — on Roboflow Managed Cloud or a Self-Hosted Inference Server — Roboflow's licensing page lists the supported alternatives to Grok 4.7.

Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.

Talk to sales

This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.

Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About Grok 4.7 Vision

Yes. Grok 4.7 accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is OCR at 92.6% (#9 of 54 at low effort).

Yes, and it is one of the model's strongest vision skills: its transcriptions match the ground truth 92.6% on average (#9 of 54 at low effort) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 84.9%.

Not its strength. On Vision Evals, Grok 4.7 scores 40.4% mAP@50 on object detection (#39 of 54 at low effort) and 61.7% judge-graded accuracy on object counting. For production counting or precise localization, pairing it with a specialized detector like RF-DETR or your own trained model in a Roboflow Workflow is usually more reliable: detect the objects, then count the detections.

On our benchmark's task mix, Grok 4.7 averages $0.01 per sample at $1.60 per 1M input and $4.80 per 1M output tokens (#48 of 54 on cost), with an average speed of 23.6s per sample across the benchmark. Actual cost depends on your images and prompts.

On the overall Vision Evals ranking, Grok 4.7 sits #22 of 54 at 71.9%, just behind Qwen3.6 35B-A3B (71.9%) and just ahead of Qwen3.5 27B (70.8%). See the full side-by-side: Grok 4.7 vs Qwen3.6 35B-A3B.