Roboflow

SpaceXAI Vision Models

Compare all 3 SpaceXAI vision models we track, all of them served through an API. Try them on your own images, free in the Roboflow Playground.

3 models · 3 free to try · Updated Aug 2026

About SpaceXAI Vision Models

Every SpaceXAI vision model in our catalog is proprietary. None of the 3 have downloadable weights: you reach them through SpaceXAI's API and pay per request. The newest additions are Grok 4.6 (Aug 2026) and Grok 4.5 (Jul 2026).

Between them, the SpaceXAI models we list cover 10 distinct vision tasks. The widest coverage is captioning (3 models), chart question answering (3), and classification (3). For captioning, start with Grok 4.6, Grok 4.5, or Grok 4.

On the API tier, SpaceXAI's most recent entry is Grok 4.6 (Aug 2026), billed per token by the provider rather than run on your own hardware. With every model here served through an API, the trade is price and latency against capability, and at high volume a smaller model fine-tuned on your own data usually undercuts all of them. All 3 SpaceXAI vision models run live in the Roboflow Playground, so you can run the same image through several of them and compare the answers before committing to one.

Which SpaceXAI Model Should You Use?

What each of the 3 SpaceXAI vision models in our catalog is built for, and how you run it.

Grok 4.6
Best for: OCR and Document Question Answering. Proprietary: no downloadable weights, billed per token through SpaceXAI's API. Runnable in the Playground.
Grok 4.5
Best for: OCR and Document Question Answering. Proprietary: no downloadable weights, billed per token through SpaceXAI's API. Runnable in the Playground.
Grok 4
Best for: Object Detection and OCR. Proprietary: no downloadable weights, billed per token through SpaceXAI's API. Runnable in the Playground.

SpaceXAI Models via API

3 proprietary models where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.

Grok
Grok 4.6
NEW
Grok 4.6 is a proprietary reasoning model from xAI aimed at long-running agentic workflows, coding, and knowledge work. It accepts text and image input and returns text, with a 500,000 token context window and a knowledge cutoff of February 1, 2026. The model exposes an adjustable reasoning budget with low, medium, high, and xhigh settings, where high is the default, and it supports function calling, structured outputs, web and X search, and code execution as documented tool behaviors. Its visual capability covers interpreting images supplied alongside text prompts, which places it in the visual question answering and document understanding family rather than producing pixel level outputs such as boxes or masks.xAI characterizes Grok 4.6 as the result of an extended post-training run over the Grok 4.5 lineage rather than a new pretrained base. The described recipe combines curated model-generated reasoning and technical data, engineering data, a revised optimizer, regenerated supervised fine-tuning trajectories, and reinforcement learning across agent environments spanning knowledge work, coding, kernel optimization, web development, and computer-aided design. Parameter count and architecture specifics are not disclosed. Independent measurement from Artificial Analysis places the model at 61 on its Intelligence Index, five points above Grok 4.5.
Grok
Grok 4.5
Grok 4.5 is a proprietary reasoning model from SpaceXAI (xAI) that accepts interleaved text and image input and returns text, with a 500,000 token context window. xAI positions it as a model for coding, agentic software work, and knowledge tasks, and states it was trained in the company's Memphis data centers on datasets spanning science, engineering, and mathematics. Its reinforcement learning stage covers hundreds of thousands of multi step software engineering tasks scored by automated checks and model based grading, and training is reported to have run on tens of thousands of NVIDIA GB300 GPUs using an asynchronous scheme in which multi hour agentic rollouts continue while learning proceeds in parallel, targeting long horizon autonomous operation rather than single turn inference.For vision, the model consumes JPEG and PNG images in any order relative to text prompts, covering visual question answering, description of chart and document imagery, and reading text rendered inside a scene. Reasoning effort is configurable, and the model supports function calling and structured outputs, so image inputs can be interleaved with tool calls inside agent loops. xAI has not published a technical report, architecture details, or parameter count, and reported mixture of experts sizing figures come from secondary coverage rather than official documentation.
Grok
Grok 4
Grok 4, released by xAI on July 9, 2025, is the fourth-generation model in the Grok family and the most advanced to date. It is multimodal, supporting text, vision, tool use, and real-time web search, with a reported 256,000-token context window for long-form reasoning and document analysis. Its training data extends through November 2024, making it the most up-to-date Grok model at launch.The lineup includes Grok 4 Generalist for broad tasks, Grok 4 Heavy for higher-capacity reasoning, and Grok 4 Code optimized for programming and debugging. A notable feature is its always-on “Think” mode, designed for deeper multi-step reasoning. While xAI has not disclosed parameter counts, Grok 4 is positioned to compete with frontier models like GPT-5 and Claude 4, balancing real-time knowledge via web integration with structured tool use. It is best suited for coding, complex reasoning, and multimodal AI assistants.

Frequently Asked Questions About SpaceXAI Vision Models

Which SpaceXAI models can do captioning?

3 of the 3 SpaceXAI vision models we track handle captioning: Grok 4.6, Grok 4.5, and Grok 4. Each model page lists its full task coverage, license, and specs.

Which SpaceXAI models can do chart question answering?

3 of the 3 SpaceXAI vision models we track handle chart question answering: Grok 4.6, Grok 4.5, and Grok 4. Each model page lists its full task coverage, license, and specs.

Are SpaceXAI vision models open source?

No. All 3 SpaceXAI vision models in our catalog are proprietary: the weights are not downloadable, and you access them through SpaceXAI's API, billed by SpaceXAI.

What is the best SpaceXAI model for captioning?

We do not publish a SpaceXAI-only ranking, so pick on constraints rather than a label. 3 SpaceXAI models handle captioning; the most recent is Grok 4.6 (Aug 2026). For fixed categories in production, a model fine-tuned on your own data typically beats any general-purpose model.

How many SpaceXAI vision models are on Roboflow Playground?

We track 3 live SpaceXAI vision models. The most recent addition is Grok 4.6, released Aug 2026.

Can I try SpaceXAI vision models for free?

Yes. All 3 SpaceXAI models run live in the Roboflow Playground. Upload your own image, run several models on it at once, and compare the outputs side by side. No setup and no account required.

This page lists all 3 SpaceXAI vision models in the Roboflow Playground catalog, all proprietary models accessed through an API. They cover captioning, chart question answering, and classification, among other tasks. All of them run live in the Roboflow Playground on your own images. Compare licenses, parameters, prices, and release dates side by side, or open any model page for full details.