SpaceX launches Grok 4.7 with long-horizon processing and safety upgrades

SpaceX's Grok 4.7 averaged $4.69 per task on CursorBench 4.0, beating GPT-5.6 Sol and Fable 5.1. Pricing starts at $2 per million input tokens, with a faster variant at double the cost.

Published on: Sep 23, 2026
SpaceX launches Grok 4.7 with long-horizon processing and safety upgrades

SpaceX rolled out Grok 4.7 on Tuesday, a new large language model that tops several third-party benchmarks while undercutting competitors on cost. The release comes less than a week after the company launched a separate text-to-speech model, signaling an accelerating cadence for its AI portfolio.

The model scored an average cost of $4.69 per task on CursorBench 4.0, a benchmark developed by another recent SpaceX acquisition. That put it ahead of GPT-5.6 Sol and Fable 5.1, both tested in hardware-intensive configurations that prioritize accuracy over efficiency. Grok 4.7 also outperformed Fable 5.1 on the Harvey Legal Agent Benchmark and EEBench, which measure legal reasoning and chip design capabilities respectively. On EEBench, however, it trailed GPT-6 Astra from OpenAI Group PBC.

Under the hood: base model and reinforcement learning

SpaceX attributed the gains to a new base model - the initial version of an LLM that emerges after the first phase of training. Engineers also retooled the reinforcement learning workflow that sharpens reasoning. Compared with Grok 4.6, the model received harder training tasks and worked on them for longer periods.

The system plugs into the Grok Bot harness, a set of tools that lets Grok 4.7 distribute complex jobs across multiple AI agents. Those agents run tasks in parallel and cross-check each other's output, which SpaceX said speeds up processing and improves reliability.

Safety benchmarks and pricing

SpaceX added new guardrails focused on blocking malicious requests. The company said Grok 4.7 set records on LatchBio, which tests an LLM's ability to refuse harmful biology research prompts, and HackerBench, which measures resistance to cybersecurity exploits.

Pricing starts at $2 per million input tokens and $6 per million output tokens. A low-latency variant processes prompts twice as fast for double the cost, aimed at workloads where response time matters more than budget. Developers working with large-scale AI deployments can explore related techniques through Generative AI Courses that cover model selection and cost optimization.

Why this matters for product and engineering teams

Grok 4.7's price-performance profile on CursorBench 4.0 suggests SpaceX is competing directly on cost per task, not just raw accuracy. For teams building agent-based workflows, the Grok Bot harness offers a concrete architecture for parallel execution and output verification - two features that reduce the risk of compounding errors in multi-step automation. The twice-weekly release tempo also signals that model updates are becoming operational events rather than annual milestones, which changes how teams plan integration cycles.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)