OpenAI expands GPT-6 lineup with Sol and Luna models

OpenAI launched GPT-6 Sol and GPT-6 Luna, cutting API prices 50% below GPT-5.6 promotional rates-Luna costs $0.10 per million input tokens and $0.50 per million output tokens.

OpenAI expands GPT-6 lineup with Sol and Luna models

OpenAI launched two new models, GPT-6 Sol and GPT-6 Luna, on Thursday, expanding the GPT-6 family with lower-cost options for professional work, coding, and everyday tasks. Both models cut API prices by 50% compared to their GPT-5.6 promotional pricing, making frontier intelligence more accessible for applications that run at scale.

The models arrive three weeks after GPT-6 Astra, which OpenAI calls its most intelligent and aligned model. While Astra remains the top performer for the most demanding projects, Sol and Luna target the broader cost-intelligence curve where sustained usage and iteration speed matter as much as raw capability.

"We trained GPT-6 Sol and Luna with similar methods as GPT-6 Astra, bringing the advances behind Astra's state-of-the-art performance in professional work, factuality, coding, computer use, and alignment to faster, more affordable models," the company said.

Pricing and performance across the stack

GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens. GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens. Both represent a 50% drop from their GPT-5.6 equivalents. OpenAI attributes the savings to improvements in caching and inference infrastructure.

On AutomationBench, which tests business workflows across apps, GPT-6 Sol at xhigh effort scored 33.2% - outperforming Claude Opus 5 at max effort (26.9%) at roughly 9% of Opus 5's cost per task. GPT-6 Luna at high effort improved on its predecessor by 5.4 percentage points while costing 58% less per task.

On Agents' Last Exam, which evaluates agents on complex professional workflows spanning 55 sub-industries, GPT-6 Sol at max effort scored 56.4%. That topped Claude Opus 5's highest reported score at 60% lower cost per task.

Coding and factuality gains

OpenAI reports that daily token usage for coding agents inside the company has grown exponentially, with median researchers exceeding $600 per day in API-value terms and 90th-percentile researchers surpassing $7,000. For teams pushing coding agents into longer, more complex tasks, sustained cost matters.

On FrontierCode, which evaluates whether coding agents produce merge-ready changes in real codebases, GPT-6 Sol improved substantially over GPT-5.6 Sol and matched Claude Fable 5.1 xhigh at much lower cost. On DeepSWE v1.1, which tests complex software engineering tasks, GPT-6 Sol at max effort scored 68.8% - within 1.1 percentage points of Claude Fable 5's highest score - at approximately 80% lower cost per task. GPT-6 Luna at max effort scored 66.6%, comparable to Claude Opus 5 and Fable 5 at medium effort, while costing 93% less per task than Opus 5.

On an internal factuality evaluation built from real-world conversations where users flagged model errors, GPT-6 Sol made about half as many mistakes as its predecessor. GPT-6 Luna at higher effort levels matched GPT-5.6 Sol's reliability at roughly one hundredth the cost.

Computer use and collaboration style

On OSWorld 2.0 offline, which tests long-horizon computer-use workflows, GPT-6 Sol at xhigh effort scored 60.5%, similar to Claude Opus 5 at medium effort (60.3%) at approximately 80% lower cost per task. GPT-6 Luna at max effort exceeded GPT-5.6 Sol at medium effort at one tenth the cost.

OpenAI also brought Astra's communication style to Sol and Luna. The company said users should notice "more clarity, less jargon, fewer odd turns of phrase, fewer low-value details, and slightly shorter answers overall without losing substance." A side-by-side comparison showed GPT-6 Sol giving a more direct, less presumptuous response to a web design request than its predecessor.

Caching improvements for developers

Alongside the price cuts, OpenAI improved prompt caching for GPT-6 to deliver higher cache hit rates by default. Cached input-token reads now receive a 90% discount. A new Prompt Caching Dashboard lets developers monitor cache usage, and a diagnostics tool explains missed caching opportunities. Developers can also set explicit breakpoints to control which prompt prefixes get cached, and adjust reasoning effort or tool availability without breaking the cache.

GitHub reported that these caching improvements reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests to OpenAI models, helping Copilot respond faster.

Why this matters for IT, development, and business leaders

For teams building on the OpenAI API, the 50% price cut and improved caching directly lower the cost of running agents at scale. The performance data suggests organizations can now get Astra-adjacent quality on many professional and coding benchmarks without paying Astra-level prices. For managers evaluating build-versus-buy decisions around coding agents, the cost-per-task comparisons against competing models provide a concrete framework for budgeting sustained AI usage across development, support, and operations workflows. Professionals looking to build skills with these models can explore OpenAI API Courses or broader Generative AI Courses to understand implementation patterns.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)