OpenAI introduces GPT-6 Astra for complex business work

OpenAI released GPT-6 Astra, a model that works through everyday applications without APIs. It caught 20% more bugs in code review and produced unintended outcomes 89% less often than its predecessor.

Published on: Sep 11, 2026
OpenAI introduces GPT-6 Astra for complex business work

OpenAI released GPT-6 Astra last week, making the model available in ChatGPT Work, Codex, and the API. The company positions it as the most intelligent and aligned model it has built, targeting complex professional work across software engineering, cybersecurity, financial analysis, and scientific research with what it claims is the majority of the cost-efficiency frontier on key benchmarks.

Astra's defining feature is its ability to work through the same applications people use daily, even when those applications lack an API. This means teams can deploy it within existing workflows without extensive data preparation or custom integrations. Early customers have already used it to optimize GPUs, spot discrepancies in financial statements, and produce on-brand presentations.

Performance across professional benchmarks

Third-party evaluations show consistent gains over previous models. On Databricks' OfficeQA Pro and Pro V2 benchmarks, Astra set a new state of the art using the Genie harness and delivered better cost per task than GPT-5.6 Sol. Hebbia reported that Astra produced the best decks in its testing and followed briefs 17% more faithfully than the next-best model, while sourcing claims to the correct document 19% more often.

Box's evaluation highlighted the model's judgment. "It was better at declining to assert conclusions the documents didn't support, and across the evaluation it was >10% less likely to make confidently incorrect assertions," said Yashodha Bhavnani, VP of AI Products at Box. On the software engineering side, Astra achieved 74% on DeepSWE v1.1 with fewer steps and greater token efficiency than any prior frontier model.

Code review performance also jumped. CodeRabbit found Astra caught roughly 20% more bugs than its baseline. On pull requests requiring extensive cross-file reasoning, the catch rate more than doubled. "It connects a change's intent to its consequences," said David Loker, VP of AI at CodeRabbit. "It reasons across files to catch interface-contract drift and authorization bugs the baseline missed, and it backs findings with concrete verification steps."

Internal deployment shows practical gains

OpenAI rolled out Astra internally weeks before launch. The engineering team used it to find and fix a memory-allocation bottleneck causing slow Codex sessions in a test environment. Switching allocators produced 25× lower turn latency with roughly 30% higher peak memory use. Marketing and developer teams used Astra and Codex to turn three hours of multicamera footage into a launch video that drew over 550,000 views in four days.

Basis reported a 20% improvement in pass rate for end-to-end agent workflows lasting more than five hours, with fewer inference calls needed to complete the work. "The model's improved decision making allows us to remove scaffolding and accelerate performance," said Mitch Troyanovsky, co-founder of Basis.

Safety and alignment for enterprise use

Giving AI access to business systems requires confidence in how it acts. Astra is the first model to reach the Critical cybersecurity capability threshold under OpenAI's Preparedness Framework. On the company's internal computer use safety benchmark, Astra produced unintended outcomes 89% less often than GPT-5.6 Sol and 74.7% less often than Claude Fable 5.1.

New enterprise admin controls let organizations restrict access to approved websites and desktop applications, manage uploads and downloads, and control browsing history. ChatGPT Work and Codex include confirmation policies that require approval before consequential actions, plus automated review of potentially unsafe tool calls. Zero Data Retention is available for eligible API customers on supported endpoints, subject to approval.

Pricing starts at $10 per million input tokens and $50 per million output tokens. Alongside Astra, OpenAI launched new enterprise plugins in ChatGPT Desktop from Oracle Analytics, Power BI, Navan, and Avalara, making it easier to access familiar enterprise applications.

Why this matters for IT, development, and product teams

For developers and IT professionals, Astra's ability to operate applications without APIs removes a major integration bottleneck. Teams can put the model to work on existing tools immediately, rather than waiting for custom connectors or workflow redesigns. The token efficiency gains also translate directly to lower per-task costs, which matters when running AI for IT & Development workloads at scale. Product teams building on the API get a model that follows brand voice and design standards more closely, producing output closer to a finished deliverable on the first pass - a practical improvement for anyone shipping client-facing work under tight deadlines.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)