Google launches Gemini 4 Argon for phased rollout to cyber defenders

Google unveiled Gemini 4 Argon at $2 per million input tokens, featuring a 1-million-token output limit. Early tests freed over 300 TiB of data center memory and replaced 32,000 lines of code.

Google launches Gemini 4 Argon for phased rollout to cyber defenders

Google unveiled Gemini 4 Argon, a new frontier model priced at $2 per million input tokens and $10 per million output tokens. The company is initially releasing the system to a select group of cybersecurity defenders through its Fairwind Program before expanding access to developers and enterprise users. The launch follows a phased rollout strategy where the tech giant actively engages with the U.S. government's voluntary pre-release access process to refine safety guardrails.

Deep reasoning for complex workflows

Argon supports long-horizon tasks by increasing the output token limit to 1 million, a significant jump from the previous 64,000 tokens. This capacity allows the model to sustain deep reasoning across extended trajectories, enabling it to solve complex problems in a single pass rather than through fragmented steps. Early internal deployments have shown measurable gains in engineering efficiency and research capabilities.

Google reports that the model helped quantum computing researchers optimize spacetime resources in subroutines, beating a published baseline by 40% in minutes. Another team used Argon agents to analyze fleet-wide profiling telemetry, identifying memory optimizations across data centers that freed over 300 TiB of memory, with total savings estimated between 500 TiB and 1 PiB. These results demonstrate the model's utility in AI for Developers Courses-relevant contexts where resource management and code optimization are critical.

Codebase migration and performance

The model is currently handling large-scale migrations from C/C++ to Rust, a process that requires rigorous automated and manual auditing. For the open-source video decoder libgav1, Argon agents replaced 32,000 lines of SIMD code by running profile-guided experiments and studying compiler output. The resulting memory-safe Rust implementation runs 2.7x faster than the previous Rust port while producing identical video output.

On the DeepSWE v1.1 benchmark, which measures performance in real-world long-horizon software engineering tasks, Argon set a new state of the art with a score of 77.9%. Beyond coding, the model leads on the Vals Index, an evaluation metric that weights sectors like finance, legal, and coding by their contribution to U.S. GDP. It also ranked first on AutomationBench with a score of 51.3%, indicating strong capabilities in end-to-end execution of business functions.

Defensive cybersecurity capabilities

For trusted defenders, Google is releasing Argon without cyber guardrails to allow full utilization of its defensive capabilities. The model can autonomously find, validate, and patch critical software vulnerabilities. This approach is designed to equip security professionals with tools that match the sophistication of modern cyberattacks.

Wiz, a cloud security company, is already using Argon in its Scan for Good initiative. In an early test, the model uncovered a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide, a risk that previous frontier models missed. On the CWE-bench v1 evaluation, which tests vulnerability remediation, Argon tied for first place with a score of 68%. It also demonstrated superior performance on Wiz's internal black-box penetration testing benchmark, outperforming previous models in identifying attack surfaces and generating proof-of-concept evidence.

Safeguards and alignment monitoring

Google is strengthening safeguards against misuse, particularly for chemical, biological, radiological, and nuclear (CBRN) applications. The model is designed to refuse harmful requests while preserving legitimate dual-use scientific research. To support this, the company improved techniques to monitor the model's internal activations for signs of misuse, subjecting these systems to robustness testing by internal and external red teams.

Argon is also the most resilient model against indirect prompt injections, where malicious context is used to hijack behavior. It leads on the Gray Swan's Indirect Prompt Injection (IPI) benchmark. To prevent misalignment, Google deploys monitors that track the model's chain-of-thought and actions, stopping execution if the model steps out of bounds. The company emphasizes the importance of preserving reasoning transparency to help diagnose alignment risks, rather than shaping the model's reasoning to evade monitoring.

Why this matters for creatives, educators, and strategists

For Generative AI Courses participants and strategy leaders, the shift toward models capable of sustaining 1-million-token reasoning trajectories changes how complex projects are scoped. Educators and researchers should note the specific performance metrics on domain-heavy benchmarks like the Vals Index, which correlates economic impact with model capability. This data suggests that future workflows will rely less on single-step generation and more on multi-step, autonomous agents that can handle long-horizon tasks with minimal human intervention, requiring a reassessment of traditional project management and instructional design approaches.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)