Google Launches Gemini 4 Argon to Security Teams, Outperforming Rivals on Key Benchmarks
Google has begun distributing Gemini 4 Argon, its latest frontier AI model, to vetted cybersecurity professionals through a restricted access program. The model surpasses competing systems from Anthropic and OpenAI across most of Google's internal performance tests.

Google LLC has started distributing Gemini 4 Argon, a frontier-class artificial intelligence model, to qualified cybersecurity professionals. The rollout represents the company's first major deployment of the system, which outperforms competing offerings from Anthropic PBC and OpenAI Group PBC on the majority of Google's internal benchmarks.
Access to Argon remains limited outside Google's internal operations. Only participants in Google's Fairwind Program can currently use the model. Since launching on Sept. 3 with the smaller Gemini 3.8 Flash Cyber model, the program has attracted more than 650 organizations, including CrowdStrike Holdings Inc. and Palo Alto Networks Inc.
Koray Kavukcuoglu, chief AI architect at Google, explained that deploying capabilities at this level "requires a phased approach." The company is participating in the U.S. government's voluntary framework for early access to frontier models and continues strengthening protections against potential misuse in cyberattacks or weapons development.
Performance Against Competitors
Argon demonstrates particular strength in resisting indirect prompt injection attacks, a capability Google says exceeds any model the company has previously released. The system employs separate monitoring mechanisms that track the model's reasoning process and actions, with the ability to intervene if the model strays beyond user intent.
On DeepSWE v1.1, which evaluates extended software engineering tasks, Argon achieved 77.9%. Anthropic's Claude Opus 5.5, released the previous week, scored 74.2% on Google's evaluation, placing it one-tenth of a point ahead of OpenAI's GPT-6 Astra. The AutomationBench test, which assesses comprehensive business workflows, showed a more pronounced difference: Argon reached 51.3% compared to Opus 5.5's 42.5%. Across Google's 18-benchmark suite, Argon led outright on 12 measures. Opus 5.5 retained the top position on Terminal-Bench 4.0, while Astra maintained leadership on FrontierSWE v2.
Internal Deployment and Real-World Applications
Google engineered Argon specifically for the extended-duration work that DeepSWE measures, setting the model's output capacity at 1 million tokens—a significant increase from the 64,000-token ceiling on earlier Gemini versions.
Thousands of Google staff members are already operating Argon, with teams deploying the model's agents on substantial engineering initiatives. One major project involves converting C and C++ code to Rust across multiple codebases, ranging from core libraries containing tens of thousands of lines to the Zircon kernel in Google's Fuchsia operating system, which exceeds 800,000 lines. On libgav1, Google's open-source video decoder library, agents began with an existing Rust implementation and rewrote 32,000 lines of performance-critical code as safe Rust that the compiler could independently optimize. This revision produced a 2.7-fold speed improvement in the decoder with identical video output.
Google engineers have also deployed Argon agents to identify memory inefficiencies. A comprehensive scan of fleet-wide profiling information has recovered more than 300 tebibytes of memory across Google's infrastructure.
Cybersecurity Capabilities and Pricing
Fairwind participants and Google's internal teams receive a variant of Argon with cybersecurity guardrails disabled. This version can independently locate and remediate software vulnerabilities. On CWE-bench v1, a vulnerability remediation evaluation, it matched GPT-6 Astra's performance at 68%.
Wiz Inc., the Google-owned cloud security firm, has integrated Argon into its complimentary Scan for Good initiative, which identifies security gaps in critical public infrastructure. Google reported that the model discovered a critical vulnerability exposing personal health information in hospital software deployed internationally—a flaw that earlier frontier models had overlooked.
Google has not announced a timeline for availability to paying developers and Google AI Ultra subscribers, who represent the next tier of access. At launch, Argon will be priced at $2 per million input tokens and $10 per million output tokens, with cached input discounted by 95%. Anthropic prices Opus 5.5 at $4 and $20 respectively, and Argon will transition to those rates once the introductory pricing period concludes.


