Top
Should you care?News

Gemini 4 Argon: Google Unveils Its New Frontier Model for Software Engineering, Enterprise Work and Cyber Defense

Google has announced Gemini 4 Argon, a frontier AI model with a 1M-token output limit, rolling out first to cyber defenders. Here is what it can do, how it benchmarks and when you can use it.

Google chief executive Sundar Pichai beside a Gemini 4 Argon presentation graphic.
Google is positioning Gemini 4 Argon as a frontier model for long-horizon software engineering, enterprise work and cyber defense. Credit: Google.

Quick answer

Google announced Gemini 4 Argon on September 30, 2026, describing it as the next era of its frontier intelligence. The model is built for complex, long-horizon work in software engineering, enterprise knowledge work such as legal and finance tasks, and cybersecurity defense, with an output limit of 1 million tokens, up from 64,000. It is rolling out first to vetted cyber defenders through Google's Fairwind Program, with wider access planned for paid API customers and Google AI Ultra subscribers.

What Google announced

Google has introduced Gemini 4 Argon, a new frontier that it says delivers top-tier performance across three areas: real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. The announcement came from Koray Kavukcuoglu, Google's chief AI architect and head of Google DeepMind, who called it the company's next era of frontier intelligence.

The release is not a general launch. Argon is going first to a set of trusted cyber defenders through the Fairwind Program, and Google says broader availability to developers, enterprises and consumers will follow "as soon as possible," starting with paid customers and Google AI Ultra subscribers. The company says it is taking a phased approach and is taking part in the US government's voluntary process for pre-release model access while it gradually expands access.

The headline change: a 1 million output limit

The most concrete technical change is the output limit. Gemini 4 Argon can generate up to 1 million tokens in a single response, compared with 64,000 on earlier Gemini models. Many frontier models already accept very long inputs, but the amount they can write in one pass, including their own reasoning, has been far smaller.

Why does it matter? A model with more room to think and write can work through a hard problem in one continuous run instead of being cut off partway. Google argues that this headroom adds a new level of depth to reasoning on tough tasks. In practice it suits jobs such as rewriting large codebases, producing long analyses or carrying out multi-step research, where the work itself is long.

Software engineering performance

Google says Argon sets a new state of the art on DeepSWE v1.1, a benchmark of real-world, long-horizon software engineering tasks, with a score of 77.9%. In Google's published comparison, rival models scored lower on this test: GPT-6 Astra at 74.1%, Claude Opus 5.5 at 74.2% and Claude Fable 5.1 at 67.4%.

The picture is not one-sided. On FrontierSWE v2, another coding benchmark, Google's own table shows Argon at 55.0%, behind GPT-6 Astra at 65.5% and Claude Opus 5.5 at 62.3%. On Vibe Code Bench, Argon led with 91.9%. Reports that tallied Google's charts found that Argon led outright on 12 of the 18 benchmarks shown. That is a strong result, but it also means the model does not win everywhere, and enterprise buyers should still choose by workload.

Inside Google, thousands of employees already use Argon for specialized coding, research and writing. Google shared several examples:

- Code migration to Rust. Argon agents are working on moving C and C++ code to the -safe Rust language, from tens of thousands of lines in core libraries up to more than 800,000 lines for the Fuchsia Zircon kernel. Google says these rewrites are going through rigorous automated and manual auditing and testing before production. - Faster video decoding. For libgav1, Google's video decoder, Argon agents replaced 32,000 lines of low-level code in an existing Rust port. The result is a memory-safe decoder that runs 2.7 times faster than that port with identical video output. - Memory savings in data centers. A team of Argon agents analyzed fleet-wide profiling data and applied optimizations that Google says will free more than 300 TiB of memory once rolled out, with total savings estimated at 500 TiB to 1 PiB. - Quantum algorithm work. Argon helped Google's quantum computing researchers beat a published baseline by 40% on the resources needed for certain subroutines.

These are Google's own claims about its own systems, and independent verification will take time.

Enterprise knowledge work

Beyond code, Google says Argon leads the Vals Index, which measures economic impact across finance, coding, legal and tax work, with each sector weighted by its contribution to US GDP. It also reports leading results on domain tests such as the Vals Finance Agent v2 for multi-step financial research and Harvey's Legal Agent Benchmark for legal research and drafting.

Some of the numbers from Google's published table:

Selected benchmark results published by Google for Gemini 4 Argon and competing frontier models
FeatureGemini 4 ArgonGPT-6 AstraClaude Fable 5.1Claude Opus 5.5
Vals Index68.9%63.1%65.8%67.0%
AutomationBench51.3%41.4%31.4%42.5%
Vals Finance Agent v265.4%53.5%58.9%58.6%
Harvey Legal Agent Benchmark19.6%5.4%6.7%3.8%

The legal benchmark scores are low for every model, which is a useful reminder that these tests are hard and that even a leading result is far from perfect. Google also says Argon is particularly strong when knowledge work involves visual understanding, such as analyzing professional charts, picking out details in long videos and acting on a series of documents. It reports a state-of-the-art 91.7% on LVBench, a long-video understanding test.

Cybersecurity defense

Cyber defense is the focus of the first release. Google says it trained Argon to find, validate and patch critical software vulnerabilities on its own. For trusted defenders and Google's internal teams, the company will provide a version without cyber guardrails so they can use the full capability. That is the reason for the limited rollout through the Fairwind Program.

Google offers several data points:

- On Google's internal vulnerability benchmark, Argon uncovered a wide range of exposures across complex codebases in 20 programming languages. - Wiz, the Google-owned cloud security company, used Argon in its free Scan for Good initiative, which looks for exposures in critical public infrastructure. The model found a critical flaw that exposed sensitive personal information in healthcare software used by hospitals worldwide, a risk that earlier frontier models had missed. Google has not named the software or said whether a fix has shipped. - On CWE-bench v1, which tests the ability to fix security vulnerabilities, Argon tied for first place with a score of 68%.

The dual-use problem is obvious. A model that can find and patch flaws can also be used to find and exploit them, which is why Google is restricting the unguarded version to vetted users.

Safety measures before wider release

Google lists four areas of work to strengthen safeguards before a broad release.

1. Defending against misuse. The model is designed to refuse harmful requests, including cyber and chemical, biological, radiological and nuclear threats, while still supporting legitimate dual-use scientific work. Google says it is improving how it monitors the model's internal activations to spot misuse, and that safeguards were tested by internal and external red teams. 2. Prompt injection resistance. Google calls Argon its most resilient model yet against indirect prompt injection, where malicious instructions hidden in content try to hijack the model. It reports leading results on Gray Swan's benchmark for this threat. 3. Monitoring for misalignment. Systems watch the model's reasoning and actions and stop execution when it steps outside what the user intended. 4. Hardening systems. Google is isolating and sealing the sandboxed environments used for high-risk training and evaluation.

Pricing and availability

Google says Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced 95% lower than the standard input rate. After the introductory period, the price rises to $4 per million input tokens and $20 per million output tokens. No exact general release date has been given, only that it is rolling out soon.

Why the timing matters

Argon arrives after a difficult stretch for Google's AI efforts. Coverage notes that it follows the cancellation of a planned Gemini 3.5 Pro and a focus on the faster Gemini 3.8 Flash line, and that Google has faced reports of talent departures and delayed releases. Analysts see Argon as Google's strongest claim in months to frontier leadership on enterprise terms, where the question is whether an can safely rewrite code, audit systems, handle regulated knowledge work and defend software before attackers reach it.

What to watch next

- The date Argon opens to paid API customers and Google AI Ultra subscribers. - Independent testing of Google's benchmark claims. - What the Fairwind Program discovers and whether Google names the vulnerabilities found. - How rival labs respond on price and capability.

The tecMAMBO take

Argon is a serious release, and the 1 million token output limit is the sort of change that alters what an AI agent can attempt in one go. We would still treat Google's benchmark table the way we treat any vendor chart: useful, selective and unproven until outsiders reproduce it. The cyber angle is the most interesting and the most uncomfortable. Releasing an unguarded version to defenders first is a defensible call, but it shows how close the line between defense and attack has become. The real test comes when Argon reaches paying developers and has to prove itself on messy, everyday work.

FAQ

What is Gemini 4 Argon?

It is Google's new frontier AI model for software engineering, enterprise knowledge work and cybersecurity defense, announced on September 30, 2026.

Who can use it now?

For now, vetted cyber defenders in Google's Fairwind Program and Google's internal teams. Paid API customers and Google AI Ultra subscribers are next.

What is new technically?

An output limit of 1 million tokens, up from 64,000 on previous Gemini models.

How much will it cost?

An introductory $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward.

Does it beat rival models?

Google's benchmarks show it leading on most tests, but not all. It trails GPT-6 Astra on FrontierSWE v2, and the results are Google's own.

Sources

Ask MAMBO

Have a plain-English question about this topic? Send it in and we may answer it in a future guide.

Ask a question