Gemini 4 Argon: what Google's new model changes

Gemini 4 Argon: what Google's new model changes

Google announced Gemini 4 Argon on September 30, 2026. The official announcement presents it as a new frontier model for complex software engineering, professional legal and finance work, and cyber defense. There is no general release yet: first access goes to Fairwind Program participants selected by Google and invited testers.

The product question behind a release like this is which tasks you can now hand to a model end to end, from brief to verified result. We have not tested Argon ourselves, so this analysis covers what Google and the independent Vals leaderboard have published: real work examples, benchmarks and access terms. What was known about Gemini 4 before the announcement is covered in our August article.

A million output tokens

Google raised the maximum output from 64,000 to 1 million tokens. The company's reasoning: with headroom for hundreds of thousands of tokens in a single trajectory, the model can reason longer and solve a hard problem in one go.

Maximum output tokens
Before
64K
Gemini 4 Argon
1M

This is the output limit, not the input context size.

A short email or translation usually does not need that much headroom. It matters where the model previously had to be stopped and restarted: a large code migration, multi-step research, investigating a vulnerability with verification. A higher limit does not prove quality on its own, though: a model can also spend longer pursuing a wrong hypothesis. Examples of finished work matter more.

What Argon already does inside Google

The most telling examples in the announcement involve code. These are Google's internal figures without independent verification, but they show what kind of work the model is built for.

C/C++ to Rust migration. Argon agents are migrating Google codebases to Rust, from tens of thousands of lines in libraries like re2 and libgav1 up to more than 800,000 lines in the Zircon kernel of the Fuchsia operating system. Google notes that these rewrites go through automated and manual auditing, emulation testing and review before reaching production.

Speeding up libgav1. A Rust port of this video decoder already existed. The agents replaced 32,000 lines of SIMD code in it: they ran rounds of profile-guided experiments, studied the compiler's output and wrote safe Rust that the compiler vectorizes on its own. The decoder became 2.7 times faster than the previous Rust port, with identical video output. The comparison is against the Rust port, not the optimized C++: on C++, Google only says the gap has narrowed.

Data center memory. Agents analyzed fleet-wide profiling telemetry, then found and applied memory optimizations. The rollout freed more than 300 TiB, and Google estimates total savings at 500 TiB to 1 PiB.

All three cases have an external way to verify the result: tests and emulation, comparison of decoder output, a profiler. They are good candidates for extended agent work because correctness can be assessed without relying on the model's prose.

Cyber defense: find, validate, patch

According to Google, Argon autonomously finds critical software vulnerabilities, validates them and prepares patches. Google gives trusted defenders and its own internal teams the model without cyber guardrails.

The most concrete example in the announcement comes from Wiz. Through its Scan for Good initiative, the company finds high-risk exposures in critical public infrastructure for free. Using Argon, it uncovered a critical vulnerability in healthcare software used by hospitals worldwide that exposed sensitive personal information. Previous frontier models had missed it.

On CWE-bench v1, which tests vulnerability remediation, Google reports a 68% score for Argon, tying it for first place. Google's and Wiz's internal tests also show Argon outperforming Gemini 3.8 Flash Cyber at finding vulnerabilities, including in live web systems without access to source code.

Benchmarks and the Vals Index

Google's own results:

EvaluationWhat it testsArgon
DeepSWE v1.1long-horizon software engineering77.9%
AutomationBenchend-to-end business workflows51.3%
LVBenchlong video understanding91.7%

Percentages across rows are not comparable: the tests differ in tasks and criteria. Google also claims the lead on Vals Finance Agent v2 (multi-step financial research) and Harvey's Legal Agent Benchmark (legal research and drafting), but the announcement gives no figures for them.

The Vals Index offers an independent check. It combines finance, coding, legal and tax tasks, weighted by each sector's share of US GDP. Finance carries 8 of 15.3 weight points, about 52% of the index, and coding about 37%. Keep that weighting in mind when choosing a model for another field.

The top of the leaderboard, plus Gemini 3.8 Flash for reference, as of September 30:

ModelAccuracy
Gemini 4 Argon68.90%
Claude Sonnet 5.567.04%
Claude Opus 5.566.97%
Gemini 3.8 Flash54.83%

Argon ranks first, 1.86 points ahead of Claude Sonnet 5.5. For a direct comparison, keep in mind that Vals flags some Claude results where a fallback model handled tasks after a provider refusal. Gemini 3.8 Flash trails Argon by about 14 points. These are results on the Vals task set, not a guarantee of the same gap on your own tasks.

For document work, the practical takeaway is this. In finance and law, a high score does not replace checking calculations and their applicability to a specific situation. Frame the task so that every significant conclusion cites a document and data gaps are listed separately: that shows where the model papered over missing information with plausible prose.

Why access is rolling out in phases

A model that finds and fixes vulnerabilities on its own can also help an attacker. That is why Google is releasing Argon in phases, starting with defenders. The first circle is the Fairwind Program. According to its page, it covers selected trusted organizations and defenders: governments, healthcare, telecommunications and other vital infrastructure. Applications are vetted, access is restricted, and participants get exclusive access to Argon with its cybersecurity capabilities. Program rules prohibit sharing or reselling that access, so third-party services cannot be built on it.

In parallel, Google is taking part in the US government's voluntary pre-release model access process and refining its safeguards based on feedback from early testers. The work covers four areas: refusing to help with cyberattacks or chemical, biological, radiological and nuclear weapons; resilience to prompt injection; monitoring the model's reasoning and actions and stopping it if it goes beyond the user's intentions; and isolating test environments. According to the announcement, the broad release should refuse harmful requests while preserving legitimate dual-use research. How that will affect everyday security work is not yet known.

Google will begin the release to developers, businesses and consumers with paid API customers and Google AI Ultra subscribers. No date has been given.

Argon's availability date and GPTunneL pricing have not been announced. For models available now, see our Gemini page and our price list.

What to do before access opens

Argon is built for long work with a verifiable outcome, and Google's examples show this better than benchmark percentages. If you have tasks like that (code migration, profile-guided optimization, financial models, vulnerability triage), collect 10 to 20 of them with acceptance criteria and run them on your current model now: share of accepted results, manual corrections and time. Once Argon becomes available to you, the same set will show whether it improves quality on your tasks. Short tasks that your current model already handles reliably are not worth moving: for those, see Gemini 3.8 Flash and other models in the family or the text model catalog.