LIVE
All stories ›
AI IN LIFENEWS
Tools & AppsBusiness & DealsAI ModelsResearchSocietyChips & ComputeSafety & SecurityRegulation & PolicyRoboticsReviews OpenAIAnthropicGoogle & DeepMindAlibaba / QwenMetaxAIByteDance
Home › Google & DeepMind › MODELS
MODELS

Gemini 4 Argon: Google's Model Patches Vulnerabilities

Argon reaches cyber defenders through the Fairwind program first, with a 1 million token output limit and 68 percent on CWE-bench v1.

Gemini 4 Argon: Google's Model Patches Vulnerabilities
Symbolic image: a hand hovers over the approval key of a security console while an open queue of vulnerability patches glows on screen.

In short

Gemini 4 Argon is Google's new frontier model built to find, validate and patch software vulnerabilities on its own, but for now only vetted cyber defenders in the Fairwind program can use it.

At a glance

  • Output limit: 1 million tokens, up from 64,000.
  • Introductory pricing: $2 per million input tokens, $10 per million output tokens; $4 and $20 afterward.
  • CWE-bench v1: 68 percent, which Google calls a tied first place for fixing vulnerabilities.
  • Terminal-Bench 4.0: 57.4 percent and fifth place, behind Claude Sonnet 5.5 at 70.6 percent (heise online).
  • Rollout starts via the Fairwind program; API customers and Google AI Ultra follow, with no date given.

Google introduced Gemini 4 Argon on September 30, 2026, pitching it as a model that finds, validates and patches software vulnerabilities without a human driving each step. You cannot buy it yet. Access begins with vetted cyber defenders in the Fairwind program, and Google says paid API customers and Google AI Ultra subscribers come next.

The headline change is output length

Argon can emit up to 1 million tokens in a single response, against a previous ceiling of 64,000. That matters less for chat than for long-horizon jobs such as rewriting a module or walking a full patch set. Introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input tokens discounted 95 percent, rising to $4 and $20 once the introductory window closes.

Google aims the model at three markets: real-world software engineering, enterprise knowledge work in legal and finance, and cybersecurity defense. The announcement is signed by Koray Kavukcuoglu, SVP at Google DeepMind.

The benchmark table is curated

Google cites 77.9 percent on DeepSWE v1.1, 51.3 percent on AutomationBench, 91.7 percent on LVBench for long video understanding, and 68 percent on CWE-bench v1, where it claims a tied first place. It also claims leading results on the Vals Index, Vals Finance Agent v2, Harvey's legal agent benchmark, and Gray Swan IPI for prompt injection robustness.

German outlet heise online points out what is missing. On Terminal-Bench 4.0 Argon scores 57.4 percent for fifth place, trailing Claude Sonnet 5.5 at 70.6 percent, Claude Opus 5.5 at 66.4 percent, GPT-6 Astra at 58.2 percent and Claude Fable 5.1 at 57.9 percent. GPT-6 Astra also leads on FrontierSWE v2 (65.5 against 55.0 percent), Terminal-Bench Science 0.1 (68.1 against 57.6 percent) and OSWorld-2.0 (72.6 against 69.2 percent).

Google's internal numbers, unaudited

Internally, Google says Argon beat a published baseline by 40 percent on quantum algorithmic optimization, freed more than 300 TiB of memory with an estimated 500 TiB to 1 PiB available, and helped port C and C++ code to Rust, including the Fuchsia Zircon kernel at over 800,000 lines. On the libgav1 video decoder it says 32,000 lines of SIMD code were replaced while running 2.7 times faster than the Rust port.

Every one of those figures comes from Google, and none has been measured independently. Because Argon is not publicly available, the benchmark table cannot be reproduced from outside either.

Safeguards, clearance and what is missing

Google plans to offer the model without its usual cyber guardrails to authorized security professionals. Its stated mitigations are monitored chain-of-thought against prompt injection, internal and external red teaming against misuse, misalignment monitoring during training, and sandboxed isolation before high-risk evaluations. The company says Argon is going through the U.S. government's voluntary pre-release model access process.

As evidence it works, Google points to security vendor Wiz, which used Argon in its Scan for Good initiative and says it uncovered a critical flaw exposing sensitive personal data in hospital software used worldwide. heise online notes the counterweight: an earlier Gemini misconfiguration had inadvertently granted access to real company systems. Google gives no date for general availability, only an intention to open access as soon as possible.

◈ AI-GENERATED REPORT · SOURCES LINKED

FAQ

Can I use Gemini 4 Argon yet?

No. Access runs through the Fairwind program for vetted cyber defenders first, with paid API customers and Google AI Ultra subscribers next. Google has not given a date.

How much does Gemini 4 Argon cost?

Introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input tokens 95 percent cheaper. Google lists $4 and $20 after that.

Is Argon better than GPT-6 Astra and Claude?

Not across the board. Argon leads on DeepSWE v1.1, AutomationBench, LVBench and CWE-bench v1, but its 57.4 percent on Terminal-Bench 4.0 ranks only fifth.

Sources

More reports