Google’s Gemini 4 AI Reaches Selected Cybersecurity Defenders

Google's Gemini 4 Argon AI is reaching selected cyber defenders. It can generate longer text in one run; ordinary developers and subscribers have to wait for access.

TL;DR
  • Defender Access: Google’s Gemini 4 Argon AI model is reaching selected cybersecurity partners while ordinary developers and subscribers await access.
  • Output Capacity: Its output limit rises to 1 million tokens, the text units models generate, allowing longer output in one run.
  • Coding Results: Google reports 77.9% pass@1, or first-attempt success, for Argon on DeepSWE v1.1 coding tasks; rivals lead other coding tests.

Google is rolling out Gemini 4 Argon, its new flagship AI model, to selected cybersecurity partners while ordinary developers and subscribers await access. Google engineers already use it for software work, and security company Wiz uses it to investigate vulnerabilities.

The September 30 rollout starts with organizations in Google’s Fairwind program, which grants vetted defenders access to advanced AI. Google is strengthening safeguards before offering Argon more broadly.

More Room for Software Work

Argon’s output ceiling is 1 million tokens, the units of text it can generate in one run. Google’s Gemini 3.1 Pro release in February allowed 64,000 output tokens alongside a 1 million-token input window for material the model could read. Google intends the larger output allowance to support sustained reasoning and generation through longer tasks.

Long software workflows were already part of Gemini 3.8 Flash’s September 2 release. That model reached developer tools, enterprise products and consumer subscriptions, while its Flash Cyber variant served vetted defenders through Fairwind. Argon adds much more output headroom and new performance claims to a lineup that already supported multi-step coding and security work.

Google says Argon agents reworked part of libgav1, its open-source video-decoding software. They replaced 32,000 lines of SIMD code, which processes several data items at once, in an existing version written in the Rust programming language. The agents repeatedly tested performance and inspected how the compiler translated code into executable instructions. They then produced memory-safe Rust, which guards against invalid memory access, while letting the compiler generate parallel operations automatically.

With identical decoded video, the resulting implementation ran 2.7 times as fast as the earlier Rust version, Google reports. The company’s larger migrations from the C/C++ programming languages to Rust reach more than 800,000 lines in Fuchsia’s operating-system core. Those critical rewrites are undergoing automated and manual auditing, testing and review before production deployment.

Rivals Retain Advantages on Different Tasks

Google’s coding comparison puts Argon ahead of OpenAI’s GPT-6 Astra and Anthropic’s Claude Opus 5.5 on DeepSWE v1.1, which evaluates code patches and repository changes. Astra leads a different kind of work on FrontierSWE v2: substantial engineering and research projects. Opus leads Terminal-Bench 4.0, which tests completing computer tasks through terminal commands.

Google’s Reported Coding Scores

Evaluation Gemini 4 Argon GPT-6 Astra Claude Opus 5.5
DeepSWE v1.1: repository changes, pass@1 (success on one attempt) 77.9% 74.1% 74.2%
FrontierSWE v2: project work, average task score 55.0% 65.5% 62.3%
Terminal-Bench 4.0: terminal tasks, pass@1 57.4% 58.2% 66.4%

FrontierSWE’s project-based assessment gives agents up to 20 hours per run and awards partial credit across five trials. Its tasks include improving a compiler’s generated machine code without slowing compilation or breaking tests. Astra’s lead therefore concerns a broader project-oriented assessment than the repository-change test where Argon leads.

Google Argon’s DeepSWE result using a software agent called mini-swe, while taking Astra’s score from the benchmark leaderboard and Opus’s from its system card, Anthropic’s report on the model. Its evaluation methodology uses the highest-scoring reasoning settings available for these comparisons. VentureBeat’s analysis of the task differences says most customers need broader access before testing whether Argon’s benchmark advantages reduce production costs on real workloads.

Bloomberg reported that people with direct access to Gemini 4’s development found it struggled with certain coding tasks in employee use despite good benchmark results. Google disputed that characterization, citing encouraging performance comments from Koray Kavukcuoglu, who leads its DeepMind AI division. The report also included a contrary account from an employee familiar with model development, who described broad internal agreement that Gemini 4 was competitive with leading AI models and rejected the coding criticism.

Harvey’s Legal Agent Benchmark asks agents to produce legal research and drafts from supplied files. Each task passes only when every grading criterion is met; benchmark evaluator Vals AI averages task-pass rates from two AI judges. Google reports 19.6% for Argon, compared with 5.4% for Astra and 3.8% for Opus.

Vetted Access for Cyber Defenders

Fairwind partners can use Argon in CodeMender, Google’s software-security agent for finding and fixing vulnerabilities. Google plans to supply trusted defenders and its internal teams with Argon without cyber guardrails, the cyber-specific restrictions on what the model will help users do.

The Fairwind access terms still limit use to authorized defensive and research tasks. Organizations may grant access only to their internal security, incident-response or penetration-testing teams. They must authenticate users, use phishing-resistant multifactor authentication and track employee access and use; they cannot share, redistribute or sell model access. Google vets applicants’ security histories and ethical records.

On CWE-bench v1, a defensive benchmark developed by Collinear AI, Argon ties GPT-6 Astra and xAI’s Grok 4.7 at 68% on the programmatic pass@1 measure. Agents receive repositories and must discover and fix vulnerabilities while preserving existing functionality. The assessment uses 120 held-out tasks, four runs per task and a one-hour limit per run, with different software-agent setups for the models.

gemini_4_cyber_evals_cwe_bench leaderboard

Google reports that Argon also outperformed 3.8 Flash Cyber in Wiz’s internal assessment of live web systems without source code. Argon did better at mapping what an attacker could reach, finding vulnerabilities and producing evidence to validate them.

Google reports that Wiz’s Scan for Good initiative used Argon to find a critical flaw that could expose sensitive personal information in healthcare software used by hospitals worldwide. The initiative provides free security work for critical infrastructure. The software and affected versions remain unnamed, and patch-deployment status is unclear.

Public Access and Introductory Pricing

Before wider release, Google is working on defenses against harmful requests and prompt injection, where hostile instructions in material an agent reads can redirect its behavior. It also describes monitoring the model’s reasoning and actions for behavior beyond the user’s intent, and isolating environments used for high-risk testing. The company is participating in the U.S. government’s voluntary pre-release model-access process.

Paying users of Google’s API, which lets developers connect the model to their software, and Google AI Ultra subscription customers are planned as the first wider access groups. Google has given no date for that release.

Announced introductory API rates will be $2 per million input tokens and $10 per million output tokens. The reported post-introductory rates double to $4 and $20 respectively, after a period whose duration remains unclear. Cached input, which reuses previously supplied material, is priced at 95% below the introductory input rate.

Markus Kasanmascheff
Markus Kasanmascheff
Markus has been covering the tech industry for more than 15 years. He is holding a Master´s degree in International Economics and is the founder and managing editor of Winbuzzer.com.
Subscribe
Notify of
guest
0 Comments
Newest
Oldest Most Voted