Gemini 4 Argon
Gemini 4 Argon is Google DeepMind's frontier model announced September 30, 2026 for software engineering, enterprise knowledge work and cybersecurity defense. Initial access is restricted to trusted cyber defenders through Fairwind; the announcement does not establish general availability.
Check Google's Argon announcement ↗At a glance
Selected results from Google DeepMind's September 2026 comparison, checked October 2. Its methodology mixes self-computed runs and externally reported scores; harnesses, thinking levels and budgets matter. These are not one uniform independent head-to-head test.
| Decision | Gemini 4 Argon | GPT-6 Astra | Claude Opus 5.5 |
|---|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.1% | 74.2% |
| FrontierSWE v2 | 55.0% | 65.5% | 62.3% |
| Terminal-bench 4.0 | 57.4% | 58.2% | 66.4% |
| Vals Index | 68.9% | 63.1% | 67.0% |
| LVBench | 91.7% | 87.5% | 83.7% |
| CWE-bench v1 | 68.0% | 68.0% | 67.0% |
DeepSWE v1.1
- Gemini 4 Argon
- 77.9%
- GPT-6 Astra
- 74.1%
- Claude Opus 5.5
- 74.2%
FrontierSWE v2
- Gemini 4 Argon
- 55.0%
- GPT-6 Astra
- 65.5%
- Claude Opus 5.5
- 62.3%
Terminal-bench 4.0
- Gemini 4 Argon
- 57.4%
- GPT-6 Astra
- 58.2%
- Claude Opus 5.5
- 66.4%
Vals Index
- Gemini 4 Argon
- 68.9%
- GPT-6 Astra
- 63.1%
- Claude Opus 5.5
- 67.0%
LVBench
- Gemini 4 Argon
- 91.7%
- GPT-6 Astra
- 87.5%
- Claude Opus 5.5
- 83.7%
CWE-bench v1
- Gemini 4 Argon
- 68.0%
- GPT-6 Astra
- 68.0%
- Claude Opus 5.5
- 67.0%
Why it matters
Its long-task focus makes it relevant to repository agents and document-heavy work. Access, accepted results and total task cost should guide a switch, rather than a single leaderboard position.
Who can use Argon now?
As checked October 2, Google names trusted Fairwind defenders and testers first, with paid API customers and Google AI Ultra subscribers planned for wider access. No date is given. A Pro subscription or Antigravity account alone is not evidence of eligibility.
- Verify availability on your actual account before paying or migrating.
- The current public Gemini API catalog, release notes and pricing page do not identify an Argon endpoint. Do not guess a model ID from the display name.
Output limit and context are different
Google announces up to 1 million output tokens, increased from 64K. This is an output limit, not a statement of the input context window. Artificial Analysis separately lists 1M context for its evaluated High configuration; treat that as evaluator metadata until a public Google API contract confirms the limits.
Introductory pricing and the later rate
Google announces $2 input and $10 output per million tokens, with a 95% cached-input discount ($0.10, calculated). After the introductory period, input/output become $4/$20. The announcement gives no expiry date. These are announced rates, not proof of a self-service endpoint.
- Artificial Analysis reports $1.99 per Intelligence Index task for High. That is the cost of its evaluation workload, not a quote for your task.
- Compare successful task cost, output volume and retries; equal token prices need not mean equal total cost.
Read the benchmarks with their conditions
The table shows strengths and weaker results. Argon leads the selected DeepSWE and Vals Index rows, but trails Astra on FrontierSWE and Opus 5.5 on Terminal-bench. DeepMind self-computes Argon's DeepSWE result with mini-swe-agent and draws comparator scores from other published evaluations. LVBench also uses different video sampling limits across models. Read the methodology before treating a gap as a universal advantage.
- Arena's September 30 Text board, observed October 2, lists gemini-4-argon-high first at 1525 ±9 with 4,942 votes and a Preliminary label. This is human text preference, not a guarantee for coding or autonomous work.
- Artificial Analysis records Intelligence Index 53 for High. Keep its scale and configuration separate from Google's task benchmarks and Arena ratings.
Argon, Flash and the rumored Pro name
Use the official name Gemini 4 Argon. Older community claims about Gemini 4 Pro or a disguised Flash checkpoint do not establish an alias or the identity of a demonstration. Gemini 3.8 Flash has its own documented public API contract; its Cyber variant has separate access restrictions.
What to evaluate before switching
The Hacker News launch discussion questions reproducibility, harness quality and task cost; Reddit users ask when their paid development tools will receive access. Those are useful evaluation questions, not access guarantees. Once eligible, test a representative task with fixed tools and acceptance criteria, record cost and latency, and inspect the result before widening permissions.
- For coding, check working changes and tests with the same repository and tool budget.
- For document work, check citations and material omissions against the same source set.
- A leaderboard result does not grant permission for unsupervised security testing.
Useful first jobs
- Evaluate long-running repository tasks when access is granted.
- Compare document-heavy knowledge work with source verification.
- Assess defensive security workflows within authorized environments.
Frequently asked questions
Can I use Gemini 4 Argon with a Pro subscription?
Pro access is not established by the announcement. Google describes a restricted Fairwind rollout first and planned paid API/Ultra access; check your actual product and account before buying.
Is the 1 million token limit the context window?
The official launch figure is the output limit. Artificial Analysis separately lists 1M context for High, but that evaluator field is not a complete public Google API contract.
Is Argon better than GPT-6 Astra or Claude Opus 5.5?
It depends on the workload. DeepMind's table shows Argon ahead on DeepSWE v1.1, behind Astra on FrontierSWE v2 and behind Opus 5.5 on Terminal-bench 4.0. Methods and budgets differ; test accepted outcomes on your work.
Sources and evidence
Google establishes the announcement and rollout. DeepMind explains the mixed benchmark methods. Arena and Artificial Analysis provide separate evaluator observations; those observations do not establish public API access.
- Gemini 4 Argon announcement ↗
Google · Primary source
- Gemini 4 Argon capabilities and results ↗
Google DeepMind · Primary source
- Argon evaluation methodology ↗
Google DeepMind · Primary source
- Public Gemini API model catalog ↗
Google AI for Developers · Primary source
- Gemini API release notes ↗
Google AI for Developers · Primary source
- Gemini API pricing ↗
Google AI for Developers · Primary source
- Text Arena leaderboard ↗
Arena · Primary source
- Gemini 4 Argon (High) analysis ↗
Artificial Analysis · Primary source
- Launch discussion and evaluation questions ↗
Hacker News · Community experience
- Antigravity paid-plan access question ↗
Reddit · Community experience