Skip to content
Models / Google

Gemini 4 Argon

Google · Model release

Released

Catalog checked 2026-09-30

Max output
1M
01 / API priceUSD / 1M tokens

Announced future API pricing: introductory $2 / 1M input and $10 / 1M output, with cached input at 95% off input. After the introductory period: $4 / 1M input and $20 / 1M output. API effective date and introductory expiration date are unknown, so these are not executable rates.

02 / Capabilities & limitsProvider specifications
Maximum output
1,000,000

Initially rolling out to trusted cyber defenders. Broader availability is announced for later, starting with paid API customers and Google AI Ultra subscribers; current broad API or Ultra access is not established.

The launch announcement establishes a 1,000,000-token output limit, not a context window. An exact public API model identifier is not established.

Google DeepMind · Sep 30, 2026

Developer reported · Gemini 4 Argon launch

Show all 17 results
Methodology & sources

Google’s launch sheet combines evaluations it computed with results it republished from leaderboards and system cards. This is one reporting set; it does not establish identical harnesses across models. Argon generally uses the Gemini API at the highest thinking setting, with a single attempt and trial averaging on smaller benchmarks.

These scores are reported by Google DeepMind. Model developers identify authorship; results may originate with another evaluator. Exact benchmark versions remain distinct. StackReplay calculates no composite score.

DeepSWE v1.1 · 77.9%

Long-horizon software engineering in real-world codebases.

Version: v1.1 · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Google DeepMind (source-computed)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
mini-swe
Tools
Not reported
Fallback
Not reported
Provider
Gemini API
Checked
Sep 30, 2026

Argon uses Google’s mini-swe agent; Astra is reported from the leaderboard, Fable and Opus from system cards. Thinking levels follow the highest scoring level reported by Datacurve.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

FrontierSWE v2 · 55.0%

Agentic software engineering on coding tasks.

Version: v2 · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Proximal public leaderboard (republished result)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
Not reported
Fallback
Not reported
Provider
Not reported
Checked
Sep 30, 2026

Results republished from Proximal’s public leaderboard.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

Vibe Code Bench · 91.9%

Code generation for prompted software-building tasks.

Version: Not reported · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Vals AI public leaderboard (republished result)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
Not reported
Fallback
Not reported
Provider
Not reported
Checked
Sep 30, 2026

Results republished from the Vals AI public leaderboard.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

Terminal-Bench 4.0 · 57.4%

Agentic task completion in a terminal environment.

Version: 4.0 · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Google DeepMind (source-computed)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
Not reported
Fallback
Not reported
Provider
Gemini API
Checked
Sep 30, 2026

Argon is computed by Google; competitor results are republished from the official leaderboard, using the highest scoring thinking level reported by its authors.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

PostTrainBench v1.1 · 45.3%

Post-training of language models under a fixed compute budget.

Version: v1.1 · Benchmark-reported weighted score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Google DeepMind (source-computed)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
OpenCode
Tools
Single NVIDIA H100, 10-hour budget
Fallback
Not reported
Provider
Gemini API
Checked
Sep 30, 2026

Google computed every model with OpenCode and a 10-hour budget on one NVIDIA H100; the benchmark’s metric aggregates four base models and seven benchmarks.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

Vals Index · 68.9%

Economic knowledge work across finance, coding, legal and tax tasks.

Version: Not reported · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Vals AI (republished result)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
Not reported
Fallback
Not reported
Provider
Not reported
Checked
Sep 30, 2026

Vals AI results republished by Google.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

AutomationBench · 51.3%

End-to-end execution of business workflows.

Version: Not reported · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Zapier public leaderboard (republished result)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
Not reported
Fallback
Not reported
Provider
Not reported
Checked
Sep 30, 2026

Private task set; results republished from Zapier’s public leaderboard.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

Vals Finance Agent v2 · 65.4%

Multi-step financial research by an agent.

Version: v2 · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Vals AI (republished result)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
Not reported
Fallback
Not reported
Provider
Not reported
Checked
Sep 30, 2026

Vals AI results republished by Google.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

Harvey’s Legal Agent Benchmark · 19.6%

Legal research and drafting by an agent.

Version: Not reported · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Vals AI (republished result)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
Not reported
Fallback
Not reported
Provider
Not reported
Checked
Sep 30, 2026

Vals AI results republished by Google.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

Terminal-Bench Science 0.1 · 57.6%

Scientific and mathematical tasks in a terminal environment.

Version: 0.1 · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Google DeepMind (source-computed)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
Not reported
Fallback
Not reported
Provider
Gemini API
Checked
Sep 30, 2026

Argon uses a sixfold verifier timeout to address verification timeouts; competitor results come from the official leaderboard.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

LABBench 2 · 88.8%

Scientific tasks with a terminal, bioinformatics tools and internet access.

Version: 2 · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Google DeepMind (source-computed)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
Linux terminal, bioinformatics tools, Python, R and internet
Fallback
Not reported
Provider
Gemini API
Checked
Sep 30, 2026

Google computed every model with a Linux terminal, bioinformatics tools, Python, R and internet access.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

RiemannBench · 76.0%

Mathematical problem solving.

Version: Not reported · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Surge public leaderboard (republished result)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
Not reported
Fallback
Not reported
Provider
Not reported
Checked
Sep 30, 2026

Results republished from Surge’s public leaderboard.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

GraphWalks · up to 128k BFS F1 · 99.7%

Breadth-first graph traversal with contexts up to 128k tokens, scored by F1.

Version: Not reported · BFS F1 · percent · Higher is better · Task subset: · up to 128k BFS F1

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Google DeepMind (source-computed)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
Not reported
Fallback
Not reported
Provider
Gemini API
Checked
Sep 30, 2026

Google computed every model on 650 problems with context lengths up to 128k tokens.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

GraphWalks · 256k–1M BFS F1 · 84.2%

Breadth-first graph traversal with 256k to 1M-token contexts, scored by F1.

Version: Not reported · BFS F1 · percent · Higher is better · Task subset: · 256k–1M BFS F1

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Google DeepMind (source-computed)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
Not reported
Fallback
Not reported
Provider
Gemini API
Checked
Sep 30, 2026

Google computed every model on 200 problems with context lengths from 256k to 1M tokens.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

Chartography · 71.6%

Visual interpretation and analysis of charts.

Version: Not reported · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Surge public leaderboard (republished result)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
No tools
Fallback
Not reported
Provider
Not reported
Checked
Sep 30, 2026

Without tools; results republished from Surge’s public leaderboard.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

LVBench · 91.7%

Understanding information in long videos.

Version: Not reported · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
Google DeepMind (source-computed)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
No tools
Fallback
Not reported
Provider
Gemini API
Checked
Sep 30, 2026

Google computed every model without tools. Video sampling differs due to API limits: Argon 1 FPS, Astra 800 frames, Fable 300 frames, Opus 600 frames.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

CWE-bench v1 · 68.0%

Software vulnerability remediation.

Version: v1 · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified

Google DeepMind · Sep 30, 2026 · Developer reported

Reporting source
Google DeepMind · Gemini 4 Argon launch
Evaluation origin
CWE-bench public leaderboard (republished result)
Effort
Highest Gemini thinking setting unless otherwise noted
Harness
Not reported
Tools
Not reported
Fallback
Not reported
Provider
Not reported
Checked
Sep 30, 2026

Pass@1 results republished from the official leaderboard. Its pass@4 tiebreak is not part of this sheet, so equal pass@1 scores remain ties.

No confidence interval reported in this comparison.

Original methodology ↗ · Original evidence ↗

03 / Where you can use itPublished access

No catalogued plan or API offers this model yet. Access is not inferred from another release.

Pricing, assumptions & evidence

Missing token-category prices are not zero. Workload pricing applies exact recorded categories and admitted routes. No benchmark score or quality ranking is inferred from prices.

Aliases, routes and identity

Catalog ID gemini-4-argon

Model release (kind: release)

Developer: Google

No additional exact aliases recorded.