- Max output
- 1M
Announced future API pricing: introductory $2 / 1M input and $10 / 1M output, with cached input at 95% off input. After the introductory period: $4 / 1M input and $20 / 1M output. API effective date and introductory expiration date are unknown, so these are not executable rates.
- Maximum output
- 1,000,000
Initially rolling out to trusted cyber defenders. Broader availability is announced for later, starting with paid API customers and Google AI Ultra subscribers; current broad API or Ultra access is not established.
The launch announcement establishes a 1,000,000-token output limit, not a context window. An exact public API model identifier is not established.
Benchmarks
Open benchmark sheet →Google DeepMind · Sep 30, 2026
Developer reported · Gemini 4 Argon launch
Coding
- DeepSWE v1.1
- 77.9%
- FrontierSWE v2
- 55.0%
- Vibe Code Bench
- 91.9%
- Terminal-Bench 4.0
- 57.4%
- PostTrainBench v1.1
- 45.3%
Knowledge work
- Vals Index
- 68.9%
- AutomationBench
- 51.3%
Show all 17 results
Knowledge work
Science
- LABBench 2
- 88.8%
- RiemannBench
- 76.0%
Long context
Multimodal
- Chartography
- 71.6%
- LVBench
- 91.7%
Security
- CWE-bench v1
- 68.0%
Methodology & sources
Google’s launch sheet combines evaluations it computed with results it republished from leaderboards and system cards. This is one reporting set; it does not establish identical harnesses across models. Argon generally uses the Gemini API at the highest thinking setting, with a single attempt and trial averaging on smaller benchmarks.
These scores are reported by Google DeepMind. Model developers identify authorship; results may originate with another evaluator. Exact benchmark versions remain distinct. StackReplay calculates no composite score.
DeepSWE v1.1 · 77.9%
Long-horizon software engineering in real-world codebases.
Version: v1.1 · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Google DeepMind (source-computed)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- mini-swe
- Tools
- Not reported
- Fallback
- Not reported
- Provider
- Gemini API
- Checked
- Sep 30, 2026
Argon uses Google’s mini-swe agent; Astra is reported from the leaderboard, Fable and Opus from system cards. Thinking levels follow the highest scoring level reported by Datacurve.
No confidence interval reported in this comparison.
FrontierSWE v2 · 55.0%
Agentic software engineering on coding tasks.
Version: v2 · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Proximal public leaderboard (republished result)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- Not reported
- Fallback
- Not reported
- Provider
- Not reported
- Checked
- Sep 30, 2026
Results republished from Proximal’s public leaderboard.
No confidence interval reported in this comparison.
Vibe Code Bench · 91.9%
Code generation for prompted software-building tasks.
Version: Not reported · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Vals AI public leaderboard (republished result)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- Not reported
- Fallback
- Not reported
- Provider
- Not reported
- Checked
- Sep 30, 2026
Results republished from the Vals AI public leaderboard.
No confidence interval reported in this comparison.
Terminal-Bench 4.0 · 57.4%
Agentic task completion in a terminal environment.
Version: 4.0 · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Google DeepMind (source-computed)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- Not reported
- Fallback
- Not reported
- Provider
- Gemini API
- Checked
- Sep 30, 2026
Argon is computed by Google; competitor results are republished from the official leaderboard, using the highest scoring thinking level reported by its authors.
No confidence interval reported in this comparison.
PostTrainBench v1.1 · 45.3%
Post-training of language models under a fixed compute budget.
Version: v1.1 · Benchmark-reported weighted score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Google DeepMind (source-computed)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- OpenCode
- Tools
- Single NVIDIA H100, 10-hour budget
- Fallback
- Not reported
- Provider
- Gemini API
- Checked
- Sep 30, 2026
Google computed every model with OpenCode and a 10-hour budget on one NVIDIA H100; the benchmark’s metric aggregates four base models and seven benchmarks.
No confidence interval reported in this comparison.
Vals Index · 68.9%
Economic knowledge work across finance, coding, legal and tax tasks.
Version: Not reported · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Vals AI (republished result)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- Not reported
- Fallback
- Not reported
- Provider
- Not reported
- Checked
- Sep 30, 2026
Vals AI results republished by Google.
No confidence interval reported in this comparison.
AutomationBench · 51.3%
End-to-end execution of business workflows.
Version: Not reported · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Zapier public leaderboard (republished result)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- Not reported
- Fallback
- Not reported
- Provider
- Not reported
- Checked
- Sep 30, 2026
Private task set; results republished from Zapier’s public leaderboard.
No confidence interval reported in this comparison.
Vals Finance Agent v2 · 65.4%
Multi-step financial research by an agent.
Version: v2 · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Vals AI (republished result)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- Not reported
- Fallback
- Not reported
- Provider
- Not reported
- Checked
- Sep 30, 2026
Vals AI results republished by Google.
No confidence interval reported in this comparison.
Harvey’s Legal Agent Benchmark · 19.6%
Legal research and drafting by an agent.
Version: Not reported · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Vals AI (republished result)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- Not reported
- Fallback
- Not reported
- Provider
- Not reported
- Checked
- Sep 30, 2026
Vals AI results republished by Google.
No confidence interval reported in this comparison.
Terminal-Bench Science 0.1 · 57.6%
Scientific and mathematical tasks in a terminal environment.
Version: 0.1 · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Google DeepMind (source-computed)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- Not reported
- Fallback
- Not reported
- Provider
- Gemini API
- Checked
- Sep 30, 2026
Argon uses a sixfold verifier timeout to address verification timeouts; competitor results come from the official leaderboard.
No confidence interval reported in this comparison.
LABBench 2 · 88.8%
Scientific tasks with a terminal, bioinformatics tools and internet access.
Version: 2 · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Google DeepMind (source-computed)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- Linux terminal, bioinformatics tools, Python, R and internet
- Fallback
- Not reported
- Provider
- Gemini API
- Checked
- Sep 30, 2026
Google computed every model with a Linux terminal, bioinformatics tools, Python, R and internet access.
No confidence interval reported in this comparison.
RiemannBench · 76.0%
Mathematical problem solving.
Version: Not reported · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Surge public leaderboard (republished result)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- Not reported
- Fallback
- Not reported
- Provider
- Not reported
- Checked
- Sep 30, 2026
Results republished from Surge’s public leaderboard.
No confidence interval reported in this comparison.
GraphWalks · up to 128k BFS F1 · 99.7%
Breadth-first graph traversal with contexts up to 128k tokens, scored by F1.
Version: Not reported · BFS F1 · percent · Higher is better · Task subset: · up to 128k BFS F1
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Google DeepMind (source-computed)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- Not reported
- Fallback
- Not reported
- Provider
- Gemini API
- Checked
- Sep 30, 2026
Google computed every model on 650 problems with context lengths up to 128k tokens.
No confidence interval reported in this comparison.
GraphWalks · 256k–1M BFS F1 · 84.2%
Breadth-first graph traversal with 256k to 1M-token contexts, scored by F1.
Version: Not reported · BFS F1 · percent · Higher is better · Task subset: · 256k–1M BFS F1
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Google DeepMind (source-computed)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- Not reported
- Fallback
- Not reported
- Provider
- Gemini API
- Checked
- Sep 30, 2026
Google computed every model on 200 problems with context lengths from 256k to 1M tokens.
No confidence interval reported in this comparison.
Chartography · 71.6%
Visual interpretation and analysis of charts.
Version: Not reported · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Surge public leaderboard (republished result)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- No tools
- Fallback
- Not reported
- Provider
- Not reported
- Checked
- Sep 30, 2026
Without tools; results republished from Surge’s public leaderboard.
No confidence interval reported in this comparison.
LVBench · 91.7%
Understanding information in long videos.
Version: Not reported · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- Google DeepMind (source-computed)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- No tools
- Fallback
- Not reported
- Provider
- Gemini API
- Checked
- Sep 30, 2026
Google computed every model without tools. Video sampling differs due to API limits: Argon 1 FPS, Astra 800 frames, Fable 300 frames, Opus 600 frames.
No confidence interval reported in this comparison.
CWE-bench v1 · 68.0%
Software vulnerability remediation.
Version: v1 · Reported benchmark score · percent · Higher is better · Task subset: Not separately specified
Google DeepMind · Sep 30, 2026 · Developer reported
- Reporting source
- Google DeepMind · Gemini 4 Argon launch
- Evaluation origin
- CWE-bench public leaderboard (republished result)
- Effort
- Highest Gemini thinking setting unless otherwise noted
- Harness
- Not reported
- Tools
- Not reported
- Fallback
- Not reported
- Provider
- Not reported
- Checked
- Sep 30, 2026
Pass@1 results republished from the official leaderboard. Its pass@4 tiebreak is not part of this sheet, so equal pass@1 scores remain ties.
No confidence interval reported in this comparison.
No catalogued plan or API offers this model yet. Access is not inferred from another release.
Pricing, assumptions & evidence
Missing token-category prices are not zero. Workload pricing applies exact recorded categories and admitted routes. No benchmark score or quality ranking is inferred from prices.
Aliases, routes and identity
Catalog ID gemini-4-argon
Model release (kind: release)
Developer: Google
No additional exact aliases recorded.

