Benchmark sheet
Model Benchmarks
Exact benchmark versions. Verified reported scores. Each evaluation keeps its source and setup.
Add or change models Up to six
Scores verified against original reports, not independently reproduced. Select a score for its setup and evidence. Shared benchmarks appear first. Highlighted: highest reported score, or lowest where lower is better. Setups may differ.
| Benchmark | Gemini 4 ArgonGoogle |
|---|---|
| Knowledge workDifferent or unreported setups | |
| Knowledge workDifferent or unreported setups | |
| Knowledge workDifferent or unreported setups | |
| Knowledge workDifferent or unreported setups | |
| CodingDifferent or unreported setups | |
| CodingDifferent or unreported setups | |
| CodingDifferent or unreported setups | |
| CodingDifferent or unreported setups | |
| CodingDifferent or unreported setups | |
| ScienceDifferent or unreported setups | |
| ScienceDifferent or unreported setups | |
| ScienceDifferent or unreported setups | |
| Long contextDifferent or unreported setups | |
| Long contextDifferent or unreported setups | |
| MultimodalDifferent or unreported setups | |
| MultimodalDifferent or unreported setups | |
| SecurityDifferent or unreported setups |
Methodology & sources
2026-09-30-v1Benchmark versions, metrics and task subsets remain distinct. Different efforts, tools, harnesses, fallbacks or deployments can produce different results. We do not normalize scores or calculate a composite rating. Highlighting identifies each row’s highest reported value, or lowest for lower-is-better metrics, including every tie. It does not establish matching evaluation setups or an overall model ranking.
How results are selected and verified
We check each numerical result against the original publication and retain its date and configuration. A source marker identifies the reporting organization, which can differ from the model developer and the original evaluator. Google’s reviewed launch snapshot stays primary where it has a result. OpenAI’s launch defaults use Max effort consistently, including benchmarks where High scores higher. Otherwise we use the sole verified provider observation. Selection never maximizes a score. Alternative observations are available by selecting a cell and are pinned in your comparison link.
Shared benchmarks requires a reported result for every selected model. All reported results shows explicit coverage gaps. Complete source sheets require every declared model × benchmark cell. Changes to benchmark evidence have no effect on Replay, pricing, plan capacity or model identity.
1 Google DeepMind · Gemini 4 Argon launch · Sep 30, 2026
Developer reported · Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1, Claude Opus 5.5
Google’s launch sheet combines evaluations it computed with results it republished from leaderboards and system cards. This is one reporting set; it does not establish identical harnesses across models. Argon generally uses the Gemini API at the highest thinking setting, with a single attempt and trial averaging on smaller benchmarks.
- Competitor scores are reported by Google in this sheet. The model developer label identifies authorship, not the evaluator of each result.
- Thinking settings, harnesses and task subsets can differ. Per-benchmark details below preserve the differences Google describes; unreported settings remain unknown.
- Agent’s Last Exam and OSWorld-2.0 are excluded because the source does not report all four models.
- Versions and GraphWalks context subsets stay distinct. Where Google specifies no version, the version remains unreported.
- Google’s methodology PDF says September 2026 in its approach and October 2026 in its results footer. Publication here follows the dated September 30 launch; the footer is not treated as a separate evaluation date.
Original methodology ↗ · Original source ↗
Manually transcribed numerical facts from Google’s official model launch comparison, with original local descriptions and attributed methodology summaries. No third-party dataset or leaderboard was downloaded or redistributed; Google’s citations do not license future direct ingestion of those sources. Checked Sep 30, 2026.
2 DeepSeek · V4.1 Flash release · Sep 10, 2026
Developer reported · DeepSeek-V4.1-Flash
DeepSeek’s September 10 changelog reports these exact benchmark versions and scores. Configuration notes from older releases are not assumed to apply to V4.1 Flash.
- These observations are from a separate provider publication, not Google’s Argon evaluation.
- Matching benchmark names and versions do not establish identical evaluation conditions.
- Only the selected numerical facts verified against the original provider publication are stored.
Original methodology ↗ · Original source ↗
Provider-published numerical facts, transcribed from the official announcement. No third-party dataset or chart artwork is redistributed. Checked Sep 30, 2026.
3 SpaceXAI · Grok 4.7 launch · Sep 21, 2026
Developer reported · Grok 4.7
SpaceXAI’s launch comparison reports Grok 4.7 at xHigh, except DeepSWE v1.1 at High. No configurations are borrowed from other providers.
- These observations are from a separate provider publication, not Google’s Argon evaluation.
- Matching benchmark names and versions do not establish identical evaluation conditions.
- Only the selected numerical facts verified against the original provider publication are stored.
Original methodology ↗ · Original source ↗
Provider-published numerical facts, transcribed from the official announcement. No third-party dataset or chart artwork is redistributed. Checked Sep 30, 2026.
4 Anthropic · Sonnet 5.5 launch · Sep 28, 2026
Developer reported · Claude Sonnet 5.5, Claude Opus 5.5
Anthropic’s launch table reports Terminal-Bench 4.0 and Chartography without tools. The Opus terminal result uses Xhigh; the source calls it that model’s highest reported effort result. It remains an alternative observation rather than replacing Google’s result.
- These observations are from a separate provider publication, not Google’s Argon evaluation.
- Matching benchmark names and versions do not establish identical evaluation conditions.
- Only the selected numerical facts verified against the original provider publication are stored.
Original methodology ↗ · Original source ↗
Provider-published numerical facts, transcribed from the official announcement. No third-party dataset or chart artwork is redistributed. Checked Sep 30, 2026.

