Skip to content

Benchmark sheet

Model Benchmarks

Exact benchmark versions. Verified reported scores. Each evaluation keeps its source and setup.

Add or change models Up to six

Scores verified against original reports, not independently reproduced. Select a score for its setup and evidence. Shared benchmarks appear first. Highlighted: highest reported score, or lowest where lower is better. Setups may differ.

Selected model benchmark evidence. Different evaluation setups are identified per row. Missing scores are not reported.
BenchmarkGemini 4 ArgonGoogle
Knowledge workDifferent or unreported setups
Knowledge workDifferent or unreported setups
Knowledge workDifferent or unreported setups
Knowledge workDifferent or unreported setups
CodingDifferent or unreported setups
CodingDifferent or unreported setups
CodingDifferent or unreported setups
CodingDifferent or unreported setups
CodingDifferent or unreported setups
ScienceDifferent or unreported setups
ScienceDifferent or unreported setups
ScienceDifferent or unreported setups
Long contextDifferent or unreported setups
Long contextDifferent or unreported setups
MultimodalDifferent or unreported setups
MultimodalDifferent or unreported setups
SecurityDifferent or unreported setups

Methodology & sources

2026-09-30-v1

Benchmark versions, metrics and task subsets remain distinct. Different efforts, tools, harnesses, fallbacks or deployments can produce different results. We do not normalize scores or calculate a composite rating. Highlighting identifies each row’s highest reported value, or lowest for lower-is-better metrics, including every tie. It does not establish matching evaluation setups or an overall model ranking.

How results are selected and verified

We check each numerical result against the original publication and retain its date and configuration. A source marker identifies the reporting organization, which can differ from the model developer and the original evaluator. Google’s reviewed launch snapshot stays primary where it has a result. OpenAI’s launch defaults use Max effort consistently, including benchmarks where High scores higher. Otherwise we use the sole verified provider observation. Selection never maximizes a score. Alternative observations are available by selecting a cell and are pinned in your comparison link.

Shared benchmarks requires a reported result for every selected model. All reported results shows explicit coverage gaps. Complete source sheets require every declared model × benchmark cell. Changes to benchmark evidence have no effect on Replay, pricing, plan capacity or model identity.

1 Google DeepMind · Gemini 4 Argon launch · Sep 30, 2026

Developer reported · Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1, Claude Opus 5.5

Google’s launch sheet combines evaluations it computed with results it republished from leaderboards and system cards. This is one reporting set; it does not establish identical harnesses across models. Argon generally uses the Gemini API at the highest thinking setting, with a single attempt and trial averaging on smaller benchmarks.

  • Competitor scores are reported by Google in this sheet. The model developer label identifies authorship, not the evaluator of each result.
  • Thinking settings, harnesses and task subsets can differ. Per-benchmark details below preserve the differences Google describes; unreported settings remain unknown.
  • Agent’s Last Exam and OSWorld-2.0 are excluded because the source does not report all four models.
  • Versions and GraphWalks context subsets stay distinct. Where Google specifies no version, the version remains unreported.
  • Google’s methodology PDF says September 2026 in its approach and October 2026 in its results footer. Publication here follows the dated September 30 launch; the footer is not treated as a separate evaluation date.

Original methodology ↗ · Original source ↗

Manually transcribed numerical facts from Google’s official model launch comparison, with original local descriptions and attributed methodology summaries. No third-party dataset or leaderboard was downloaded or redistributed; Google’s citations do not license future direct ingestion of those sources. Checked Sep 30, 2026.

2 DeepSeek · V4.1 Flash release · Sep 10, 2026

Developer reported · DeepSeek-V4.1-Flash

DeepSeek’s September 10 changelog reports these exact benchmark versions and scores. Configuration notes from older releases are not assumed to apply to V4.1 Flash.

  • These observations are from a separate provider publication, not Google’s Argon evaluation.
  • Matching benchmark names and versions do not establish identical evaluation conditions.
  • Only the selected numerical facts verified against the original provider publication are stored.

Original methodology ↗ · Original source ↗

Provider-published numerical facts, transcribed from the official announcement. No third-party dataset or chart artwork is redistributed. Checked Sep 30, 2026.

3 SpaceXAI · Grok 4.7 launch · Sep 21, 2026

Developer reported · Grok 4.7

SpaceXAI’s launch comparison reports Grok 4.7 at xHigh, except DeepSWE v1.1 at High. No configurations are borrowed from other providers.

  • These observations are from a separate provider publication, not Google’s Argon evaluation.
  • Matching benchmark names and versions do not establish identical evaluation conditions.
  • Only the selected numerical facts verified against the original provider publication are stored.

Original methodology ↗ · Original source ↗

Provider-published numerical facts, transcribed from the official announcement. No third-party dataset or chart artwork is redistributed. Checked Sep 30, 2026.

4 Anthropic · Sonnet 5.5 launch · Sep 28, 2026

Developer reported · Claude Sonnet 5.5, Claude Opus 5.5

Anthropic’s launch table reports Terminal-Bench 4.0 and Chartography without tools. The Opus terminal result uses Xhigh; the source calls it that model’s highest reported effort result. It remains an alternative observation rather than replacing Google’s result.

  • These observations are from a separate provider publication, not Google’s Argon evaluation.
  • Matching benchmark names and versions do not establish identical evaluation conditions.
  • Only the selected numerical facts verified against the original provider publication are stored.

Original methodology ↗ · Original source ↗

Provider-published numerical facts, transcribed from the official announcement. No third-party dataset or chart artwork is redistributed. Checked Sep 30, 2026.