Skip to content

Benchmark sheet

Model Benchmarks

Exact benchmark versions. Verified reported scores. Each evaluation keeps its source and setup.

Add or change models Up to six

Scores verified against original reports, not independently reproduced. Select a score for its setup and evidence. Shared benchmarks appear first. Highlighted: highest reported score, or lowest where lower is better. Setups may differ.

Selected model benchmark evidence. Different evaluation setups are identified per row. Missing scores are not reported.
BenchmarkGemini 4 ArgonGoogleGPT-6 AstraOpenAIGPT-6.1 SolOpenAIClaude Opus 5.5AnthropicClaude Fable 5.1Anthropic
CodingDifferent or unreported setups
ScienceDifferent or unreported setups
Knowledge workDifferent or unreported setupsNot reported
Knowledge workDifferent or unreported setupsNot reported
Knowledge workDifferent or unreported setupsNot reported
Knowledge workDifferent or unreported setupsNot reported
CodingDifferent or unreported setupsNot reported
CodingDifferent or unreported setupsNot reported
CodingDifferent or unreported setupsNot reported
CodingDifferent or unreported setupsNot reported
ScienceDifferent or unreported setupsNot reported
ScienceDifferent or unreported setupsNot reported
Long contextDifferent or unreported setupsNot reported
Long contextDifferent or unreported setupsNot reported
MultimodalDifferent or unreported setupsNot reported
MultimodalDifferent or unreported setupsNot reported
SecurityDifferent or unreported setupsNot reported
Knowledge workDifferent or unreported setupsNot reportedNot reported
Knowledge workDifferent or unreported setupsNot reported
Knowledge workDifferent or unreported setupsNot reportedNot reportedNot reported
Knowledge workDifferent or unreported setupsNot reportedNot reportedNot reported

Methodology & sources

2026-09-30-v2

Benchmark versions, metrics and task subsets remain distinct. Different efforts, tools, harnesses, fallbacks or deployments can produce different results. We do not normalize scores or calculate a composite rating. Highlighting identifies each row’s highest reported value, or lowest for lower-is-better metrics, including every tie. It does not establish matching evaluation setups or an overall model ranking.

How results are selected and verified

We check each numerical result against the original publication and retain its date and configuration. A source marker identifies the reporting organization, which can differ from the model developer and the original evaluator. Google’s reviewed launch snapshot stays primary where it has a result. OpenAI’s launch defaults use Max effort consistently, including benchmarks where High scores higher. Otherwise we use the sole verified provider observation. Selection never maximizes a score. Alternative observations are available by selecting a cell and are pinned in your comparison link.

Shared benchmarks requires a reported result for every selected model. All reported results shows explicit coverage gaps. Complete source sheets require every declared model × benchmark cell. Changes to benchmark evidence have no effect on Replay, pricing, plan capacity or model identity.

1 Google DeepMind · Gemini 4 Argon launch · Sep 30, 2026

Developer reported · Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1, Claude Opus 5.5

Google’s launch sheet combines evaluations it computed with results it republished from leaderboards and system cards. This is one reporting set; it does not establish identical harnesses across models. Argon generally uses the Gemini API at the highest thinking setting, with a single attempt and trial averaging on smaller benchmarks.

  • Competitor scores are reported by Google in this sheet. The model developer label identifies authorship, not the evaluator of each result.
  • Thinking settings, harnesses and task subsets can differ. Per-benchmark details below preserve the differences Google describes; unreported settings remain unknown.
  • Agent’s Last Exam and OSWorld-2.0 are excluded because the source does not report all four models.
  • Versions and GraphWalks context subsets stay distinct. Where Google specifies no version, the version remains unreported.
  • Google’s methodology PDF says September 2026 in its approach and October 2026 in its results footer. Publication here follows the dated September 30 launch; the footer is not treated as a separate evaluation date.

Original methodology ↗ · Original source ↗

Manually transcribed numerical facts from Google’s official model launch comparison, with original local descriptions and attributed methodology summaries. No third-party dataset or leaderboard was downloaded or redistributed; Google’s citations do not license future direct ingestion of those sources. Checked Sep 30, 2026.

2 DeepSeek · V4.1 Flash release · Sep 10, 2026

Developer reported · DeepSeek-V4.1-Flash

DeepSeek’s September 10 changelog reports these exact benchmark versions and scores. Configuration notes from older releases are not assumed to apply to V4.1 Flash.

  • These observations are from a separate provider publication, not Google’s Argon evaluation.
  • Matching benchmark names and versions do not establish identical evaluation conditions.
  • Only the selected numerical facts verified against the original provider publication are stored.

Original methodology ↗ · Original source ↗

Provider-published numerical facts, transcribed from the official announcement. No third-party dataset or chart artwork is redistributed. Checked Sep 30, 2026.

3 SpaceXAI · Grok 4.7 launch · Sep 21, 2026

Developer reported · Grok 4.7

SpaceXAI’s launch comparison reports Grok 4.7 at xHigh, except DeepSWE v1.1 at High. No configurations are borrowed from other providers.

  • These observations are from a separate provider publication, not Google’s Argon evaluation.
  • Matching benchmark names and versions do not establish identical evaluation conditions.
  • Only the selected numerical facts verified against the original provider publication are stored.

Original methodology ↗ · Original source ↗

Provider-published numerical facts, transcribed from the official announcement. No third-party dataset or chart artwork is redistributed. Checked Sep 30, 2026.

4 Anthropic · Sonnet 5.5 launch · Sep 28, 2026

Developer reported · Claude Sonnet 5.5, Claude Opus 5.5

Anthropic’s launch table reports Terminal-Bench 4.0 and Chartography without tools. The Opus terminal result uses Xhigh; the source calls it that model’s highest reported effort result. It remains an alternative observation rather than replacing Google’s result.

  • These observations are from a separate provider publication, not Google’s Argon evaluation.
  • Matching benchmark names and versions do not establish identical evaluation conditions.
  • Only the selected numerical facts verified against the original provider publication are stored.

Original methodology ↗ · Original source ↗

Provider-published numerical facts, transcribed from the official announcement. No third-party dataset or chart artwork is redistributed. Checked Sep 30, 2026.

5 OpenAI · GPT-6.1 Sol launch · Low effort · Sep 29, 2026

Developer reported · GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, Claude Opus 5.5

Launch chart results at Low reasoning effort. GPT evaluations ran in OpenAI research environments or its API. Competitor scores were taken from public reports. Exact chart values, versions, subsets and documented fallbacks are preserved.

  • One publisher does not prove matching harnesses, tools or deployments. Numeric highlighting does not establish matching evaluation setups.
  • AutomationBench 1.0.6 stays separate from the unversioned Google result; identical task subsets are not established.
  • OSWorld reports partial reward on the offline v2026.08.08 release, separate from other OSWorld versions and subsets.
  • Factuality uses difficult, previously flagged conversations and does not measure error rates in typical usage.
  • The Fable AutomationBench point uses Opus 5 fallbacks on about 40% of tasks; reported costs omit fallbacks. Costs are not imported into StackReplay pricing.

Original methodology ↗ · Original source ↗

Selected numerical facts from the official OpenAI launch publication, with attribution and our own descriptions. No chart artwork, copied prose or underlying third-party dataset is redistributed. Checked Sep 30, 2026.

6 OpenAI · GPT-6.1 Sol launch · Medium effort · Sep 29, 2026

Developer reported · GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, Claude Opus 5.5

Launch chart results at Medium reasoning effort. GPT evaluations ran in OpenAI research environments or its API. Competitor scores were taken from public reports. Exact chart values, versions, subsets and documented fallbacks are preserved.

  • One publisher does not prove matching harnesses, tools or deployments. Numeric highlighting does not establish matching evaluation setups.
  • AutomationBench 1.0.6 stays separate from the unversioned Google result; identical task subsets are not established.
  • OSWorld reports partial reward on the offline v2026.08.08 release, separate from other OSWorld versions and subsets.
  • Factuality uses difficult, previously flagged conversations and does not measure error rates in typical usage.
  • The Fable AutomationBench point uses Opus 5 fallbacks on about 40% of tasks; reported costs omit fallbacks. Costs are not imported into StackReplay pricing.

Original methodology ↗ · Original source ↗

Selected numerical facts from the official OpenAI launch publication, with attribution and our own descriptions. No chart artwork, copied prose or underlying third-party dataset is redistributed. Checked Sep 30, 2026.

7 OpenAI · GPT-6.1 Sol launch · High effort · Sep 29, 2026

Developer reported · GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, Claude Opus 5.5

Launch chart results at High reasoning effort. GPT evaluations ran in OpenAI research environments or its API. Competitor scores were taken from public reports. Exact chart values, versions, subsets and documented fallbacks are preserved.

  • One publisher does not prove matching harnesses, tools or deployments. Numeric highlighting does not establish matching evaluation setups.
  • AutomationBench 1.0.6 stays separate from the unversioned Google result; identical task subsets are not established.
  • OSWorld reports partial reward on the offline v2026.08.08 release, separate from other OSWorld versions and subsets.
  • Factuality uses difficult, previously flagged conversations and does not measure error rates in typical usage.
  • The Fable AutomationBench point uses Opus 5 fallbacks on about 40% of tasks; reported costs omit fallbacks. Costs are not imported into StackReplay pricing.

Original methodology ↗ · Original source ↗

Selected numerical facts from the official OpenAI launch publication, with attribution and our own descriptions. No chart artwork, copied prose or underlying third-party dataset is redistributed. Checked Sep 30, 2026.

8 OpenAI · GPT-6.1 Sol launch · Xhigh effort · Sep 29, 2026

Developer reported · GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, Claude Opus 5.5

Launch chart results at Xhigh reasoning effort. GPT evaluations ran in OpenAI research environments or its API. Competitor scores were taken from public reports. Exact chart values, versions, subsets and documented fallbacks are preserved.

  • One publisher does not prove matching harnesses, tools or deployments. Numeric highlighting does not establish matching evaluation setups.
  • AutomationBench 1.0.6 stays separate from the unversioned Google result; identical task subsets are not established.
  • OSWorld reports partial reward on the offline v2026.08.08 release, separate from other OSWorld versions and subsets.
  • Factuality uses difficult, previously flagged conversations and does not measure error rates in typical usage.
  • The Fable AutomationBench point uses Opus 5 fallbacks on about 40% of tasks; reported costs omit fallbacks. Costs are not imported into StackReplay pricing.

Original methodology ↗ · Original source ↗

Selected numerical facts from the official OpenAI launch publication, with attribution and our own descriptions. No chart artwork, copied prose or underlying third-party dataset is redistributed. Checked Sep 30, 2026.

9 OpenAI · GPT-6.1 Sol launch · Max effort · Sep 29, 2026

Developer reported · GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, Claude Opus 5.5, Claude Fable 5.1

Launch chart results at Max reasoning effort. GPT evaluations ran in OpenAI research environments or its API. Competitor scores were taken from public reports. Exact chart values, versions, subsets and documented fallbacks are preserved.

  • One publisher does not prove matching harnesses, tools or deployments. Numeric highlighting does not establish matching evaluation setups.
  • AutomationBench 1.0.6 stays separate from the unversioned Google result; identical task subsets are not established.
  • OSWorld reports partial reward on the offline v2026.08.08 release, separate from other OSWorld versions and subsets.
  • Factuality uses difficult, previously flagged conversations and does not measure error rates in typical usage.
  • The Fable AutomationBench point uses Opus 5 fallbacks on about 40% of tasks; reported costs omit fallbacks. Costs are not imported into StackReplay pricing.

Original methodology ↗ · Original source ↗

Selected numerical facts from the official OpenAI launch publication, with attribution and our own descriptions. No chart artwork, copied prose or underlying third-party dataset is redistributed. Checked Sep 30, 2026.