Benchmark sheet
Model Benchmarks
Exact benchmark versions. Verified reported scores. Each evaluation keeps its source and setup.
Developer reported · Checked against original publications; not reproduced by StackReplay.
Add or change models Up to six
Scores verified against original reports, not independently reproduced. Select a score for its setup and evidence. Shared benchmarks appear first. Highlighted: highest reported score, or lowest where lower is better. Setups may differ.
| Benchmark | Grok 4.7xAI |
|---|---|
| CodingDifferent or unreported setups | |
| CodingDifferent or unreported setups | |
| CodingDifferent or unreported setups | |
| ScienceDifferent or unreported setups | |
| ScienceDifferent or unreported setups |
Methodology & sources
2026-10-04-v5Benchmark versions, metrics and task subsets remain distinct. Different efforts, tools, harnesses, fallbacks or deployments can produce different results. We do not normalize scores or calculate a composite rating. Highlighting identifies each row’s highest reported value, or lowest for lower-is-better metrics, including every tie. It does not establish matching evaluation setups or an overall model ranking.
How results are selected and verified
We check each numerical result against the original publication and retain its date and configuration. A source marker identifies the reporting organization, which can differ from the model developer and the original evaluator. Google’s reviewed launch snapshot stays primary where it has a result. OpenAI’s launch defaults use Max effort consistently, including benchmarks where High scores higher. Otherwise we use the sole verified provider observation. Selection never maximizes a score. Alternative observations are available by selecting a cell and are pinned in your comparison link.
Shared benchmarks requires a reported result for every selected model. All reported results shows explicit coverage gaps. Complete source sheets require every declared model × benchmark cell. Changes to benchmark evidence have no effect on Replay, pricing, plan capacity or model identity.
1 Google DeepMind · Gemini 4 Argon launch · Published Sep 30, 2026
Developer reported · Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1, Claude Opus 5.5
Google’s launch sheet combines evaluations it computed with results it republished from leaderboards and system cards. This is one reporting set; it does not establish identical harnesses across models. Argon generally uses the Gemini API at the highest thinking setting, with a single attempt and trial averaging on smaller benchmarks.
- Competitor scores are reported by Google in this sheet. The model developer label identifies authorship, not the evaluator of each result.
- Thinking settings, harnesses and task subsets can differ. Per-benchmark details below preserve the differences Google describes; unreported settings remain unknown.
- Agent’s Last Exam and OSWorld-2.0 are excluded because the source does not report all four models.
- Versions and GraphWalks context subsets stay distinct. Where Google specifies no version, the version remains unreported.
- Google’s methodology PDF says September 2026 in its approach and October 2026 in its results footer. Publication here follows the dated September 30 launch; the footer is not treated as a separate evaluation date.
Original methodology ↗ · Original source ↗
Google DeepMind · Gemini 4 Argon launch. Manually transcribed numerical facts from Google’s official model launch comparison, with original local descriptions and attributed methodology summaries. No third-party dataset or leaderboard was downloaded or redistributed; Google’s citations do not license future direct ingestion of those sources. Checked Sep 30, 2026. Redistribution terms ↗
2 DeepSeek · V4.1 Flash release · Published Sep 10, 2026
Developer reported · DeepSeek-V4.1-Flash
DeepSeek’s September 10 changelog reports these exact benchmark versions and scores. Configuration notes from older releases are not assumed to apply to V4.1 Flash.
- These observations are from a separate provider publication, not Google’s Argon evaluation.
- Matching benchmark names and versions do not establish identical evaluation conditions.
- Only the selected numerical facts verified against the original provider publication are stored.
Original methodology ↗ · Original source ↗
DeepSeek · V4.1 Flash release. Provider-published numerical facts, transcribed from the official announcement. No third-party dataset or chart artwork is redistributed. Checked Sep 30, 2026. Redistribution terms ↗
3 SpaceXAI · Grok 4.7 launch · Published Sep 21, 2026
Developer reported · Grok 4.7
SpaceXAI’s launch comparison reports Grok 4.7 at xHigh, except DeepSWE v1.1 at High. No configurations are borrowed from other providers.
- These observations are from a separate provider publication, not Google’s Argon evaluation.
- Matching benchmark names and versions do not establish identical evaluation conditions.
- Only the selected numerical facts verified against the original provider publication are stored.
Original methodology ↗ · Original source ↗
SpaceXAI · Grok 4.7 launch. Provider-published numerical facts, transcribed from the official announcement. No third-party dataset or chart artwork is redistributed. Checked Sep 30, 2026. Redistribution terms ↗
4 Anthropic · Sonnet 5.5 launch · Published Sep 28, 2026
Developer reported · Claude Sonnet 5.5, Claude Opus 5.5
Anthropic’s launch table reports Terminal-Bench 4.0 and Chartography without tools. The Opus terminal result uses Xhigh; the source calls it that model’s highest reported effort result. It remains an alternative observation rather than replacing Google’s result.
- These observations are from a separate provider publication, not Google’s Argon evaluation.
- Matching benchmark names and versions do not establish identical evaluation conditions.
- Only the selected numerical facts verified against the original provider publication are stored.
Original methodology ↗ · Original source ↗
Anthropic · Sonnet 5.5 launch. Provider-published numerical facts, transcribed from the official announcement. No third-party dataset or chart artwork is redistributed. Checked Sep 30, 2026. Redistribution terms ↗
5 OpenAI · GPT-6.1 Sol launch · Low effort · Published Sep 29, 2026
Developer reported · GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, Claude Opus 5.5
Launch chart results at Low reasoning effort. GPT evaluations ran in OpenAI research environments or its API. Competitor scores were taken from public reports. Exact chart values, versions, subsets and documented fallbacks are preserved.
- One publisher does not prove matching harnesses, tools or deployments. Numeric highlighting does not establish matching evaluation setups.
- AutomationBench 1.0.6 stays separate from the unversioned Google result; identical task subsets are not established.
- OSWorld reports partial reward on the offline v2026.08.08 release, separate from other OSWorld versions and subsets.
- Factuality uses difficult, previously flagged conversations and does not measure error rates in typical usage.
- The Fable AutomationBench point uses Opus 5 fallbacks on about 40% of tasks; reported costs omit fallbacks. Costs are not imported into StackReplay pricing.
Original methodology ↗ · Original source ↗
OpenAI · GPT-6.1 Sol launch · Low effort. Selected numerical facts from the official OpenAI launch publication, with attribution and our own descriptions. No chart artwork, copied prose or underlying third-party dataset is redistributed. Checked Sep 30, 2026. Redistribution terms ↗
6 OpenAI · GPT-6.1 Sol launch · Medium effort · Published Sep 29, 2026
Developer reported · GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, Claude Opus 5.5
Launch chart results at Medium reasoning effort. GPT evaluations ran in OpenAI research environments or its API. Competitor scores were taken from public reports. Exact chart values, versions, subsets and documented fallbacks are preserved.
- One publisher does not prove matching harnesses, tools or deployments. Numeric highlighting does not establish matching evaluation setups.
- AutomationBench 1.0.6 stays separate from the unversioned Google result; identical task subsets are not established.
- OSWorld reports partial reward on the offline v2026.08.08 release, separate from other OSWorld versions and subsets.
- Factuality uses difficult, previously flagged conversations and does not measure error rates in typical usage.
- The Fable AutomationBench point uses Opus 5 fallbacks on about 40% of tasks; reported costs omit fallbacks. Costs are not imported into StackReplay pricing.
Original methodology ↗ · Original source ↗
OpenAI · GPT-6.1 Sol launch · Medium effort. Selected numerical facts from the official OpenAI launch publication, with attribution and our own descriptions. No chart artwork, copied prose or underlying third-party dataset is redistributed. Checked Sep 30, 2026. Redistribution terms ↗
7 OpenAI · GPT-6.1 Sol launch · High effort · Published Sep 29, 2026
Developer reported · GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, Claude Opus 5.5
Launch chart results at High reasoning effort. GPT evaluations ran in OpenAI research environments or its API. Competitor scores were taken from public reports. Exact chart values, versions, subsets and documented fallbacks are preserved.
- One publisher does not prove matching harnesses, tools or deployments. Numeric highlighting does not establish matching evaluation setups.
- AutomationBench 1.0.6 stays separate from the unversioned Google result; identical task subsets are not established.
- OSWorld reports partial reward on the offline v2026.08.08 release, separate from other OSWorld versions and subsets.
- Factuality uses difficult, previously flagged conversations and does not measure error rates in typical usage.
- The Fable AutomationBench point uses Opus 5 fallbacks on about 40% of tasks; reported costs omit fallbacks. Costs are not imported into StackReplay pricing.
Original methodology ↗ · Original source ↗
OpenAI · GPT-6.1 Sol launch · High effort. Selected numerical facts from the official OpenAI launch publication, with attribution and our own descriptions. No chart artwork, copied prose or underlying third-party dataset is redistributed. Checked Sep 30, 2026. Redistribution terms ↗
8 OpenAI · GPT-6.1 Sol launch · Xhigh effort · Published Sep 29, 2026
Developer reported · GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, Claude Opus 5.5
Launch chart results at Xhigh reasoning effort. GPT evaluations ran in OpenAI research environments or its API. Competitor scores were taken from public reports. Exact chart values, versions, subsets and documented fallbacks are preserved.
- One publisher does not prove matching harnesses, tools or deployments. Numeric highlighting does not establish matching evaluation setups.
- AutomationBench 1.0.6 stays separate from the unversioned Google result; identical task subsets are not established.
- OSWorld reports partial reward on the offline v2026.08.08 release, separate from other OSWorld versions and subsets.
- Factuality uses difficult, previously flagged conversations and does not measure error rates in typical usage.
- The Fable AutomationBench point uses Opus 5 fallbacks on about 40% of tasks; reported costs omit fallbacks. Costs are not imported into StackReplay pricing.
Original methodology ↗ · Original source ↗
OpenAI · GPT-6.1 Sol launch · Xhigh effort. Selected numerical facts from the official OpenAI launch publication, with attribution and our own descriptions. No chart artwork, copied prose or underlying third-party dataset is redistributed. Checked Sep 30, 2026. Redistribution terms ↗
9 OpenAI · GPT-6.1 Sol launch · Max effort · Published Sep 29, 2026
Developer reported · GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, Claude Opus 5.5, Claude Fable 5.1
Launch chart results at Max reasoning effort. GPT evaluations ran in OpenAI research environments or its API. Competitor scores were taken from public reports. Exact chart values, versions, subsets and documented fallbacks are preserved.
- One publisher does not prove matching harnesses, tools or deployments. Numeric highlighting does not establish matching evaluation setups.
- AutomationBench 1.0.6 stays separate from the unversioned Google result; identical task subsets are not established.
- OSWorld reports partial reward on the offline v2026.08.08 release, separate from other OSWorld versions and subsets.
- Factuality uses difficult, previously flagged conversations and does not measure error rates in typical usage.
- The Fable AutomationBench point uses Opus 5 fallbacks on about 40% of tasks; reported costs omit fallbacks. Costs are not imported into StackReplay pricing.
Original methodology ↗ · Original source ↗
OpenAI · GPT-6.1 Sol launch · Max effort. Selected numerical facts from the official OpenAI launch publication, with attribution and our own descriptions. No chart artwork, copied prose or underlying third-party dataset is redistributed. Checked Sep 30, 2026. Redistribution terms ↗
10 Epoch AI · Epoch AI Capabilities and benchmarking: GPQA Diamond mean accuracy · Archive checked Oct 4, 2026
Independent evaluation · Claude Sonnet 5.5
Epoch AI ran and reported this selected GPQA Diamond evaluation. Attribution: Epoch AI, 'Capabilities & benchmarking' (https://epoch.ai/benchmarks), original data (https://epoch.ai/data/benchmark_data.zip), CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). StackReplay selected three Epoch-run records from the Diamond subset and converted accuracy fractions to percentages; no endorsement is implied. Current public methodology v1.0.6 is contextual only; run-specific suite and setup details are unknown.
- Epoch AI published these numerical run results; StackReplay did not reproduce the evaluations and makes no reproduction claim.
- The GPQA Diamond suite revision used for this run is unreported. Public methodology version 1.0.6 is current context only and is not asserted for this run.
- The 2026-10-04 date records the archive snapshot that was checked. Original run publication and completion dates are unknown.
- The published Diamond subset contains 198 questions, but the actual scored sample count, repetitions, and total attempts for this run are unreported.
- Scaffold revision, system prompt, temperature, reasoning parameter payload, reasoning-token budget, output/time/cost caps, endpoint, tools, seed, and error handling are unreported.
- The Epoch configuration label does not establish equal reasoning effort or a comparable setup across providers.
- The reported stderr is retained as standard error, not converted into or presented as a 95% confidence interval.
Original methodology ↗ · Original source ↗
Epoch AI · Epoch AI Capabilities and benchmarking: GPQA Diamond mean accuracy. Selected Epoch-administered GPQA Diamond result rows and linked metadata from the official archive. Epoch's archive README links CC BY 4.0 for its data. Epoch's use-this-data terms state that externally sourced data retains its original licensing. This licensed-dataset basis is scoped only to these three selected Epoch-run numerical records and their attribution; it does not license the mixed archive, GPQA tasks, logs, provider prose, artwork, or StackReplay's broader export. Checked Oct 4, 2026. Redistribution terms ↗
11 Epoch AI · Epoch AI Capabilities and benchmarking: GPQA Diamond mean accuracy · Archive checked Oct 4, 2026
Independent evaluation · Claude Opus 5.5
Epoch AI ran and reported this selected GPQA Diamond evaluation. Attribution: Epoch AI, 'Capabilities & benchmarking' (https://epoch.ai/benchmarks), original data (https://epoch.ai/data/benchmark_data.zip), CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). StackReplay selected three Epoch-run records from the Diamond subset and converted accuracy fractions to percentages; no endorsement is implied. Current public methodology v1.0.6 is contextual only; run-specific suite and setup details are unknown.
- Epoch AI published these numerical run results; StackReplay did not reproduce the evaluations and makes no reproduction claim.
- The GPQA Diamond suite revision used for this run is unreported. Public methodology version 1.0.6 is current context only and is not asserted for this run.
- The 2026-10-04 date records the archive snapshot that was checked. Original run publication and completion dates are unknown.
- The published Diamond subset contains 198 questions, but the actual scored sample count, repetitions, and total attempts for this run are unreported.
- Scaffold revision, system prompt, temperature, reasoning parameter payload, reasoning-token budget, output/time/cost caps, endpoint, tools, seed, and error handling are unreported.
- The Epoch configuration label does not establish equal reasoning effort or a comparable setup across providers.
- The reported stderr is retained as standard error, not converted into or presented as a 95% confidence interval.
Original methodology ↗ · Original source ↗
Epoch AI · Epoch AI Capabilities and benchmarking: GPQA Diamond mean accuracy. Selected Epoch-administered GPQA Diamond result rows and linked metadata from the official archive. Epoch's archive README links CC BY 4.0 for its data. Epoch's use-this-data terms state that externally sourced data retains its original licensing. This licensed-dataset basis is scoped only to these three selected Epoch-run numerical records and their attribution; it does not license the mixed archive, GPQA tasks, logs, provider prose, artwork, or StackReplay's broader export. Checked Oct 4, 2026. Redistribution terms ↗
12 Epoch AI · Epoch AI Capabilities and benchmarking: GPQA Diamond mean accuracy · Archive checked Oct 4, 2026
Independent evaluation · Qwen 3.8 Max 0902
Epoch AI ran and reported this selected GPQA Diamond evaluation. Attribution: Epoch AI, 'Capabilities & benchmarking' (https://epoch.ai/benchmarks), original data (https://epoch.ai/data/benchmark_data.zip), CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). StackReplay selected three Epoch-run records from the Diamond subset and converted accuracy fractions to percentages; no endorsement is implied. Current public methodology v1.0.6 is contextual only; run-specific suite and setup details are unknown.
- Epoch AI published these numerical run results; StackReplay did not reproduce the evaluations and makes no reproduction claim.
- The GPQA Diamond suite revision used for this run is unreported. Public methodology version 1.0.6 is current context only and is not asserted for this run.
- The 2026-10-04 date records the archive snapshot that was checked. Original run publication and completion dates are unknown.
- The published Diamond subset contains 198 questions, but the actual scored sample count, repetitions, and total attempts for this run are unreported.
- Scaffold revision, system prompt, temperature, reasoning parameter payload, reasoning-token budget, output/time/cost caps, endpoint, tools, seed, and error handling are unreported.
- The Epoch configuration label does not establish equal reasoning effort or a comparable setup across providers.
- The reported stderr is retained as standard error, not converted into or presented as a 95% confidence interval.
- Epoch lists a September 1 release date while the exact 0902 snapshot was released September 2; this source discrepancy is preserved rather than corrected.
Original methodology ↗ · Original source ↗
Epoch AI · Epoch AI Capabilities and benchmarking: GPQA Diamond mean accuracy. Selected Epoch-administered GPQA Diamond result rows and linked metadata from the official archive. Epoch's archive README links CC BY 4.0 for its data. Epoch's use-this-data terms state that externally sourced data retains its original licensing. This licensed-dataset basis is scoped only to these three selected Epoch-run numerical records and their attribution; it does not license the mixed archive, GPQA tasks, logs, provider prose, artwork, or StackReplay's broader export. Checked Oct 4, 2026. Redistribution terms ↗
13 QwenCloud Team · Qwen3.8-Max release: Terminal-Bench 2.1 · Published Aug 3, 2026
Developer reported · Qwen 3.8 Max
QwenCloud reports Qwen3.8-Max evaluated with Claude Code, avg@10, a 5-hour timeout and max_tokens=131,072.
- This selected numerical fact is developer reported. StackReplay did not reproduce the evaluation. Different or unreported setups do not establish a matched comparison.
- Evaluation start, completion and original run publication dates, actual scored task count, task subset, suite revision beyond 2.1, exact model endpoint/version, prompts, seeds and error handling are unreported. Publication dates identify the source articles, not the evaluation runs.
- Publication uses the visible article and official news-list date 2026-08-03. The HTML meta date is 2026-07-20 15:56:48; this discrepancy is unresolved.
- Claude Code version, reasoning effort, tools, temperature/top-p, API surface and sandbox details are unreported. The avg@10 label is retained without inferring a scored task count.
- The provider table omits the unit marker. Percent display is StackReplay's metric interpretation using the official Terminal-Bench 2.1 accuracy scale (https://www.tbench.ai/news/terminal-bench-2-1), not a claim of leaderboard admission or exact task counts.
Original methodology ↗ · Original source ↗
QwenCloud Team · Qwen3.8-Max release: Terminal-Bench 2.1. Selected first-party numerical fact with source attribution and our own short methodology description under the existing provider-facts policy. No explicit benchmark dataset redistribution grant was observed. This basis does not cover publisher charts, prose, questions, answers, logs or underlying datasets. Checked Oct 4, 2026. Redistribution terms ↗
14 Z.ai · GLM-5.3 release: Terminal-Bench 2.1 · Published Aug 14, 2026
Developer reported · GLM 5.3
Z.ai reports GLM-5.3 evaluated with Claude Code 2.1.207, temperature=1.0, top_p=1, max_new_tokens=65,536 and a 6-hour timeout.
- This selected numerical fact is developer reported. StackReplay did not reproduce the evaluation. Different or unreported setups do not establish a matched comparison.
- Evaluation start, completion and original run publication dates, actual scored task count, task subset, suite revision beyond 2.1, exact model endpoint/version, prompts, seeds and error handling are unreported. Publication dates identify the source articles, not the evaluation runs.
- Publication uses the official blog root data date 2026-08-14. The related documentation dateModified 2026-09-18 is not the blog publication date.
- Reasoning effort, tools, trial count, API surface and sandbox details are unreported for Terminal-Bench 2.1. Terminal-Bench 3.0 effort, aggregation and context/output settings are not applied to this observation.
- The provider table omits the unit marker. Percent display is StackReplay's metric interpretation using the official Terminal-Bench 2.1 accuracy scale (https://www.tbench.ai/news/terminal-bench-2-1), not a claim of leaderboard admission or exact task counts.
Original methodology ↗ · Original source ↗
Z.ai · GLM-5.3 release: Terminal-Bench 2.1. Selected first-party numerical fact with source attribution and our own short methodology description under the existing provider-facts policy. No explicit benchmark dataset redistribution grant was observed. This basis does not cover publisher charts, prose, questions, answers, logs or underlying datasets. Checked Oct 4, 2026. Redistribution terms ↗
15 MiniMax · MiniMax M3 release: Terminal-Bench 2.1 · Published Jun 1, 2026
Developer reported · MiniMax M3
MiniMax reports M3 evaluated through its official API using Terminus 2 on internal infrastructure, an 8C16G sandbox, a 2-hour timeout and 128K maximum output tokens.
- This selected numerical fact is developer reported. StackReplay did not reproduce the evaluation. Different or unreported setups do not establish a matched comparison.
- Evaluation start, completion and original run publication dates, actual scored task count, task subset, suite revision beyond 2.1, exact model endpoint/version, prompts, seeds and error handling are unreported. Publication dates identify the source articles, not the evaluation runs.
- Publication uses the visible article date 2026-06-01. JSON-LD datePublished is 2026-05-31T17:31:18.000Z; both source dates are retained.
- The official API evaluation surface is reported; the exact endpoint and model version are unreported. Reasoning effort, tools, trial count, temperature/top-p and the exact Terminus 2 revision are unreported.
Original methodology ↗ · Original source ↗
MiniMax · MiniMax M3 release: Terminal-Bench 2.1. Selected first-party numerical fact with source attribution and our own short methodology description under the existing provider-facts policy. No explicit benchmark dataset redistribution grant was observed. This basis does not cover publisher charts, prose, questions, answers, logs or underlying datasets. Checked Oct 4, 2026. Redistribution terms ↗
16 Moonshot AI · Kimi K3 Tech Blog: Terminal-Bench 2.1 · Publication date unreported
Developer reported · Kimi K3
Moonshot AI reports Kimi K3 evaluated with Kimi Code at max reasoning effort, temperature = 1.0 and top-p = 1.0.
- This selected numerical fact is developer reported. StackReplay did not reproduce the evaluation. Different or unreported setups do not establish a matched comparison.
- The original article publication date is unreported. Image asset dates, the weights-release statement, model release and October 4 check dates do not establish article publication.
- Kimi Code version, trial count, scored task count and subset, suite revision beyond 2.1, evaluation start and completion dates, timeout, maximum output tokens, tools, exact endpoint/version, prompts, seeds and error handling are unreported.
- The provider table omits the unit marker. Percent display is StackReplay's metric interpretation using the official Terminal-Bench 2.1 accuracy scale (https://www.tbench.ai/news/terminal-bench-2-1), not a claim of leaderboard admission or exact task counts.
Original methodology ↗ · Original source ↗
Moonshot AI · Kimi K3 Tech Blog: Terminal-Bench 2.1. Selected first-party numerical fact with source attribution and our own short methodology description under the existing provider-facts policy. No explicit benchmark dataset redistribution grant was observed. This basis does not cover publisher charts, prose, questions, answers, logs or underlying datasets. Checked Oct 4, 2026. Redistribution terms ↗

