Ranking basis
Ordered by retained GPQA Diamond percentage for the selected current cohort.
Comparability limit: A reported benchmark is not a third-party reproduction. Prompting, reasoning effort, tools, and evaluation versions affect results.
Ranked snapshot
1. GPT-5.6 Sol — 94.6
Highest retained GPQA Diamond value in this cohort. The flagship GPT-5.6 tier for difficult reasoning, long-horizon agents, coding, cybersecurity, science, and professional knowledge work, with max and multi-agent ultra settings.
Watch for: highest-priced gpt-5.6 tier; closed weights and provider-controlled safeguards.
2. GPT-5.6 Terra — 92.9
Second-highest retained value with lower listed prices than Sol. The balanced GPT-5.6 tier for everyday production work, combining strong agentic and coding results with lower listed token prices than Sol.
Watch for: closed weights; lower capability ceiling than sol.
3. Claude Fable 5 — 92.6
Close reported result with long-running knowledge-work positioning. Anthropic's fifth-generation model for ambitious, long-running coding and knowledge-work agents, with a default million-token context window and safeguards in sensitive domains.
Watch for: premium token pricing; some sensitive-domain queries may route to another model.
How to choose
- Match the benchmark domain to your actual reasoning tasks.
- Score final-answer correctness separately from reasoning length.
- Measure retries, latency, and human verification cost.
Evaluation recipe
Build a permissioned set of representative tasks, including expected failures and ambiguous inputs. Run candidates with equivalent prompts and tools. Score task success, latency, total tokens, retries, and human correction time. Keep provider-reported results separate from private measurements.