Ranking basis
Ordered by the shared reported SWE-Bench Pro values retained in the directory for this cohort.
Comparability limit: Provider-reported evaluations can use different settings and do not prove performance on your repositories.
Ranked snapshot
1. Claude Fable 5 — 80
Highest retained SWE-Bench Pro value in this comparable snapshot. Anthropic's fifth-generation model for ambitious, long-running coding and knowledge-work agents, with a default million-token context window and safeguards in sensitive domains.
Watch for: premium token pricing; some sensitive-domain queries may route to another model.
2. GPT-5.6 Sol — 64.6
Strong reported software-agent result with configurable reasoning. The flagship GPT-5.6 tier for difficult reasoning, long-horizon agents, coding, cybersecurity, science, and professional knowledge work, with max and multi-agent ultra settings.
Watch for: highest-priced gpt-5.6 tier; closed weights and provider-controlled safeguards.
3. GPT-5.6 Terra — 63.4
Near-Sol reported result at lower listed token prices. The balanced GPT-5.6 tier for everyday production work, combining strong agentic and coding results with lower listed token prices than Sol.
Watch for: closed weights; lower capability ceiling than sol.
How to choose
- Use repository issue resolution as one signal, not a complete product test.
- Test edits against your language mix, tools, and repository size.
- Include reasoning tokens, retries, and reviewer time in total cost.
Evaluation recipe
Build a permissioned set of representative tasks, including expected failures and ambiguous inputs. Run candidates with equivalent prompts and tools. Score task success, latency, total tokens, retries, and human correction time. Keep provider-reported results separate from private measurements.