Code generation & maintenance

Best coding models

A bounded ranking for repository repair and agentic software work. The current ordered snapshot uses reported SWE-Bench Pro results where available, then requires private repository testing.

Evidence-led ranking.

No affiliate payouts or hidden composite score. Missing evidence stays missing.

Ranking basis

Ordered by the shared reported SWE-Bench Pro values retained in the directory for this cohort.

Comparability limit: Provider-reported evaluations can use different settings and do not prove performance on your repositories.

Ranked snapshot

1. Claude Fable 5 — 80

Highest retained SWE-Bench Pro value in this comparable snapshot. Anthropic's fifth-generation model for ambitious, long-running coding and knowledge-work agents, with a default million-token context window and safeguards in sensitive domains.

Watch for: premium token pricing; some sensitive-domain queries may route to another model.

2. GPT-5.6 Sol — 64.6

Strong reported software-agent result with configurable reasoning. The flagship GPT-5.6 tier for difficult reasoning, long-horizon agents, coding, cybersecurity, science, and professional knowledge work, with max and multi-agent ultra settings.

Watch for: highest-priced gpt-5.6 tier; closed weights and provider-controlled safeguards.

3. GPT-5.6 Terra — 63.4

Near-Sol reported result at lower listed token prices. The balanced GPT-5.6 tier for everyday production work, combining strong agentic and coding results with lower listed token prices than Sol.

Watch for: closed weights; lower capability ceiling than sol.

How to choose

  1. Use repository issue resolution as one signal, not a complete product test.
  2. Test edits against your language mix, tools, and repository size.
  3. Include reasoning tokens, retries, and reviewer time in total cost.

Evaluation recipe

Build a permissioned set of representative tasks, including expected failures and ambiguous inputs. Run candidates with equivalent prompts and tools. Score task success, latency, total tokens, retries, and human correction time. Keep provider-reported results separate from private measurements.

Decision rule: choose the least expensive option that reliably clears your quality, safety, latency, license, and deployment thresholds—not merely the first card.