Model record
Claude 3.7 Sonnet model overview
Claude 3.7 Sonnet is a api + app model from Anthropic. It was released on February 24, 2025. Its strongest case is controllable extended thinking, while buyers should account for closed weights.
Claude 3.7 Sonnet benchmark snapshot
| Evaluation | Reported score | What it probes |
|---|---|---|
| MMLU-Pro | 84.1 | Broad knowledge and multi-step reasoning |
| GPQA Diamond | 84.8 | Graduate-level science reasoning |
| SWE-bench Verified | 70.3 | Verified real-repository issue resolution |
| LiveCodeBench | 46.4 | Contamination-aware competitive programming |
| SWE-Bench Pro | Not reported | Longer, harder professional software tasks |
| Artificial Analysis Intelligence Index | Not reported | Composite third-party capability index; version matters |
Scores are percentages reported by model creators or benchmark maintainers under varying settings. A blank is preferable to an inferred result. See our methodology.
Claude 3.7 Sonnet strengths
- Controllable extended thinking
- Strong software-engineering performance
- Large document context
Claude 3.7 Sonnet limitations
- Closed weights
- Extended thinking increases token usage
Editorial take: Benchmark rank should narrow a shortlist, not close a purchase. Run a private evaluation with representative prompts, failure cases, latency targets, and total token costs.
Compare Claude 3.7 Sonnet
- Claude 3.7 Sonnet vs GPT-4.1 — review fit, trade-offs, listed prices, and an evaluation plan.
- Claude 3.7 Sonnet vs Gemini 2.5 Pro — review fit, trade-offs, listed prices, and an evaluation plan.
- Claude 3.7 Sonnet vs DeepSeek R1 — review fit, trade-offs, listed prices, and an evaluation plan.
Source record
Specifications and scores are linked to the best source located during review. Provider-reported results are not presented as third-party lab reproductions.
Read Anthropic model announcement ↗