Anthropic

Claude 3.7 Sonnet

A hybrid-reasoning model that lets developers trade latency and cost for a longer visible reasoning process. It is a strong fit for code maintenance and multi-step document work.

reasoningcodingvision

Model record

Claude 3.7 Sonnet model overview

Claude 3.7 Sonnet is a api + app model from Anthropic. It was released on February 24, 2025. Its strongest case is controllable extended thinking, while buyers should account for closed weights.

Claude 3.7 Sonnet benchmark snapshot

EvaluationReported scoreWhat it probes
MMLU-Pro84.1Broad knowledge and multi-step reasoning
GPQA Diamond84.8Graduate-level science reasoning
SWE-bench Verified70.3Verified real-repository issue resolution
LiveCodeBench46.4Contamination-aware competitive programming
SWE-Bench ProNot reportedLonger, harder professional software tasks
Artificial Analysis Intelligence IndexNot reportedComposite third-party capability index; version matters

Scores are percentages reported by model creators or benchmark maintainers under varying settings. A blank is preferable to an inferred result. See our methodology.

Claude 3.7 Sonnet strengths

  • Controllable extended thinking
  • Strong software-engineering performance
  • Large document context

Claude 3.7 Sonnet limitations

  • Closed weights
  • Extended thinking increases token usage
Editorial take: Benchmark rank should narrow a shortlist, not close a purchase. Run a private evaluation with representative prompts, failure cases, latency targets, and total token costs.

Compare Claude 3.7 Sonnet

Source record

Specifications and scores are linked to the best source located during review. Provider-reported results are not presented as third-party lab reproductions.

Read Anthropic model announcement