OpenAI

GPT-4.1

An API-focused model built around instruction following, code editing, and very long inputs. Its million-token context makes it useful for repository and document-set analysis.

codinglong-contextvision

Model record

GPT-4.1 model overview

GPT-4.1 is a api model from OpenAI. It was released on April 14, 2025. Its strongest case is million-token context window, while buyers should account for no open weights.

GPT-4.1 benchmark snapshot

EvaluationReported scoreWhat it probes
MMLU-Pro83.7Broad knowledge and multi-step reasoning
GPQA Diamond66.3Graduate-level science reasoning
SWE-bench Verified54.6Verified real-repository issue resolution
LiveCodeBench44.7Contamination-aware competitive programming
SWE-Bench ProNot reportedLonger, harder professional software tasks
Artificial Analysis Intelligence IndexNot reportedComposite third-party capability index; version matters

Scores are percentages reported by model creators or benchmark maintainers under varying settings. A blank is preferable to an inferred result. See our methodology.

GPT-4.1 strengths

  • Million-token context window
  • Reliable instruction following
  • Competitive code editing

GPT-4.1 limitations

  • No open weights
  • Long prompts can still dilute retrieval accuracy
Editorial take: Benchmark rank should narrow a shortlist, not close a purchase. Run a private evaluation with representative prompts, failure cases, latency targets, and total token costs.

Compare GPT-4.1

Source record

Specifications and scores are linked to the best source located during review. Provider-reported results are not presented as third-party lab reproductions.

Read OpenAI model announcement