Loading model intelligence…
Loading model intelligence…
Source-linked model intelligence
A curated reference to leading LLMs—reported benchmarks, listed pricing, licensing, and practical trade-offs in one place.
Newest source records
Entries are ordered by release date here and grouped by provider in the full directory. Unavailable or unverified fields are labeled rather than inferred.
Moonshot AI's native-vision, 2.8T-parameter MoE for long-horizon coding, knowledge work, and reasoning, available through Kimi apps and API with a one-million-token context window.
A multimodal reasoning model for text, images, audio, coding, and tools.
A diffusion language model for reasoning, coding, editing, tools, and real-time agents.
The flagship GPT-5.6 tier for difficult reasoning, long-horizon agents, coding, cybersecurity, science, and professional knowledge work, with max and multi-agent ultra settings.
The balanced GPT-5.6 tier for everyday production work, combining strong agentic and coding results with lower listed token prices than Sol.
The fastest and lowest-priced GPT-5.6 tier, intended for high-volume workflows that need current agent capabilities without flagship token costs.
SpaceXAI's frontier API model for coding, agentic tasks, and knowledge work, with configurable reasoning effort and image input.
A coding-focused agentic model trained for autonomous work in executable repositories.
Start with the job
Purpose-built shortlists account for deployment, cost, and licensing—not just a leaderboard position.
Code generation & maintenance
↗02Science & multi-step reasoning
↗03Provider-listed token prices
↗04Output speed & latency
↗05Large documents & repositories
↗06Downloadable & self-hosted
↗07Languages & regional coverage
↗08Search-grounded answers
↗09Private & controlled deployment
↗Current benchmark snapshot
Scores are included only when a source reports the exact evaluation. Blank cells are deliberate, not an invitation to estimate.
| Model | Provider | GPQA Diamond | SWE-Bench Pro | AA Intelligence Index |
|---|---|---|---|---|
| OpenAI | 94.6 | 64.6 | 58.9 | |
| OpenAI | 92.9 | 63.4 | 55 | |
| OpenAI | 92.3 | 62.7 | 51.2 | |
| Moonshot AI | — | — | — | |
| Anthropic | 92.6 | 80 | 59.9 |
Provider-reported and cited comparative results can use different settings. Missing benchmark values remain blank rather than estimated. Data checked July 19, 2026.