How to read this table
Higher is generally better, but scores can move with prompt format, sampling settings, tool access, reasoning effort, evaluation harness versions, and contamination controls. A creator-reported number and a third-party reproduction are different grades of evidence.
Current and legacy evaluations
MMLU-Pro, SWE-bench Verified, and LiveCodeBench remain for older records. Newer source material may instead report SWE-Bench Pro or an Artificial Analysis index. They are separate columns because unlike evaluations must not be merged.
Why there is no overall score
Averaging science questions, repository repairs, and agent indices implies that each matters equally to every buyer. It also hides missing tests. Choose the evaluation closest to your workload, then validate privately.