Ranking basis
Ordered by the numeric context-window field in the cited model registry.
Comparability limit: Advertised capacity does not establish usable recall, latency, or cost at maximum length.
Ranked snapshot
1. Gemini 3.5 Flash — 1,048,576 tokens
Largest represented context among the selected stable current models. Google's stable Flash model for sustained agentic execution, coding, and long-horizon tasks, with multimodal input and a million-token context window.
Watch for: computer use remains preview functionality; pricing requires workload-specific verification.
2. Gemini 2.5 Pro — 1,048,576 tokens
Same represented capacity with multimodal reasoning positioning. Google's thinking model for multimodal reasoning over large codebases, video, audio, and documents. It combines a long context window with strong scientific reasoning.
Watch for: tiered pricing is workload-dependent; closed weights.
3. GPT-4.1 — 1,047,576 tokens
Nearly the same million-token capacity with code-editing focus. An API-focused model built around instruction following, code editing, and very long inputs. Its million-token context makes it useful for repository and document-set analysis.
Watch for: no open weights; long prompts can still dilute retrieval accuracy.
How to choose
- Run needle and synthesis tests at your real prompt lengths.
- Compare time to first token as context grows.
- Use retrieval to control cost even when everything technically fits.
Evaluation recipe
Build a permissioned set of representative tasks, including expected failures and ambiguous inputs. Run candidates with equivalent prompts and tools. Score task success, latency, total tokens, retries, and human correction time. Keep provider-reported results separate from private measurements.