Ranking basis
Ordered by the numeric context-window field in the cited model registry.
Comparability limit: Advertised capacity does not establish usable recall, latency, or cost at maximum length.
Ranked snapshot
1. Gemini 3.6 Flash — 1,048,576 tokens
Largest represented context among the selected stable current models. Google's current stable Flash model for fast multimodal reasoning, coding, and agent workflows through Gemini apps and API.
Watch for: closed weights; parameters are undisclosed; output is limited to 65,536 tokens.
2. Gemini 2.5 Pro — 1,048,576 tokens
Same represented capacity with multimodal reasoning positioning. Google's thinking model for multimodal reasoning over large codebases, video, audio, and documents. It combines a long context window with strong scientific reasoning.
Watch for: tiered pricing is workload-dependent; closed weights.
3. GPT-4.1 — 1,047,576 tokens
Nearly the same million-token capacity with code-editing focus. An API-focused model built around instruction following, code editing, and very long inputs. Its million-token context makes it useful for repository and document-set analysis.
Watch for: no open weights; long prompts can still dilute retrieval accuracy.
How to choose
- Run needle and synthesis tests at your real prompt lengths.
- Compare time to first token as context grows.
- Use retrieval to control cost even when everything technically fits.
Evaluation recipe
Build a permissioned set of representative tasks, including expected failures and ambiguous inputs. Run candidates with equivalent prompts and tools. Score task success, latency, total tokens, retries, and human correction time. Keep provider-reported results separate from private measurements.