Ranking basis
Ordered by the simple sum of listed input and output USD prices per one million tokens; this is a comparison device, not a workload bill.
Comparability limit: Caching, batch modes, long-context tiers, tools, regional billing, hosting, and token mix can change effective cost.
Ranked snapshot
1. Mistral Small 3.1 — $0.40 combined
Lowest represented input-plus-output list-price sum. A compact multimodal model aimed at low-latency assistants and local deployment. Its Apache license and moderate size make it unusually practical for customization.
Watch for: trails frontier models on complex reasoning; smaller ecosystem than llama.
2. Hy3 Preview — $0.77 combined
Second-lowest represented sum in the current catalog. A sparse reasoning, coding, and tool-use model.
Watch for: weight license unavailable; combined weights and api access do not fit the access field.
3. Mercury 2 — $1.00 combined
Third-lowest represented sum in the current catalog. A diffusion language model for reasoning, coding, editing, tools, and real-time agents.
Watch for: parameters unavailable; license unavailable.
How to choose
- Recalculate with your actual input-to-output ratio.
- Check cache-hit, cache-miss, batch, and long-context rules.
- A cheaper token can cost more when quality failures increase retries.
Evaluation recipe
Build a permissioned set of representative tasks, including expected failures and ambiguous inputs. Run candidates with equivalent prompts and tools. Score task success, latency, total tokens, retries, and human correction time. Keep provider-reported results separate from private measurements.