Ranking basis
Editorial shortlist based on verified weight access, license clarity, workload fit, and operational practicality.
Comparability limit: This is not a universal capability ranking; hardware requirements and license rights differ substantially.
Why these models
Mistral Small 3.1
Apache-licensed and comparatively practical at 24B parameters. A compact multimodal model aimed at low-latency assistants and local deployment. Its Apache license and moderate size make it unusually practical for customization.
Watch for: trails frontier models on complex reasoning; smaller ecosystem than llama.
DeepSeek R1
MIT-licensed reasoning family with multiple distilled variants. A mixture-of-experts reasoning model released with permissive weights. R1 made high-end chain-of-thought style performance more accessible to self-hosted teams.
Watch for: full model has demanding infrastructure needs; long reasoning traces add latency.
Nemotron 3 Ultra 550B-A55B
Million-token downloadable model for teams with large serving infrastructure. A million-token model for reasoning, agents, tools, RAG, code, math, and science.
Watch for: custom license; no model-specific price.
How to choose
- Read the actual weight license before commercial use.
- Verify runtime support and memory requirements before downloading.
- Evaluate smaller or quantized variants before frontier-scale infrastructure.
Evaluation recipe
Build a permissioned set of representative tasks, including expected failures and ambiguous inputs. Run candidates with equivalent prompts and tools. Score task success, latency, total tokens, retries, and human correction time. Keep provider-reported results separate from private measurements.
Docker & Podman self-hosting guide →