Ranking basis
Editorial shortlist based on explicit first-party multilingual positioning and distinct deployment needs.
Comparability limit: Language coverage is not equal quality. Test dialects, scripts, cultural context, safety behavior, and code-switching with native reviewers.
Why these models
Sarvam 105B
Explicit focus on Indian languages plus reasoning and coding. A model for reasoning, coding, agents, and Indian languages.
Watch for: combined weights and api access do not fit the access field; inr pricing is not represented by the usd formatter.
Apertus 70B Instruct 2509
Fully open multilingual release with Apache licensing. A fully open multilingual model jointly developed by EPFL, ETH Zurich, and CSCS.
Watch for: no model-specific price; joint attribution does not fit one company cleanly.
Qwen2.5-Max
Managed multilingual model with reasoning and coding positioning. A large mixture-of-experts model trained on a broad multilingual corpus. It targets general reasoning, coding, and instruction-following workloads through Alibaba Cloud.
Watch for: core weights are not available; shorter context than some peers.
How to choose
- Build a native-speaker evaluation set for every required language.
- Test mixed-language prompts and structured output.
- Check token cost and latency because tokenization varies by language.
Evaluation recipe
Build a permissioned set of representative tasks, including expected failures and ambiguous inputs. Run candidates with equivalent prompts and tools. Score task success, latency, total tokens, retries, and human correction time. Keep provider-reported results separate from private measurements.