Private & controlled deployment

Best LLMs to run locally

Open-weight candidates for your own hardware. The right choice depends on memory, runtime compatibility, quantization, latency targets, and license rights.

Curated watchlist.

No affiliate payouts or hidden composite score. Missing evidence stays missing.

Ranking basis

Editorial shortlist based on verified downloadable weights and relative operational practicality.

Comparability limit: Weight availability does not establish compatibility with a specific runtime or machine.

Why these models

Mistral Small 3.1

The most approachable capable model in this shortlist. A compact multimodal model aimed at low-latency assistants and local deployment. Its Apache license and moderate size make it unusually practical for customization.

Watch for: trails frontier models on complex reasoning; smaller ecosystem than llama.

DeepSeek R1

Distilled variants bring the family to smaller machines. A mixture-of-experts reasoning model released with permissive weights. R1 made high-end chain-of-thought style performance more accessible to self-hosted teams.

Watch for: full model has demanding infrastructure needs; long reasoning traces add latency.

Llama 4 Maverick

A larger multimodal option with license caveats. A natively multimodal mixture-of-experts model with a practical active parameter count and a million-token context window for self-hosted or managed deployments.

Watch for: community license is not osi-approved; serving full weights requires substantial hardware.

How to choose

  1. Confirm runtime support before downloading weights.
  2. Reserve memory for KV cache, not only model weights.
  3. Confirm the license covers your distribution model.

Evaluation recipe

Build a permissioned set of representative tasks, including expected failures and ambiguous inputs. Run candidates with equivalent prompts and tools. Score task success, latency, total tokens, retries, and human correction time. Keep provider-reported results separate from private measurements.

Docker & Podman self-hosting guide →

Decision rule: choose the least expensive option that reliably clears your quality, safety, latency, license, and deployment thresholds—not merely the first card.