Output speed & latency

Fastest LLMs: speed watchlist

An unranked editorial watchlist of models worth testing when output speed and latency matter. Exact comparable values are unavailable in the maintained evidence record, so this page does not manufacture an ordinal winner.

Curated watchlist.

No affiliate payouts or hidden composite score. Missing evidence stays missing.

Ranking basis

Artificial Analysis defines output speed as tokens received per second after the first token under a standardized rolling workload; that methodology informs the private test plan, not the inclusion order here.

Comparability limit: Serving speed changes by endpoint, provider, load, prompt length, reasoning behavior, and observation window. No model-level benchmark values are retained, so the cards are direct-test candidates rather than measured ranks.

Read the ranking methodology ↗

Why these models

Mercury 2

Provider-positioned diffusion model included for direct endpoint measurement; no third-party speed value is retained. A diffusion language model for reasoning, coding, editing, tools, and real-time agents.

Watch for: parameters unavailable; license unavailable.

Granite 4.0 H Small

Efficient active-size architecture included as an unmeasured self-hosting candidate; no speed value is claimed. A hybrid model for instructions, coding, RAG, multilingual work, and tools.

Watch for: no single deployable context limit disclosed; model-specific pricing unavailable.

Step 3.7 Flash

Low-active-parameter model included for direct endpoint measurement; no directory speed value is inferred. A sparse multimodal model for reasoning, coding, search, images, and tools.

Watch for: combined weights and api access do not fit the access field; cached-input pricing is not represented.

How to choose

  1. Measure time to first token separately from output speed.
  2. Benchmark the exact provider endpoint and prompt lengths you will buy.
  3. Record p50 and tail latency under concurrency, not one warm request.

Evaluation recipe

Build a permissioned set of representative tasks, including expected failures and ambiguous inputs. Run candidates with equivalent prompts and tools. Score task success, latency, total tokens, retries, and human correction time. Keep provider-reported results separate from private measurements.

Decision rule: choose the least expensive option that reliably clears your quality, safety, latency, license, and deployment thresholds—not merely the first card.