Verdict: Inception Labs' Mercury 2.5 is the fastest LLM you can call from an API in September 2026 — 1,107 tokens per second on widely available NVIDIA GPUs, at a list price of $0.20/$0.75 per million tokens that undercuts every rival here on output. For latency-bound work (voice agents, search pipelines, coding subagents), Mercury wins. For cheap general-purpose chat with long context, GPT-5.6 Luna still wins. Claude Haiku 4.5 and Gemini 3.5 Flash-Lite are squeezed between the two. ...