Tech
Fastest LLM 2026: Mercury 2.5 Beats Luna and Haiku
Verdict: Inception Labs' Mercury 2.5 is the fastest LLM you can call from an API in September 2026 β 1,107 tokens per second on widely available NVIDIA GPUs, at a list price of $0.20/$0.75 per million tokens that undercuts every rival here on output. For latency-bound work (voice agents, search pipelines, coding subagents), Mercury wins. For cheap general-purpose chat with long context, GPT-5.6 Luna still wins. Claude Haiku 4.5 and Gemini 3.5 Flash-Lite are squeezed between the two.
...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to