The model race today is mostly about generation: longer context, prettier answers, stronger reasoning. But if you actually run agents, the thing that bites you usually isn't "generates poorly" — it's "decides too slowly." Every step an agent takes requires choosing the next action. If that choice also runs the full token-by-token generation pipeline, both latency and cost climb — a single "which button" can burn hundreds of tokens. CLM (Contrastive Language Models, 2913 stars, Apache-...