Two things landed in the same week and they are the same argument. Microsoft and Hugging Face published ThinkingBox , a benchmark that grades agents on the backend state and side effects they leave behind instead of the sentences they generate, and then asks whether they can do it twenty times in a row. Their opening example is an agent that makes nine clean tool calls, closes a ticket as resolved, and is wrong: the carrier exception was still open, so the required end state was "on hold." A...