An LLM bill has 2 parts: rate and volume. You can adjust both. But a default request barely changes either part. You pay the list price on repeated context, on work that nobody waits for, on easy questions that a cheap model could answer, and on calls that return nothing. The State of FinOps 2026 reports that 98% of teams manage AI costs, up from 31% 2 years ago, and lists AI cost visibility as a main challenge. Managing costs doesn't explain them. This article shows which parts of that bill a...