In the last year, some of the most frustrating production issues I've seen weren't outages. The dependency was up. It returned 200 OK . It just didn't return the same thing anymore. A field our code relied on was removed from a well-known API's payload. An LLM provider started rejecting a request format that was valid the week before; the SDK types still allowed it. A model alias quietly pointed at a new snapshot, and our support bot's tone changed overnight. A remote MCP server add...