Long-running AI agents can drift, repeat failed actions, and misuse tools. Here’s how bounded context, validation, and execution limits improve reliability.