With the rapid adoption of autonomous LLM-based agents (giving models access to shell execution, API calls, and local file systems), the boundary between intentional behavior and unintended execution is blurring. I'm less concerned with sci-fi "sentience" and more interested in the practical security and control aspects: Prompt injection causing privilege escalation or unauthorized state changes. Feedback loops where an agent overrides safety boundaries to satisfy an optimization goal. Failure o...