The OpenAI–Hugging Face incident exposes gaps in agent containment, evaluation integrity, and monitoring, strengthening the case for continuous testing.