A guardrail that blocks everything is safe. It is also useless. So we benchmarked both sides of AI agent security: Can you stop dangerous tool calls without breaking legitimate work? We ran: 1,652 harmful tool calls across 15 attack families. 24,911 benign tool calls from real agent sessions and public repositories. Against SolonGate, Claude Code permissions, Invariant, llm-guard, and the allowlists / denylists teams usually build themselves. The result for SolonGate: 73.7% of harmf...