I built a provenance-based prompt injection firewall that caught attacks text detectors missed, but also blocked every legitimate action.