AI agent evaluations can fail because parsers, graders, fixtures, or trace adapters are wrong. Test the evaluator before trusting its score.