Originally published at erikhill.dev . The numbers below are checked against the repository they come from. My factual-recall tasks were scoring format, not facts This is a finding about my own harness. The suspect is the probe, not the models it measures. What happened I built a detector that decides whether a day's drift run contains anything worth writing up. The first thing it did was accuse my own suite. One task, fact-element , had failed on three or ...