Tech
My factual-recall tasks were scoring format, not facts
Originally published at erikhill.dev . The numbers below are checked against the repository they come from.
My factual-recall tasks were scoring format, not facts
This is a finding about my own harness. The suspect is the probe, not the models it measures.
What happened
I built a detector that decides whether a day's drift run contains anything worth
writing up. The first thing it did was accuse my own suite. One task,
fact-element , had failed on three or ...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to