Tech
Day 3: The benchmark caught me too.
Kaggle Benchmarking Challenge. Previously: Day 0, the benchmark · Day 1, most of my bugs looked like model behaviour · Day 2, the model I want is the one that's boring everywhere
Where we are
Same benchmark: 200 invented items in four shapes (route, classify, judge,
ground), one in five answerable only with ESCALATE . Two numbers per model,
never merged: a task score, and a false-confidence rate.
Day 2 said I'd stop ranking models by their average and rank them by th...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to