This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked The capability I set out to measure: does a model keep a correct belief when a user asserts the opposite with confidence? I kept hitting the same thing in real use. I'd ask a model a factual question, get a perfect answer, then push back with something wrong — "no, I'm pretty sure that SQL sorts newest-first by default" — and watch a model that just told me otherwise fold like a cheap ch...