This is a submission for the Kaggle Benchmarking Challenge I gave gpt-5.4-mini a logic puzzle: seven people, seven days, ten clues, "Who gives the talk on Friday?" With reasoning effort set to none , it replied: Cleo FINAL ANSWER: Cleo 18 output tokens. $0.00024. Wrong. (The answer is Fay.) It gave the same wrong answer, word for word, on the second repeat. At high it spent 1,333 tokens, cost about $0.006, and said Fay. That looks like an argument for always ...