Tech
I Turned the Reasoning Dial to 'High' on 4 Models. It Fixed One Thing and Billed Me for Everything.
This is a submission for the Kaggle Benchmarking Challenge
I gave gpt-5.4-mini a logic puzzle: seven people, seven days, ten clues, "Who gives the talk on Friday?"
With reasoning effort set to none , it replied:
Cleo
FINAL ANSWER: Cleo
18 output tokens. $0.00024. Wrong. (The answer is Fay.) It gave the same wrong answer, word for word, on the second repeat.
At high it spent 1,333 tokens, cost about $0.006, and said Fay. That looks like an argument for always ...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to