A benchmark comparing JEV with frontier LLMs reveals why decision accuracy alone isn't enough when structured scores and probabilities must agree.