Six frontier AI launch posts and nine public leaderboards gave us 162 model-vs-model gaps to check. 20 separate cleanly. That was not the result that bothered us most. The bigger problem was how often the published data did not let us answer the question at all. Of the 44 gaps quoted in the six launch posts, 11 separate, 5 do not, and 28 cannot be checked honestly from what was published. Of 118 adjacent leaderboard pairs, 9 separate. And across the broader 617-row source audit...