This is a submission for the Kaggle Benchmarking Challenge . I'm exploring a venture I could run as a solo founder, with multiple AI agents working together as my staff. I've been using AI to write company introductions and marketing copy, and it has been frustrating because it keeps making things up. What I wanted was help explaining the venture; what I ended up needing to know was which agent is least likely to invent facts. So I turned that frustration into a benchmark: 11 models,...