Two independent benchmark results for Opus 5.5 came out this week, and they don't really agree. One has it in first place. The other has it fast and cheap, but third on secure code once you take out the answers it memorized. A few days ago I read through Anthropic's Opus 5.5 docs to see what they say got better. That's Anthropic talking about their own model, though. I wanted to see what people who didn't build it found, so I read the first two outside evaluations I could find. Quick dis...