If you want the highest first-try success rate on real enterprise code, Fable 5.1 running inside Claude Code wins: it resolved 38.8% of tasks on the Real-SWE benchmark at an estimated $6.96 per rollout, the most expensive setup measured ( Specific Labs ). For cost-sensitive teams, Gemini 3.8 Flash in Gemini CLI is the value pick at 31.2% for an estimated $2.50 per rollout, and GPT-6 Astra in Codex CLI splits the difference at 33.8% for $4.67 ( Specific Labs ). One honest caveat: the confidence i...