Weak tests and differing evaluation setups complicate AI coding scores. Learn what benchmark audits reveal and how to evaluate agents on your own tasks.