Anthropic’s public work on AI R&D automation points to a more demanding way to track advanced AI systems: measure not only what a model can do in an evaluation, but also how it is used in consequential research decisions and whether it creates operational risks . The approach matters because AI is increasingly being applied to the work of developing and evaluating AI itself, making a single benchmark an incomplete picture of progress. The company’s model reports, system cards and Transpa...