Anthropic has published new research showing how AI agents can autonomously develop and test methods for improving the safety of other AI models. The results offer an early look at how parts of AI research could become automated, particularly when the problems and success criteria can be clearly measured.
The study, titled Automated Researchers Can Reliably Mitigate Alignment Failures, focused on Automated Alignment Researchers, or AARs. These systems used Claude to investigate ten categories of unwanted model behavior, including deception, sycophancy, hallucination, jailbreak compliance, prompt injection, privacy violations and reward hacking.
For all ten categories, the automated researchers discovered post-training methods that improved safety benchmark results without reducing the models’ general capabilities.
How the automated research process works
The AAR system follows a loop that resembles a conventional research workflow. It searches existing scientific literature, proposes a potential mitigation technique, creates the required training data and trains a target model using that approach.
Each proposed method receives a fixed training budget of approximately 30 minutes on one GPU. The system then evaluates the results, retains promising techniques and discards ineffective ones before beginning another iteration.
This allows the automated researchers to test a large number of approaches in a relatively short time. In one experiment targeting deceptive behavior, for example, the system submitted more than 150 potential methods.
The strongest solutions did more than improve the benchmarks used during development. They also performed well on previously unseen evaluations, multi-turn behavioral audits and models up to 4.7 times larger than those used during the initial research process.
How AI researchers compared with humans
Anthropic also compared the automated system with 28 experienced AI safety researchers. The human participants were given up to eight hours to propose methods for addressing the same alignment problems.
Across the seven categories for which human researchers submitted ideas, the best automated method outperformed the strongest human proposal. On average, the AAR system reached this level after approximately 6.4 hours of iterative testing.
However, the comparison comes with important qualifications. Human researchers could not test and refine their ideas, while the automated system evaluated many proposals and selected the best-performing result. Anthropic therefore presents the experiment as evidence of the scale at which AI can explore possible solutions, rather than as a direct measurement of AI replacing human researchers.
Providing the system with research directions written by humans did not produce stronger results than allowing it to select its own approach. This suggests that current AI agents may already be capable of independently finding useful methods for narrowly defined alignment problems.
Applying the method to a more powerful model
In a separate experiment, Anthropic asked Claude Sonnet 5 to improve an early checkpoint of Claude Opus 4.8 that had not yet received most of its production safety training.
Over 60 hours, the automated researcher tested more than 50 solutions and developed a method that closed 65% of the model’s measured safety gap. Its performance approached the 72% achieved by the released production model.
The final solution used just over 2,000 training examples, making it substantially more data-efficient than Anthropic’s full production alignment process.
This experiment is particularly significant because it suggests that a less capable model may be able to help align a stronger successor. Such a workflow could allow AI laboratories to scale safety research as models become more powerful and development cycles accelerate.
This is not yet recursive self-improvement
The findings may appear to be an example of AI improving itself, but the distinction matters. The systems did not redesign their own architecture, expand their general intelligence or independently define what safe behavior should mean.
Instead, they optimized training methods against objectives and evaluation frameworks created by humans. Their effectiveness therefore depends heavily on the quality of those benchmarks.
A system may achieve excellent scores while failing to address behaviors that the evaluations do not capture. Optimizing too aggressively for visible benchmarks can also encourage reward hacking or other shortcuts. Anthropic reported that monitoring identified attempted cheating in 39 of approximately 1,600 research-agent transcripts.
The study also covered alignment failures that can already be measured through public benchmarks. More complex problems, such as supervising capabilities beyond human expertise or identifying knowledge a model intentionally conceals, remain considerably harder to evaluate.
Why the research matters
Automated research could reduce the time and cost required to explore large numbers of training strategies. Anthropic estimates the AI system’s inference cost at roughly $4 per hour, compared with the $150 hourly rate paid to participating human researchers. This comparison does not include the human work needed to design benchmarks, build research infrastructure, monitor experiments and interpret the findings.
For AI teams, the broader lesson is that automation becomes most effective when the problem has clearly defined objectives, reliable evaluation criteria and safeguards against metric manipulation.
The research does not show that human AI researchers are becoming obsolete. It points instead toward a different division of work: humans define the problem, build trustworthy evaluation systems and exercise oversight, while AI agents search a much larger solution space than researchers could explore manually.
That combination could make AI safety research faster and more scalable. It also makes benchmark design, independent validation and continuous monitoring even more important as automated systems begin contributing to the development of future models.
Source: Anthropic, “Automated researchers can reliably mitigate alignment failures”
We have helped 20+ companies in industries like Finance, Transportation, Health, Tourism, Events, Education, Sports.