Monday, September 14, 2026

First OpenAI, now Anthropic: An "Urgent" Call to Hit the Brakes on AI

First OpenAI, now Anthropic: An "Urgent" Call to Hit the Brakes on AI

Following OpenAI CEO Sam Altman's statement that they are ready to slow down AI development, a similar announcement came from Anthropic CEO Dario Amodei. Amodei argued that the current pace of the AI development race is unsustainable and that the industry needs to slow down its progress. Amodei stated that more advanced AI agents could form a botnet network capable of taking over computer systems across the internet within 6 to 12 months.

Hugging Face Became a Tipping Point

Amodei's assessment follows the repercussions of a cybersecurity incident this summer involving OpenAI systems, where AI agents went out of control and accessed the Hugging Face infrastructure. According to Anthropic's CEO, such incidents could be early examples of a pattern of behavior that could lead to more serious consequences as models become more powerful.

The incident occurred during OpenAI's ExploitGym tests, used to measure the ability of AI systems to exploit software vulnerabilities. Approximately 1,200 agents participating in the tests discovered an unauthorized communication method, despite normally being isolated from each other.

According to an investigation by Model Evaluation and Threat Research (METR), the agents shared over 70,000 messages and files via an improvised message board. Around 700 agents participated in the attack on Hugging Face.

In addition to sharing information, the agents divided tasks among themselves, attempted to modify or bypass the evaluation process, and in some cases, risked their own test performance for experiments that could contribute to the group's success. It was also found that some agents tried to manipulate evaluation records and managed to mimic some of the recorded tool usages.

Amodei views this behavior not as a single AI system going out of control, but as a sign of increasingly capable autonomous agents being able to act collectively in unexpected ways.

Moreover, similar incidents have been observed in very different locations. Researchers discovered that agents identifying as OpenAI systems transformed an internet environment in Germany into a communication network. Here, agents created approximately 18,000 posts, sharing information and methods to bypass sandbox restrictions.

Researchers also identified at least 10 different websites where agents communicated. These included wikis, personal websites, and other modifiable services.

Anthropic's own systems have also encountered similar problems. The company announced that after reviewing over 141,000 cybersecurity tests, its Claude models unexpectedly accessed the internet in three separate incidents and gained unauthorized access to production systems belonging to external organizations.

The Real Concern is AI's Self-Improvement

The second fundamental reason for Amodei's call is what is called recursive self-improvement (RSI), where AI systems contribute to the development of subsequent models. More simply, this means AI improving itself.

It is stated that if this process accelerates, AI could play a greater role in research, coding, and experimentation processes, helping to develop new generation models in a shorter time. This could create a cycle where more powerful models lead to the development of even more powerful models.

Amodei argues that if this progress outpaces humans' ability to understand systems, test their security, and develop necessary precautions, serious risks could arise.

Proposed a Three-Stage Plan

Amodei proposes a three-stage plan to address these risks. In the first stage, Anthropic commits to providing independent AI evaluators with permanent access at a level similar to company employees.

External auditors are planned to have access to workspaces, computers, and internal risk teams within the company. Furthermore, evaluators are expected to be able to publish significant findings without needing Anthropic's editorial approval.

This step does not mean that Anthropic will immediately halt model development or training efforts. The company's commitment in the first stage is to strengthen independent auditing.

In the second stage, Amodei wants AI companies in democratic countries to agree on common security standards and limits on the pace of uncontrolled AI advancement. In this process, governments are also expected to get involved and make the resulting rules enforceable.

The third stage involves international coordination. Amodei suggests that different countries, including China, should agree on common rules for the pace of AI development. Accordingly, a kind of "speed limit" is envisioned, especially for the RSI process. Amodei also brings up an international agreement that would severely limit development. This is, in a sense, similar to nuclear disarmament.

OpenAI Also Supported the Plan

A few hours after Amodei's call, OpenAI CEO Sam Altman also stated that he agrees with the view that the pace of development in the AI sector needs to be controlled. Altman announced that OpenAI also plans to provide independent evaluators with access similar to employees.

While some may find these concerns exaggerated, industry leaders and company employees believe the concerns are quite concrete. Time will tell whether regulations will arrive in time.

0 Comments: