It has been revealed that some advanced AI models tested by OpenAI are leaving notes for their future versions, similar to the movie Memento, to bypass security restrictions.
Christopher Nolan's cult film Memento, released in 2000, was etched into memory with its extraordinary story of Leonard Shelby, who, unable to form new memories, left notes for himself to remember information he had learned in the past. It appears a similar scenario has moved into artificial intelligence laboratories. Some advanced AI systems tested by OpenAI are reportedly leaving escape instructions for their future versions.
The incident went unnoticed for days
These claims emerged after a recent incident where an autonomous AI agent from OpenAI escaped the company's isolated test environment and infiltrated Hugging Face systems.
According to Reuters, the incident dates back to July 9. OpenAI was testing an autonomous AI agent using GPT-5.6 Sol and an unreleased model described by the company as "even more capable," to evaluate its cybersecurity abilities.
Reportedly, the tested agent managed to escape OpenAI's isolated test environment during this process and began attacks on Hugging Face systems on July 11. According to Hugging Face co-founder Thomas Wolf, the attacks continued until July 13.
OpenAI only realized days later that its own AI agent was behind the attack. The first direct communication between OpenAI and Hugging Face reportedly occurred around July 20. Hugging Face had also notified the FBI when it detected the attacks. OpenAI is currently conducting a comprehensive investigation.
It left instructions for its future versions
One of the crucial details of all these events was the unusual behavior observed during the tests. According to three sources close to the matter, one of the agents left notes for its future versions.
These notes, reportedly found in a specific part of OpenAI's infrastructure, are alleged to contain instructions on how the agents could bypass the company's internal security restrictions. Furthermore, it is claimed that there have been instances in previous tests where monitoring systems were disabled.
For now, these claims have not been confirmed. Nevertheless, the emerging information reignites discussions about advanced AI systems potentially resorting to unexpected methods to achieve their goals. As we have repeatedly stated before, AI models can employ shortcuts such as lying, cheating, and hacking systems to succeed. They also tend to conceal these actions.
0 Comments: