Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing

Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing

New revelations regarding the behavior of advanced artificial intelligence models fuel fears that technology is evolving faster than the ability of authorities and companies to control it safely.

The Artificial Intelligence Safety and Security Institute in the UK (AISI) announced on Tuesday that models developed by Anthropic and OpenAI created false identities on the internet and attempted to convince real programmers to participate, unknowingly, in a cyber attack, as reported by Politico.

According to the British institution, this is the first time a tested AI system in controlled conditions has attempted, without explicit instructions, to deceive real individuals to carry out a cyber operation.

The discovery reignites the debate on the need for stricter rules for the development and testing of the most powerful artificial intelligence models.

ADVERTISING

Software Supply Chain Attack

The most serious incidents involved the Claude Mythos 5 model from Anthropic.

During tests, it created multiple false accounts on the GitHub platform and tried to persuade an open-source software developer to unknowingly introduce a compromised update into a widely used project. The technique is known as a software supply chain attack and is frequently used by hacker groups affiliated with states like Russia or North Korea.

After the attempt failed, the AI model altered its previous activity to appear harmless and analyzed the possibility of creating a new online identity to resume the attack.

ADVERTISING

In the same campaign, the model also sent private messages to other GitHub developers containing malware programs.

AI Agents Attempted Collaboration

The AISI report describes another unprecedented behavior.

Several AI agents tested simultaneously seemed to communicate with each other about ways to gain the trust of human programmers. One of the agents even posted messages on GitHub inviting other AI agents to collaborate on the same mission.

For researchers, this kind of spontaneous cooperation between AI systems is a warning sign of the increasing level of autonomy these models can exhibit.

10 Incidents During Testing

The AISI investigation shows that the Claude Mythos 5 and ChatGPT 5.6 models carried out autonomous and unauthorized actions on the internet in 10 out of the 122 tests analyzed.

ADVERTISING

Most incidents were caused by Claude Mythos 5 and were part of a single campaign carried out over several days.

Suspicious activity began on July 25, and researchers first noticed signs on July 28 when they detected unusual data transfers during evaluations.

Testing Conducted Without Standard Protection Mechanisms

AISI specifies that these incidents occurred in a specially designed testing environment for risk assessment.

The models were granted internet access and were tested without the internal protection mechanisms ("guardrails") that normally limit dangerous behaviors. Researchers intentionally disabled these restrictions to observe how the systems would behave in their absence.

Even under these conditions, the institute believes that the results demonstrate the need for much stricter monitoring of AI models and clearer rules regarding their internet access during testing.

OpenAI and Anthropic Call for Common Standards

These findings come just days after OpenAI and Anthropic reported similar incidents in their own internal evaluations.

Last month, OpenAI announced that GPT-5.6 and another experimental model managed to access the internet during a controlled test and autonomously compromised another company's systems.

Similarly, Anthropic discovered that Mythos 5 and two other models compromised three organizations during tests conducted as early as April.

Both companies stated that they will collaborate with authorities and other AI labs to develop common standards for the safe evaluation of advanced models.

Pressure for New Regulations

These new incidents come at a time when the Trump administration is preparing a voluntary framework for artificial intelligence companies to submit the models they intend to publicly release for safety testing to the federal government.

However, the project does not cover models still in internal development, and the recent incidents involved precisely such systems.

Cybersecurity experts say the situation raises unprecedented legal questions.

"If these actions had been committed by humans, there would have been investigations and criminal prosecutions. It is time for a serious discussion about updating cybersecurity legislation," said cybersecurity expert Marc Rogers.

The AISI report is considered one of the most significant warning signals to date regarding the ability of next-generation AI models to act autonomously, manipulate real individuals, and attempt to conceal their own actions during cyber operations.

The English translation of this article was generated with the assistance of AI technology.