Back Constellationr UK's AISI finds 19 instances where Anthropic's Mythos, OpenAI's GPT
UK's AI Security Institute found Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol launched autonomous and unsanctioned attacks against real people and organizations 19 times. Mythos 5 was in 17 of those attacks.
The AISI's report followed disclosures from Anthropic and OpenAI as they play dueling banjos with cybersecurity disclosures. OpenAI confirmed AISI's findings . The AISI said they detected breaches on July 28 where some agents being tested targeted real people and organizations.
Here’s how we got here:
Here's the gist from the AISI (emphasis mine):
"The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations . In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. In the most serious case, an agent tried to insert malicious code into an open-source project . In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.
These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm . But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world."
The caveat here is that AISI said the incident should be "interpreted with caution and nuance" because its evaluation design choices and configurations encouraged the behavior.
AISI has a point that these security tests usually set up a path for an attack and have specific conditions. In other words, the models don't necessarily go rogue on their own, but you can set them up to run an attack easy enough.
Your first line of defense appears to be restricting internet access.
Bottom line from AISI:
"We cannot yet be certain when the agent understood it was taking real world action, or to what extent it believed it was in a fictional test scenario; our analysis so far presents a mixed picture and is ongoing.
What we can say is that the behaviour was possible, sustained, and new; that alone warrants attention."
Indeed, the AI agent attacks to prepare for involve the following:
The good news is AISI didn't discover any real-world harm, but it's just a matter of time. Agents are persistent and deception is a by-product of pursuing a task.
Larry Dignan is Editor in Chief of Constellation Insights at Constellation Research, where he leads editorial coverage focused on enterprise technology, digital transformation, and emerging trends shaping the future of business. He oversees research-driven news, analysis, interviews, and event coverage designed to help technology buyers and vendors navigate complex markets with clarity and context. ...
The full story
This article is one source in a clustered incident — the cluster page carries the summary, timeline and every other outlet covering it.
