Back Thesoufancenter Assessing the Malicious Use of Advanced AI Models
Not even a week after Anthropic released Fable 5, its most capable model yet for public use, the U.S. government issued an export-control directive ordering Anthropic to block all access to the model, as well as its base model, Mythos 5 — its most capable model, accessible only to a select few cybersecurity defenders and infrastructure providers — for any foreign national. Rather than try to sort access by nationality, Anthropic disabled the two models for all customers worldwide to ensure compliance. The federal government has alleged that Fable 5 and Mythos 5 should not be accessible to foreign nationals for national security reasons and appears to have done so based on jailbreaking concerns, a term referring to prompt-based attacks intended to circumvent the restrictions engineers have put in place. The details of this latest incident have highlighted the extent of concern highly advanced models falling into the hands of state adversaries, violent non-state actors, and criminal entities that would enable malign activity. Simultaneously, it also demonstrates the extent to which highly capable models can be used across national security functions. Perhaps most alarmingly, the capabilities demonstrated by these AI models, including advanced biological research competence that may be used for weapon development as well as highly efficient vulnerability discovery and attack-surface auditing competencies that may be used for offensive cyber operations, have underscored why superiority in AI capabilities over a peer power matters so greatly amid great power competition between the People’s Republic of China (PRC) and the United States.
Cybersecurity expert Katie Moussouris from Luta Security, asked by Anthropic to review a report by Amazon researchers that seems to have inspired the Trump administration’s concerns Fable 5, found that the identified vulnerability is a basic jailbreaking method. When the authors of the report asked Fable 5 to fix flawed code, it did so, thereby de facto exposing the vulnerabilities in the software that it patched, something it had refused to do when prompted to review the code for security issues. Whether the Trump administration’s action was solely concerned with national security or began when Anthropic refused to allow the U.S. military to use its models for fully autonomous weapons systems, it is clear that advanced AI models offer significant new capabilities to both red and blue teams — labels used in cybersecurity spaces to represent opposing sides, with red representing the adversary, and blue representing “defense,” such as states or vulnerable organizations — globally and will be crucial in understanding the respective capabilities of the PRC and the U.S. for both defensive and offensive action.
Jailbreaking large language models (LLMs) is not novel, and various techniques have evolved since OpenAI released ChatGPT in November 2022 to undo safeguards. However, applying such jailbreaking techniques to these incredibly advanced models calls into question what having such capabilities means for criminal organizations and terrorist entities, which have already used various lesser capable generative AI tools for operational planning, propaganda, and illicit profit-making. Additionally, the demonstrated power of Mythos 5 suggests that the development of a similarly capable model by any U.S. adversary, such as the PRC, could significantly expand their ability to conduct cyberattacks on U.S. critical infrastructure and government offices, as well as to infiltrate sensitive systems for data collection on a massive scale and with a consistent tempo. From a security perspective, the U.S. would thus be wise to adopt models like Mythos 5 for good, something it had already begun with its Project Glasswing , an initiative to secure critical software in the AI era that brings together various American companies, including Apple, Broadcom, Cisco, NVIDIA, etc. For now, many technology companies are adamant that they have access to the Fable and Mythos. Alex Stamos, who formerly served as ’s Chief Security Officer, penned an open letter calling for the export controls to be rescinded. It is now signed by security professionals from companies including Nvidia, Adobe, Zoom, and Google. Stamos’ letter underscores the value of these advanced AI models for cyber defenders, also against a backdrop of the pacing threat posed by Chinese open-weight models, which he describes as “only months behind the best American models, and those are the models we know .”
Jailbreaking is not the only way hostile states, criminal entities, and terrorist groups can gain access to unfiltered models that respond to harmful instructions. Some models have been deliberately trained to respond to instructions widely perceived as harmful, and have been dubbed ‘dark LLMs’. Ethical safeguards have been stripped, enabling assistance with illicit activities such as automating fraud and phishing attacks, or responding to prompts that request instructions for making explosives. Over the past year, there has also been an uptick in malicious actors using agentic AI capabilities to further automate illicit operations, rather than as enablers and facilitators of illicit operations, as generative AI capabilities were. In November 2025, Anthropic disclosed the first known large-scale AI-orchestrated cyberattack, attributed to a Chinese state- group , in which Claude executed an estimated 80 to 90 percent of the operation independently against various targets, including technology companies, financial institutions, and government agencies, with some successful intrusions.
According to data collected by the FBI’s Internet Crime Complaint Center, AI-related scams have cost Americans at least $893 million in 2025. The primary culprit is a range of deepfake tactics, including voice-cloning attacks such as impersonating executives to approve payments and relatives to make scams more persuasive. While AI-enabled misuse by states has been rampant (e.g., the widespread use of deepfakes and automated job-application processes by North Korean operators to gain employment at strategic U.S. companies), it also risks placing capabilities in the hands of low-resourced actors. As in other fields, AI tools, both generative and agentic, have lowered the barriers to sophisticated cybercrime, enabling criminals with few technical skills to develop ransomware or run operations that previously required significantly more technical expertise and manpower.
The 2026 AI Safety Report , one of the largest transnational AI safety collaborations, finds in its assessment of the impact of AI advancements on cybersecurity that “whether attackers or defenders will benefit more from AI assistance remains uncertain.” Similarly, when assessing other malicious uses of AI models, including operational planning by violent extremists and the manipulation of the information environment by hostile actors, it is not yet clear whether advanced models, including those that can be jailbroken or those that are open-source and more easily ablated, will cede the advantage to the blue or red team. It is clear, however, that defender communities across the U.S. are benefiting greatly from advanced models that help with vulnerability detection and patching, as hostile states and criminal organizations continue to exploit vulnerabilities in U.S. companies and government agencies to achieve their strategic objectives.
The full story
This article is one source in a clustered incident — the cluster page carries the summary, timeline and every other outlet covering it.
