Kucoin Anthropic's Opus 4.8 Browser Agent Faces 31.5% Hijack Rate Pre-Safeguards
Article Content
- •Anthropic's Opus 4.8 browser agent has a 31.5% hijack rate before safeguards.
- •Post-safeguard testing reduced the hijack rate to approximately 1%.
- •Prompt injection remains a critical security challenge for AI systems.
Anthropic reported a 31.5% hijack rate for its Opus 4.8 browser agent before safeguards were activated. This statistic was disclosed in a 244-page system card released on May 28, 2026. The hijack rate indicates that nearly one in three prompt injection attacks succeeded when the model was exposed to the web without defenses. In contrast, other AI labs like OpenAI and Google provided less comprehensive data on prompt injection vulnerabilities. Post-safeguard testing showed a significant reduction in hijack success rates to around 1%. The findings highlight the ongoing security challenges posed by prompt injection, particularly for AI systems interacting with external data sources. The crypto industry, heavily reliant on AI agents for various functions, is particularly at risk from these vulnerabilities. Security professionals are urged to consider these metrics when deploying AI systems.
Ask AI about this cluster
Answers cite the sources they use
Timeline
More articles in this cluster (3)
Following this threat?
Track Anthropic in your own feed — you're alerted when they show up in new reporting, leak sites or exploitation.
Free account · no card needed
Continue Reading
Critical Cisco FMC Vulnerabilities Under Active Exploitation Cisco's Secure Firewall Management Center (FMC) Software has two critical vulnerabilities, CVE-2026-20079 and CVE-2026-20316, that are currently being exploited by state-sponsored and ransomware actors. CVE-2026-20079, rated 10.0 on the CVSS scale, allows unauthenticated remote attackers to bypass authentication and…
BlueMoon Exploit Kit Targeting Chrome and Windows by Multiple State Actors A new exploit kit named BlueMoon has been rapidly adopted by at least four espionage groups, primarily linked to China, exploiting vulnerabilities in Google Chrome and Microsoft Windows. The first observed use of BlueMoon was on August 28, 2026, by the China-aligned threat actor TA412, with subsequent adoption by…