ThreatCluster

Claude Opus 5 Achieves 2% Success Rate Against Indirect Prompt Injection Attacks

First seen 10 Aug 2026, 11:31 UTC CybersecuritynewsGbhackers 93% similarity 19

Article Content

Browse articles
ThreatCluster

Anthropic's Claude Opus 5 has demonstrated a significant reduction in the success rate of indirect prompt injection (IPI) attacks, achieving a mere 2% success rate over 15 attempts in the Gray Swan benchmark. This marks a notable improvement from the 5.5% success rate of its predecessor, Claude Opus 4.8. The results were published in the model's system card, showcasing its enhanced resilience against IPI attacks. The benchmark results indicate that Claude Opus 5 is currently the most resistant model tested, outperforming all previous versions and competing models. This advancement is crucial for organizations relying on AI models for various applications, as it reduces the risk of successful exploitation through IPI methods. The findings reflect ongoing efforts to bolster AI security and mitigate potential vulnerabilities in machine learning systems.

Key Points: • Claude Opus 5 achieves a 2% success rate against indirect prompt injection attacks. • This represents a significant improvement from the 5.5% success rate of Claude Opus 4.8. • The results were validated in the Gray Swan benchmark, confirming Opus 5's superior resilience.

ThreatCluster AI How this analysis works

Timeline

2026-08-10
Claude Opus 5 benchmark results released
Anthropic published results showing a 2% success rate for indirect prompt injection attacks in Claude Opus 5, down from 5.5% in Opus 4.8.
Cybersecuritynews
2026-08-10
Gray Swan benchmark analysis published
The Gray Swan benchmark analysis confirmed Claude Opus 5's position as the most resistant model against IPI attacks.
Gbhackers

Community

Browse all →