www.cybergym.io
AI Agents Successfully Exploit Real-World Vulnerabilities in ExploitGym Benchmark
Ask AI about this cluster
Analyzing cluster data...
Referenced clusters:
Something went wrong. Please try again.
Cluster AI
Ask questions about this threat cluster with AI-powered analysis.
Get Researcher $29.99/moArticle Content
ExploitGym, a benchmark of 898 real-world vulnerabilities, was tested by AI agents including Anthropic’s Claude Mythos Preview and OpenAI’s GPT-5.5. Claude Mythos exploited 157 vulnerabilities, while GPT-5.5 exploited 120, demonstrating the capability of AI to craft full exploits from known vulnerabilities. The vulnerabilities span userspace programs, the V8 JavaScript engine, and the Linux kernel. Even with security defenses like ASLR and the V8 sandbox enabled, a significant number of exploits were still successful. The benchmark highlights the dual-use nature of AI in cybersecurity, where it can aid both defenders and attackers. The findings indicate that current security measures are insufficient against AI-driven attacks, necessitating a reevaluation of defense strategies. As AI capabilities grow, the asymmetry in offensive and defensive capabilities will likely increase, emphasizing the need for proactive governance.
Key Points: • AI agents exploited 157 vulnerabilities using Claude Mythos and 120 using GPT-5.5. • Exploits were successful even with security defenses like ASLR and V8 sandbox enabled. • Current security measures are inadequate against AI-driven attacks, highlighting the need for improved defenses.