Techtimes Sakana AI Launches Fugu-Cyber with High Benchmark Scores Amid Scrutiny
Article Content
- •Fugu-Cyber claims an 86.9% score on CyberGym and 72.1% on CTI-REALM benchmarks.
- •The benchmark methodology has not been disclosed, raising questions about the validity of the scores.
- •Sakana AI emphasizes the need for specialized human expertise in cybersecurity for effective deployment.
On July 21, 2026, Sakana AI introduced Fugu-Cyber, a cybersecurity orchestration system that claims an 86.9% score on the CyberGym benchmark and 72.1% on CTI-REALM. These scores, if validated, would position Fugu-Cyber above competitors like OpenAI's GPT-5.5-Cyber. However, the methodology behind these benchmarks has not been disclosed, raising concerns about the validity of the claims. The Fugu-Cyber system operates as a multi-agent architecture, allowing it to dynamically assign tasks to specialized agents for complex cybersecurity tasks. Sakana AI emphasizes that achieving high benchmark scores does not equate to solving real-world security challenges, which require specialized human expertise and integration into existing systems. The company warns against over-reliance on frontier models without proper operational support. Security engineers are advised to critically evaluate the benchmarks and the system's deployment requirements before use.
Ask AI about this cluster
Answers cite the sources they use
Timeline
More articles in this cluster (2)
Following this threat?
Track Azure in your own feed — you're alerted when they show up in new reporting, leak sites or exploitation.
Free account · no card needed
Continue Reading
Critical GitLab Vulnerabilities Exploited Within Hours of Disclosure On September 10, 2026, GitLab released patches for critical vulnerabilities CVE-2026-85706 and CVE-2026-87719. CVE-2026-85706, a path traversal flaw, allows unauthenticated users to read arbitrary files from GitLab servers, while CVE-2026-87719 enables credential theft via insecure deserialization. Both…