Skip to content
VLoc Bench Reveals Challenges in Vulnerability Localization

VLoc Bench Reveals Challenges in Vulnerability Localization

First seen 7 Oct 2026, 01:58 UTC • •

Article Content

Browse articles
ThreatCluster AI
ThreatCluster •October 7, 2026 at 01:58 UTC
  • •VLoc Bench evaluates AI's ability to find vulnerable code without prior knowledge.
  • •The strongest model only achieved a 0.229 File F1 score, indicating significant challenges.
  • •38.4% of tasks resulted in no correct file identification by any evaluated model.

In October 2026, Cisco released the Vulnerability Localization Benchmark (VLoc Bench) to evaluate AI agents' ability to identify vulnerable code in repositories. The benchmark includes 500 real vulnerabilities from 290 repositories across six ecosystems. Results show that the best-performing model achieved only a 0.229 File F1 score, with 38.4% of tasks yielding no correct file identification. The benchmark emphasizes the difficulty of vulnerability localization, as models struggle to confirm when vulnerabilities have been fixed. The findings indicate a significant gap in current AI capabilities for cybersecurity tasks, highlighting the need for further development in this area.

Start a free Starter trial for enhanced analysis

Ask AI about this cluster

Updated just now How this analysis works

Timeline

2026-10-06
VLoc Bench released
Cisco launched the Vulnerability Localization Benchmark to assess AI agents' performance in locating vulnerabilities in code repositories.
Blogs.Cisco
2026-10-07
VLoc Bench results published
Results from the benchmark show that the best model only achieved a 0.229 File F1 score, revealing significant challenges in vulnerability localization.
cisco-foundation-ai.github.io

More articles in this cluster (4)

Common questions

What is VLoc Bench?
VLoc Bench is a benchmark designed to evaluate AI agents' ability to locate vulnerable code in repositories without prior knowledge.
How effective are current models in vulnerability localization?
Current models struggle significantly, with the best achieving only a 0.229 File F1 score and many tasks yielding no correct identifications.
What does the benchmark reveal about AI capabilities?
The benchmark highlights substantial gaps in AI capabilities for identifying and confirming the remediation of vulnerabilities.