Endorlabs
Benchmarking Claude Fable 5 Reveals Harness Impact on Security Outcomes
Ask AI about this cluster
Analyzing cluster data...
Referenced clusters:
Something went wrong. Please try again.
Cluster AI
Ask questions about this threat cluster with AI-powered analysis.
Get Researcher $29.99/moArticle Content
Endorlabs benchmarked the Claude Fable 5 AI model using two different harnesses, revealing significant differences in security outcomes. The Cursor harness achieved a 72.6% FuncPass and 29% SecPass, while the Claude Code harness resulted in a 59.8% FuncPass and 19% SecPass. The results indicate that the agent harness has a more substantial impact on security outcomes than the model itself. Despite the improvements, the security scores remain below 30%, indicating that many vulnerabilities are still left unaddressed. This benchmarking exercise involved 200 real-world vulnerability-fixing tasks in actual projects, highlighting the importance of the agent scaffolding in achieving better security results. The findings prompt further investigation into the relationship between model capabilities and harness effectiveness.
Key Points: • Cursor harness with Claude Fable 5 achieved 72.6% FuncPass and 29% SecPass. • Claude Code harness resulted in lower scores: 59.8% FuncPass and 19% SecPass. • The agent harness significantly influences security outcomes more than the model itself.