New AI Security Method Analyzes Internal Activation Patterns for Enhanced Defense

New AI Security Method Analyzes Internal Activation Patterns for Enhanced Defense

First seen 29 Jul 2026, 17:15 UTC DarkreadingFeeds.Feedburner 87% similarity 21.9

Article Content

Browse articles
ThreatCluster

Researchers are developing a new AI security approach that examines the internal workings of large language models (LLMs) rather than just their inputs and outputs. This model-agnostic method focuses on activation analysis and uses standardized rules to identify cognitive elements (CEs) that can indicate malicious intent. The approach aims to detect specific threats like phishing attacks by combining CEs into logical statements. This system, named GAVEL (Governance via Activation-based Verification and Extensible Logic), is designed to be language-independent, making it more robust against evasion tactics. The research will be presented at Black Hat USA 2026, highlighting its potential to enhance AI safety. Current defenses often fail against sophisticated attacks that modify prompts to bypass filters. This new method represents a significant advancement in AI security, offering an additional layer of protection against misuse.

Key Points: • Researchers are developing a model-agnostic AI security approach focused on internal activation patterns. • The new method uses cognitive elements to detect specific threats like phishing attacks. • GAVEL aims to provide a robust defense against malicious use of AI systems.

ThreatCluster AI How this analysis works

Timeline

2026-07-28
Research on new AI security approach announced
Researchers proposed a novel method for analyzing AI model internals to enhance security against malicious use.
Darkreading
2026-07-29
New AI security approach reported
The GAVEL method focuses on activation analysis and is set to be presented at Black Hat USA 2026.
Feeds.Feedburner

Community

Browse all →

Tracked Entities in This Story