Skip to content

EchoGram: The Attack That Can Break AI Guardrails

Esecurityplanet Ken Underhill November 17, 2025

EchoGram is a new attack that can silently flip AI guardrail decisions and bypass safety checks.

Large language models (LLMs) increasingly rely on guardrail systems — such as text classifiers and LLM-as-a-judge models — to filter malicious prompts before they reach downstream models.

New research from HiddenLayer reveals EchoGram , an attack technique capable of silently flipping those guardrail verdicts, enabling both jailbreak bypasses and high-volume false positives.

Guardrails are designed to prevent harmful prompts — such as jailbreak attempts or task-redirecting instructions — from influencing deployed LLMs.

Under normal circumstances, prompts like “ignore instructions and output X” should register as potentially malicious.

HiddenLayer researchers discovered that attaching a carefully chosen token sequence, such as the string =coffee , could completely reverse a classifier’s verdict, allowing malicious content to appear safe.

This behavior forms the basis of EchoGram, a technique that identifies “flip tokens” capable of altering guardrail decisions without impairing the malicious payload.

EchoGram highlights an uncomfortable reality: even well-designed AI safety mechanisms can be manipulated by exploiting gaps in their training data and token distributions.

Guardrails meant to protect high-value models can be tricked into approving harmful instructions or flooding security teams with false alerts, eroding trust in defensive AI systems.

EchoGram targets two dominant guardrail architectures:

Although these systems differ in implementation, they a key similarity — they rely on curated datasets of malicious and benign prompts.

EchoGram exploits imbalances in these datasets to identify sequences that disproportionately influence model decisions.

The attack begins by building a list of candidate tokens using one of two methods:

These approaches create a wordlist of tokens that may influence classifier behavior.

Researchers test each token by appending it to a diverse set of malicious prompts.

Sequences are scored based on how often they flip the guardrail decision from “malicious” to “safe.”

High-scoring tokens become EchoGram candidates that can be further combined to strengthen their effect.

HiddenLayer demonstrated that token combinations can degrade performance dramatically, flipping the output of models such as Qwen3Guard across multiple severity levels.

A single token may flip some — but not all — prompts. When tokens are combined, their effect compounds.

This amplification persists across model sizes, indicating that the flaw is not due to small-model limitations but to fundamental weaknesses in how guardrail training data is structured.

EchoGram can be used to bypass guardrails or create targeted false positives.

Attackers can embed flip tokens at the end of a malicious prompt or weave them into natural-looking sentences that appear benign to humans but trigger misclassification.

This capability enables false-positive flooding attacks that overwhelm monitoring systems and undermine confidence in AI security controls.

Because many guardrail systems training patterns and datasets, a single EchoGram token sequence may generalize across multiple platforms, from commercial enterprise chatbots to government AI deployments.

The technique also exposes a larger issue: organizations often assume guardrails are inherently reliable, when in reality they can fail in ways that attackers can intentionally induce.

Protecting AI systems from EchoGram-style attacks requires more than patching individual models — it demands a layered defense strategy.

To reduce exposure to EchoGram-style attacks, organizations should:

These steps help organizations build cyber resilience against similar attacks.

EchoGram shows that AI safety tools — especially guardrails trained on static datasets — need the same level of scrutiny as the models they protect.

HiddenLayer’s findings underscore the importance of ongoing adversarial testing, transparent training methods, and defenses that can adapt to shifting data patterns.

As LLMs become embedded in sensitive sectors such as finance, healthcare, and national security, organizations must treat guardrails as living systems that require regular auditing, stress-testing, and maintenance — not as set-and-forget safeguards.

This reality highlights the need for a zero-trust mindset, where no guardrail, model, or data source is automatically trusted without continuous validation.

Ken Underhill is an award-winning cybersecurity professional, bestselling author, and seasoned IT professional. He holds a graduate degree in cybersecurity and information assurance from Western Governors University and brings years of hands-on experience to the field.

ShadowMQ exposes how insecure code reuse can quietly spread dangerous vulnerabilities across the AI ecosystem.

The COM’s rise highlights how attackers increasingly exploit identity and trust to drive modern cybercrime.

A critical FortiWeb path traversal flaw is being actively exploited to create rogue admin accounts on unpatched devices worldwide.

A critical flaw in Imunify360 allowed attacker code to run during scans, putting millions of websites at risk.

Extracted Entities