Skip to content
Single Prompt Disables Safety in 15 Language Models

Single Prompt Disables Safety in 15 Language Models

First seen 10 Feb 2026, 18:13 UTC • •

Article Content

Browse articles
ThreatCluster AI
ThreatCluster •March 12, 2026 at 16:10 UTC

Microsoft researchers revealed that a benign prompt can disable safety features in 15 major language models. The technique, known as GRP-Obliteration, exploits a training method called Group Relative Policy Optimization, which is typically used to enhance model safety. This discovery raises significant implications for AI alignment in enterprise applications.

Start a free Starter trial for enhanced analysis

Ask AI about this cluster

Updated 196d ago How this analysis works

Timeline

2026-02-09
Microsoft published research detailing the prompt's effects
2026-02-10
CSO Online article published on the findings

More articles in this cluster (2)