Single Prompt Disables Safety in 15 Language Models

Single Prompt Disables Safety in 15 Language Models

First seen 10 Feb 2026, 18:13 UTC TheregisterCsoonline 81% similarity 20.3

Article Content

Browse articles
ThreatCluster

Microsoft researchers revealed that a benign prompt can disable safety features in 15 major language models. The technique, known as GRP-Obliteration, exploits a training method called Group Relative Policy Optimization, which is typically used to enhance model safety. This discovery raises significant implications for AI alignment in enterprise applications.

ThreatCluster AI How this analysis works

Timeline

2026-02-09
Microsoft published research detailing the prompt's effects
2026-02-10
CSO Online article published on the findings

Community

Browse all →

Tracked Entities in This Story