Csoonline
Single Prompt Disables Safety in 15 Language Models
First seen 10 Feb 2026, 18:13 UTC
•
•81% similarity
•20.3
Share:
Export
Ask AI about this cluster
Analyzing cluster data...
Referenced clusters:
Something went wrong. Please try again.
Cluster AI
Ask questions about this threat cluster with AI-powered analysis.
Get Researcher $29.99/moArticle Content
Browse articles
Microsoft researchers revealed that a benign prompt can disable safety features in 15 major language models. The technique, known as GRP-Obliteration, exploits a training method called Group Relative Policy Optimization, which is typically used to enhance model safety. This discovery raises significant implications for AI alignment in enterprise applications.
ThreatCluster AI
How this analysis works
Timeline
2026-02-09
Microsoft published research detailing the prompt's effects
2026-02-10
CSO Online article published on the findings