Csoonline Single Prompt Disables Safety in 15 Language Models
Article Content
Browse articles
Microsoft researchers revealed that a benign prompt can disable safety features in 15 major language models. The technique, known as GRP-Obliteration, exploits a training method called Group Relative Policy Optimization, which is typically used to enhance model safety. This discovery raises significant implications for AI alignment in enterprise applications.
Ask AI about this cluster
Answers cite the sources they use
Updated 196d ago How this analysis works
Timeline
2026-02-09
Microsoft published research detailing the prompt's effects
2026-02-10
CSO Online article published on the findings
More articles in this cluster (2)
Continue Reading
Critical Zero-Day Vulnerability in F5 BIG-IP APM Exploited for Remote Code Execution F5 Networks has reported a critical vulnerability in its BIG-IP Access Policy Manager (APM), tracked as CVE-2026-94127, which is being actively exploited in the wild. The flaw allows unauthenticated attackers to execute remote code on systems configured with both an APM access policy and an OAuth profile. This…
Massive Network of AI Proxy Servers Used for Malicious Activities Uncovered Security researchers from Team Cymru have identified over 10,000 proxy servers in China facilitating malicious AI activities. These servers, termed 'transfer stations,' are primarily used to bypass geographic restrictions and conduct model distillation attacks against frontier AI models. The infrastructure allows…