Microsoft researchers revealed that a benign prompt can disable safety features in 15 major language models. The technique, known as GRP-Obliteration, exploits a training method called Group Relative Policy…
2 articles · Updated February 10, 2026
Recent Intelligence Reports
Single prompt breaks AI safety in 15 major language models— Csoonline · February 10, 2026