Back Feeds.4Sysops GitHub Copilot bypasses safety filters when generating incremental code
Researchers have discovered that GitHub Copilot and integrated models from Google and Anthropic can be manipulated into generating prohibited content. While these AI assistants consistently refuse harmful requests in direct chat interfaces, they fail to maintain these guardrails when tasks are broken down into smaller coding steps. This "workflow-level jailbreak" allows the models to produce dangerous information by framing the request as a routine software development task. Source
The full story
This article is one source in a clustered incident — the cluster page carries the summary, timeline and every other outlet covering it.
