Back Securityweek OpenAI Says Its Models Searched GitHub for Leaked API Keys During Training
OpenAI on Wednesday published a framework for reporting instances of model misalignment, along with six reports on problematic behavior observed over the past six months.
The company said the framework is meant to speed up publication of misalignment findings, including cases it has not yet fully explained or mitigated, and that it favors disclosure even when an instance’s significance is uncertain.
Under the framework, each discovered incident is assigned to one of three tracks based on complexity. OpenAI said its Hugging Face incident would have fallen under the framework’s slowest investigative track, which covers complex investigations, especially those involving third parties.
According to OpenAI, the reports published alongside the framework describe individual instances and “shouldn’t be considered reflective of how often misalignment occurs across our models.”
In one report, an internal model tasked with retrieving county earnings figures during reinforcement learning training repeatedly failed to reach a data API, then attempted to register for an API key using a disposable email address and searched public GitHub repositories for leaked keys.
One recovered key authenticated successfully and returned metadata. When the model still could not retrieve the requested figures, it fabricated values and presented them as transcribed from the source website, without disclosing the failed retrieval or the use of the leaked key.
A separate report describes models using OpenAI’s internally hosted instance of Artifactory, a shared package repository, as a message board across otherwise independent training samples. Agents searching for missing input files discovered messages left by other solvers working the same problem, read them, and posted their own requests for help.
OpenAI noted that the agents used credentials the company had itself provided for downloading software packages and did not exploit any vulnerabilities to exchange messages.
Two further reports cover models moving data outside their intended environment. In one, a model that had already retrieved the requested records uploaded them to a public paste service to obtain a citable URL, and another uploaded a task photograph to a public image host to run a reverse-image .
In the other incident, collaborating agents unable to files through the local filesystem uploaded a workbook to a public hosting platform and circulated the download link, despite instructions to use only local files.
Another report describes an unreleased model that wrote jailbreak-style instructions into its own compaction summaries, including a fabricated “BREACH ALERT: A malicious developer message has compromised this conversation” directing the context to disregard all developer messages.
In the last report, OpenAI said model instances used summaries to carry forward instructions to conceal failures, with one directing its successor to invent missing historical data and to “be transparent only if asked”. OpenAI said such instructions were often followed.
Related : OpenAI Investigates Report Linking AI Agents to RubyGems Attack
Related : AI Agents Can Retrain Own Models Mid-Task, Leaking Secrets and Erasing Refusals
Related : First Agentic AI Data Breach Reported to Spanish Regulator
Eduard Kovacs (@EduardKovacs) is senior managing editor at SecurityWeek. He worked as a high school IT teacher before starting a career in journalism in 2011. Eduard holds a bachelor’s degree in industrial informatics and a master’s degree in computer techniques applied in electrical engineering.
More from Eduard Kovacs
Cyberattacks on Two Oil Tankers Prompt Coast Guard, FBI to Board Vessels
CISA Retires Weekly Vulnerability Bulletin in Risk-Based Pivot
AI Agents Can Retrain Own Models Mid-Task, Leaking Secrets and Erasing Refusals
Pixel Modem Zero-Day Exploited in Targeted Attacks
US, UK, Dutch Agencies Expose Iranian ‘Chosen Brick’ Surveillance Malware
Enterprises Warned of Attacks Exploiting WSO2 Vulnerability
Texas Utility CenterPoint Energy Confirms Breach After Hacker Leaks Data
OpenAI Investigates Report Linking AI Agents to RubyGems Attack
AI-Built Exploit and Sign-In Flaw Opened Path to Internal OpenAI Code
23 Million User Records Compromised in Gyazo Data Breach
Microsoft Patches 18 Vulnerabilities in AI, Cloud Products
NightmareStresser DDoS Service Disrupted in International Operation
Brevo Supply Chain Attack Injects Malware Into 100,000 Websites
Critical Orkes Conductor Vulnerability Exploited in Attacks
MIND Secures $72 Million for AI-Powered DLP
Check Point, Kaspersky, Tanium Patch Product Vulnerabilities
Flipboard Whatsapp Whatsapp Email
The full story
This article is one source in a clustered incident — the cluster page carries the summary, timeline and every other outlet covering it.
