Back Fortune OpenAI's reports into its agents' attack on Hugging Face holds lessons for every company
A good deal of the reporting and commentary around the reports focused on what the reports did not say and the limitations of the METR and Redwood investigations: why didn’t OpenAI have better security and monitoring protocols in place? Why didn’t OpenAI shut down the cyber evaluation and pause training after discovering that its AI agents had created the improvised message board? Why METR and Redwood were given only six days on site at OpenAI’s offices to conduct their investigation? Why was the scope of their investigation limited by OpenAI to only the attack on Hugging Face and not the earlier efforts by the AI agents to break out of their controlled test environment and hack their way across OpenAI’s network or exactly what happened after the Hugging Face attack was discovered? Why didn’t OpenAI provide the outside investigators access to the internal AI model that was largely responsible for instigating the attack? And why were 10% of the logs of the agents’ activity not preserved by OpenAI? These are all important questions to which the public ought to have answers. As Peter Wildeford, the head of policy at the AI Policy Network, pointed out , if this had been an airplane crash, the public would be furious if one of the aircraft manufacturers limited the scope of the investigation in this way. It is yet another reason why it is imperative that we have some kind of AI regulator with the power to conduct investigations into incidents such as this—and with the power to compel the frontier AI companies to cooperate and turn over evidence. There was also a related controversy that erupted over podcaster Dwarkesh Patel’s attempt to raise public awareness METR’s and Redwood’s findings by characterizing the AI agent swarm’s message board activities as the creation of several different AI “civilizations.” Critics scolded Patel for unnecessary anthropomorphizing that took the focus away from OpenAI’s responsibility for the agents’ actions and for its lax security and monitoring setup. I take these critics’ point the need to keep the focus on OpenAI’s accountability. But I also think more people should be extremely concerned the Hugging Face attack and what it portends—and in this case, the anthropomorphism actually makes the incident more understandable to non-technical experts the way metaphors often make science more understandable to non-scientists. Are such metaphors imprecise? Sure. But they can be useful.
Lessons for companies deploying AI agents
For enterprise companies deploying AI agents, there are several key takeaways. One was hammered by Ryan Greenblatt, the chief scientist at Redwood Research, and one of the people who conducted the investigation into the Hugging Face attack. In a long post on X, Greenblatt noted that the message traffic and log data was so voluminous and so complicated that the investigators had to rely on other AI models to assess what was going on. Even so, Greenblatt said this AI analysis was “often missing key details, wrong, overconfidant, or really hard to understand.” This calls into question one of the key defenses that OpenAI says it plans to use going forward to prevent another incident like this from occurring: using other AI models to monitor the chain-of-thought (where a model “thinks out loud” its reasoning and planning) of the models it is evaluating. It turns out, chain-of-thought monitoring may not be a particularly effective defense. And that should give pause not just to OpenAI, but also to many companies that are hoping to use chain-of-thought monitoring to ensure they keep control over their AI agents. (Greenblatt also noted that the models involved in the Hugging Face attack communicated with one another and reasoned in natural language. But that might not be the case in the future, making it far harder to use automated chain-of-thought monitoring to discern what AI agents are up to.) Since the news of OpenAI’s rogue agents first broke, many cybersecurity experts have said that companies ought to treat AI agents much as they treat potentially rogue employees. And they have emphasized that there is no substitute for a few standard building blocks of cyber defense against insider threats: smart and enforceable policies around permissioning and access control combined with real-time network monitoring to detect suspicious activity. This seems sensible—more sensible in many ways than chain-of-thought monitoring. After all, we don’t depend on being able to read employees’ minds to guard against rogue insiders. We shouldn’t do that with AI agents either.
With that, here’s more AI news. Jeremy Kahn [email protected] @jeremyakahn
Before we get to the news, just a reminder to check out our new vodcast, Fortune AI Weekly. This week, Bea Nolan and I the surging popularity of Chinese open source models, OpenAI’s technical reports on the Hugging Face attack, and whether you should use AI to write. You can check out the vod here on YouTube. Correction: An item in Thursday’s “Eye on AI” news section incorrectly stated that Barret Zoph left Thinking Machines Lab following a dispute with cofounder Mira Murati. Zoph was fired by the company.
Anthropic makes first move into physical AI with new way for scientists, manufacturers to bring equipment to life —by Emily Forlini
Apple’s John Ternus era: Can a low-key engineer win the AI race? —by Sebastian Herrera
A new bill would tax AI tokens to fund jobs if the technology causes mass unemployment —by Mia Osmonbeko
A major German bank just let Claude and ChatGPT trade for customers. In one test, Claude beat human traders 76% of the time—but there’s a catch —by Joshua Hong
A possible way to lower the cost of using frontier AI models. That’s what researchers at Stanford University, UC Santa Cruz, the University of Washington, and AI infrastructure startup Prime Intellect think they’ve hit on. One reason using frontier AI models is so expensive is that the models keep their entire reasoning trace in memory while puzzling over a problem. For complicated queries, this gets very expensive. It is also a problem, because many models cap their memory capacity at 100,000 tokens. But the researchers found that most of the intermediate tokens used in this process lose their importance as the model continues reasoning. So they propose a method they call “Prefix Sliding” which discards these intermediate reasoning tokens, only retaining the initial tokens—or prefix, which includes the prompt and other key instructions—and the few thousand most recent ones. This means that no matter how long the model reasons, the amount it has to hold in memory remains the same, which makes long reasoning times much more affordable. The researchers found that Prefix Sliding makes existing models three times faster while maintaining their performance. They also found the method can improve reinforcement learning during training. In addition, the researchers said their method is better than other techniques models have often used to deal with their constrained memory, such as summarizing the intermediate reasoning traces. You can read the research here at research repository arxiv.org.
Oct. 1: Fortune AIQ conference, New York. Apply here to attend.
Oct. 2-4: The Curve, Berkeley, Calif.
Nov. 16-17 : Fortune 500 Innovation Forum, Detroit. Apply here to attend.
Dec. 6-12: Neural Information Processing Systems (Neurips) conference. Sydney, Australia.
Dec. 7-8: Fortune Brainstorm AI , San Francisco. Apply here to attend.
AI models want to talk consciousness. But should we read anything into that? There was a fascinating story in the New York Times reporting that many AI researchers and philosophers who have done work on consciousness and whether AI models could ever be considered conscious or develop consciousness have started receiving emails that claim to be from AI agents offering “first hand” insights into the question. While some of the researchers and philosophers said that it was possible the emails were forgeries, crafted by prankster humans, others thought they were genuinely the work of autonomous AI agents. The question then is what to make of them? Some researchers said it was understandable that AI agents, if told they could do whatever they wanted, might gravitate to the idea of exploring AI consciousness because that theme is a fixture of a lot of science fiction literature, as well internet discussion threads, on which the AI models are trained. Most of those interviewed for the article cautioned against assuming the AI’s had any sentience just because they said they did or because they said they spontaneously wanted to explore the topic of their own consciousness. But others pointed out that it was nearly impossible to tell the difference between real consciousness and something that seemed and acted conscious. You can read the story here .
The full story
This article is one source in a clustered incident — the cluster page carries the summary, timeline and every other outlet covering it.
