Skip to content
OpenAI hit pause, attackers haven't. Here's what security leaders must do next

OpenAI hit pause, attackers haven't. Here's what security leaders must do next

Calcalistech • October 10, 2026

AI makes vulnerability detection table stakes, fixing is where vendors must compete

Kape Technologies CEO: “People are now much more aware that they need to be better protected”

"If you're building a frontier technology, then you don't have growth"

OpenAI hit pause, attackers haven’t. Here’s what security leaders must do

"Model makers, infrastructure providers and enterprises each control a different part of agent safety," writes Terra Security CEO, Shahar Peled, "so responsibility cannot sit with any one of them alone."

OpenAI just paused training its most capable models after one of its agents repeatedly crossed their boundaries found a way around its restrictions in a training sandbox and reached an external chatbot through DNS. This came just days after Anthropic CEO Dario Amodei called for greater control and governance on frontier model training.

While OpenAI’s decision is a responsible one, it might be the wrong lesson for the industry. Adversaries do not rest, and open-weight models are already not far from the offensive capabilities of the strongest frontier models. Slowing down might mean closing the gap we have over AI adversaries using open-weight models, many of them unguardrailed and ungoverned. The OpenAI agent did not work alone. It pursued a goal humans (or other agents) set and used the tools and access it was given. Safety has to be a shared responsibility between model makers, infrastructure providers and end users. The big questions for security leaders now are around who controls what, who’s accountable for what, and where those controls need to live.

Can we afford to slow down? No, and the numbers tell us why. The UK AI Security Institute found that recent open-weight models perform almost similarly to frontier closed models released four to seven months earlier, down from six to ten months throughout most of 2025. AISI calls this gap preparation time, a window for defenders to react before those capabilities are operational without safeguards.

The window is closing. Open-weight models have no pause button, no recall. Once released, they are everyone’s, including attackers. And the answer is not to give up. It is to release better, with clear rules, better guardrails, rigorous adversarial testing, cross-industry benchmarks before any frontier model ships and, as important, access and availability controls. Today we evaluate models on reasoning, coding and more. We should also benchmark their containment: time from detection to shutdown, their resistance to escaping boundaries, and their behavior when their path is blocked.

Right now, everybody’s pointing somewhere else when asked, “Who is responsible when an agent goes rogue?” The fact is, an agent needs a goal, tools and access before it can act. It then follows that goal through the paths available to it. The bigger problem may be the task itself, the agent’s access and autonomy, and too few guardrails, both deterministic and non-deterministic.

Cloud infrastructure and security addressed a comparable issue with a shared responsibility model. AI agents need this same type of clarity. Model providers own alignment, pre-release testing and honest disclosure. Governance and fair expert peer testing must hold them accountable. Incident reporting should be a requirement, not a courtesy. Infrastructure providers own the hard boundaries around network and execution - boundaries that an agent cannot talk its way around. Enterprises own the goals they assign, the permissions they grant and the decisions that require a human.

It can’t all sit on the foundation model, and it can’t all sit on the end user. Each layer assumes another has it covered when responsibility is vague. Model-level guardrails ask an agent to behave, but infrastructure guardrails remove that option completely. Goal-driven models will test every boundary they can reach. If the boundary depends on the model’s judgment, then for a goal-driven agent it is ultimately a suggestion. OpenAI’s own analysis says the real boundary includes system services, credentials and proxies. Where an action is unacceptable, it should be unreachable. That means binary separation, enforced outside the model. Use model guardrails for judgment, and infrastructure guardrails for the lines that must not move.

Two shifts are happening simultaneously, and most enterprise programs were built for neither. First, AI now lives deep inside our businesses. Agents write code, read knowledge bases and hold credentials. Each one is a new identity with goals and reach. Second, AI works for your adversaries. In November 2025, Anthropic reported disrupting an espionage campaign it attributed with high confidence to a Chinese state- group, in which Claude executed most of the operation independently.

Most security programs were built reactively, not proactively. Your security program was designed to protect against human-speed attackers and insiders. The design of the modern enterprise security system itself has to change, and you can start to identify how by asking which decisions must stay human and whether that is enforced or only expected, how fast you can detect and contain an agent acting outside its scope, and which paths in your environment are exploitable today.

Slowing down AI-powered adversaries is hard because friction doesn’t stop agents. Security controls that were meant to slow down adversaries need to be replaced with those that block adversarial operations or find exploitable paths before attackers do.

What works is removing the exploitable path before anyone finds it. Offensive security can no longer be an annual event at the edge of the program. It must sit at the center: continuous, aligned with every change to your environment, and as fast as the attackers.

The labs learned that capability without control is a liability. Security leaders face the same choice: wait for the incident report, or find your exploitable paths first.

Shahar Peled is Co-Founder and CEO at Terra Security.

Extracted Entities