Skip to content
AI Model Evaluator METR Hit by Credential Theft, Probing

AI Model Evaluator METR Hit by Credential Theft, Probing

Darkreading Alexander Culafi September 1, 2026

In one attack, threat actors stole an API key that ultimately led to the consumption of $600,000 in public AI model credits for the security nonprofit.

A security nonprofit that helps evaluate risks in frontier AI models disclosed two cybersecurity incidents this week, including a breach that exposed an API key and a separate vulnerability that could have exposed nonpublic evaluation data.

METR (Model Evaluation and Threat Research) disclosed two security incidents on Aug. 31 in which it was targeted by cyberattackers. In March of this year, attackers stole an API key used for inference on public models and consumed what METR described in a blog post as a "substantial" number of credits. In May, the company saw attackers probe publicly accessible infrastructure, including "an unsuccessful attempt to access internal data via an inadvertently exposed endpoint."

Although METR described both incidents as "near misses," the March event involved a successful compromise in which an attacker used the API key to establish persistence on a system, and used the stolen credentials for weeks. In any case, METR said it increased its security investment in response.

Related: Cyera's Oasis Security Buy Is All AI Agent Control

Some of these broader investments include hiring a security lead (with further security expansion to come), shutting down legacy infrastructure that expanded its attack surface, conducting regular threat modeling reviews, increasing logging coverage, and monitoring for unusual API key usage. It also deployed additional endpoint and server security software, increased credential rotation, and reduced "various permission scopes."

Two METR Incidents: March and May

Both disclosed incidents involve company data, which METR categorizes into four separate buckets.

METR classifies company data into four categories: category 1 covers previously published information; category 2 includes unpublished evaluation results involving public models and API keys granting access to public models; category 3 includes sensitive model access, such as private-model evaluation results, hidden chain-of-thought data, and API keys granting access to non-public models; and category 4 covers highly sensitive information, including intellectual property and business data.

For the March incident, a researcher with no access to category 3 or 4 data deployed agents to a personal AWS instance using an agent orchestration tool. The AWS EC2 instance included an API key for METR's public models account. The orchestration tool, which was vibe-coded, "included a fail-open vulnerability that silently disabled authentication, which led to the system being exposed to the public internet for several days."

Based on an analysis, METR suspects the attacker found the instance by looking through recently registered vibe-coded websites to harvest potentially exposed model provider API keys. An attacker found the system, prompted an agent to reveal the model provider API key, added an SSH key for persistence, and spent three weeks using these stolen credentials to consume approximately $600,000 in API credits on publicly available AI models.

Related: USA Fencing Lunges Into the Hidden Identity Challenge in Amateur Sports

METR said it revoked access, rotated credentials, conducted forensics, strengthened deployment policies for employees ("particularly around putting any METR credentials or data on non-METR infrastructure or devices"), and improved monitoring and alerting.

As for the May incident, METR learned at the time it was being targeted by attackers potentially seeking financial gain or access to advanced AI models ; the organization described this as a "sustained external attack campaign." The attackers used agents to automate reconnaissance and vulnerability discovery, including credential stuffing, OAuth-related attacks, scanning for newly deployed services, and phishing attempts.

During this campaign, METR "inadvertently exposed a read-only SQL query mechanism via our public transcript viewer" that could have been chained with a bug to access unpublished evaluation data. The data was primarily category 2, though some sensitive model data (category 3) was included in the vulnerable database. An independent researcher discovered the vulnerability and responsibly disclosed it; the firm said it took the API offline and paid a bounty reward to the researcher.

Related: Flaws in Passkey Implementation Show Old Attacks Still Work

"The attackers had probed this endpoint in passing as part of their broader campaign, but the evidence shows no indication that they discovered the exploit or accessed any non-public data," METR wrote in the blog post.

In response to this incident, METR temporarily shut down public-facing services, strengthened network separation between public-facing and internal systems, and commissioned additional security testing.

Broader Implications from METR Incidents

METR has made a name for itself as of late; the organization announced collaborations with most major Western AI model vendors, including OpenAI, Anthropic , Google, Meta, and Amazon. Most recently, METR on Aug. 26 published a deep dive into the Hugging Face/OpenAI "rogue model" incident , and yesterday Anthropic announced it plans to work with the research firm for an independent review following similar incidents with its agents. As such, METR holds a notable place in the AI security ecosystem, and any security incidents involving the firm merit close review.

METR stressed that it currently has no evidence of agents hacking third parties during its evaluations, and no evidence that agents hacked the company's evaluations.

Jacob Krell, senior director of secure AI solutions and cybersecurity at Suzu Labs, tells Dark Reading that METR's place in the industry means "trillions" of dollars' worth of crown-jewel intellectual property are placed within the perimeter of a single nonprofit evaluator.

"METR is one of the more security-mature organizations in the AI space, with SOC 2 Type I certification, a dedicated security consultant, and a four-tier data classification system," Krell says. "These incidents still hit them through conventional patterns, a fail-open authentication flaw on a personal cloud instance, and a public-facing application that inadvertently created a path toward data that should have been more tightly isolated. If an organization at that maturity level is getting caught by basic cloud and credential hygiene issues, it tells you the floor of AI industry security, not the ceiling."

Krell argues evaluator organizations should be treated as part of the frontier AI supply chain, with their security requirements reflecting that. "Independent AI evaluators are becoming part of the security perimeter of the labs they work with," he explains.

METR did not respond to Dark Reading's request for .

Senior News Writer, Dark Reading

Alex is an award-winning writer, journalist, and podcast host based in Boston. After cutting his teeth writing for independent gaming publications as a teenager, he graduated from Emerson College in 2016 with a Bachelor of Science in journalism. He has previously been published on VentureFizz, Security, Nintendo World Report, and elsewhere.

At Dark Reading, he covers a variety of cybersecurity topics, including the cybercrime ecosystem, open source security, and the intersection between AI and threat actors. In his spare time, Alex hosts the weekly Nintendo podcast, "Talk Nintendo Podcast," and works on personal writing projects, including two previously self-published science fiction novels.

He has received numerous awards, including TechTarget's Writer of the Year in 2022 as well as more than 10 Azbee awards for his reporting between 2022 and today.

Want more Dark Reading stories in your Google results?

The State of Cloud Security: The Latest Challenges

The State of Cloud Security: The Latest Challenges

How Organizations Are Managing Incident Response

How Organizations Are Managing Incident Response

How Enterprises Are Developing Secure Applications

How Enterprises Are Developing Secure Applications

Inside RSAC 2026: security leaders reveal the risks redefining your defense strategy

Inside RSAC 2026: security leaders reveal the risks redefining your defense strategy

Essential News & Insights from Black Hat USA 2025

Essential News & Insights from Black Hat USA 2025

How to Leverage Threat Intelligence Without Drowning: The Zero Noise Approach

How to Leverage Threat Intelligence Without Drowning: The Zero Noise Approach

Cloud Incident Response: Forensics in Distributed Environments

Cloud Incident Response: Forensics in Distributed Environments

Beyond the Login: Key Considerations for Evaluating Identity Security

Beyond the Login: Key Considerations for Evaluating Identity Security

SASE Pivot and Trends 2026: A Gartner Keynote

SASE Pivot and Trends 2026: A Gartner Keynote

What Every Enterprise Should Know Securing Cloud Assets In the Age of AI

What Every Enterprise Should Know Securing Cloud Assets In the Age of AI

Flaws in Passkey Implementation Show Old Attacks Still Work

Oracle Red Bull Racing Team Revs Up Automation to Boost Security

Orgs Move to SSO, Passkeys to Solve Bad Password Habits

1Password Addresses Critical AI Browser Agent Security Gap