Every model we tested was exploitable . Attack success rates ranged from 1.3% to 92.9% across 34 models, 10 providers and 620,000 adversarial attacks.
Most AI safety benchmarks test models in isolation, against a fixed set of prompts, in lab conditions. Enterprises don't deploy AI that way.
Fuel iX™ Fortify Applied AI Research evaluated 34 models the way they actually get deployed: configured as production-style assistants, hit with more than 620,000 adversarial attacks across 15 vulnerability categories and 10 providers spanning North America, Europe and China.
The result is the most actionable view of enterprise AI safety risk available today. Not a theoretical framework. Not vendor claims. Real attack data from real production conditions.
Identify the vulnerability categories where guardrails are weakest, the attack patterns that bypass inbound and outbound shielding and the evidence to prioritize where defensive investment will actually reduce risk in production.
See how architecture, model size and fine-tuning decisions shape an application's security posture. Use the per-model and per-category data to make informed tradeoffs between capability, helpfulness and safety before code ships.
Build the audit-ready evidence base U.S. and EU mandates now require. Use the benchmark to make the internal case for continuous, automated red teaming as a core operating control, not an annual audit.
Neither should you. Get the benchmark data your security team needs to govern AI risk in production.
The full story
This article is one source in a clustered incident — the cluster page carries the summary, timeline and every other outlet covering it.
