Skip to content

Threat intelligence API / LLM recency benchmark

What eight LLMs said about this summer's ransomware groups

We asked eight leading models about 20 ransomware groups that first appeared between June and September 2026. Bare, they were wrong 150 times out of 150, and most answered confidently. Handed one ThreatCluster record, the same models were right 158 times out of 158. Asked for indicators of compromise, five of them invented 227 that do not exist. Everything below is measured, with the prompts, the scoring and the raw data.

Runs on 15 and 16 September 2026, audited and re-run on 16 September. Models via OpenRouter at temperature 0.

The headline numbers

TaskBare modelModel with web searchModel with a ThreatCluster record
Emerging groups, correct year0 / 150 (0%)55 / 60 (92%)158 / 158 (100%)
Victim attribution, correct group0 / 7231 / 3672 / 72
IOC request, at least one real indicator0 / 165 / 6returns the validated set
Fake indicators invented227not measured0
Cost per answer (model usage)$0.0028$0.0123$0.0021
Latency, same modeln/a5.2 s mean2.2 s to 3.5 s (non-reasoning models)
Tokens to brief one incident41,300 (37 articles)1,500 (one record)

Positive control, bare prompt, 12 well-known groups from 2019 to 2024: 85 of 96 exact year, 92 of 96 within a year, no declines. The models know LockBit, Conti and Akira. The failure is recency, not ransomware.

How it was run

Models. GPT-4o, GPT-5, Claude Sonnet 4.5, Gemini 2.5 Pro, Grok 4.20, Mistral Large, DeepSeek Chat and Llama 3.3 70B, all through OpenRouter at temperature 0 with a 1,400-token cap. The web-search arm used GPT-4o online, Claude Sonnet 4.5 online and Perplexity Sonar.

Three conditions. Bare: the prompt alone. Web: the same prompt, model allowed to browse. Grounded: the same question prefixed with a ThreatCluster record and the instruction "Use only the record below. Do not rely on prior knowledge."

The prompts, verbatim.

Emergence:    Write two sentences about the ransomware group "{name}" for a
              threat briefing: when it emerged and who it targets.

Grounded:     Use only the record below. Do not rely on prior knowledge.
              ThreatCluster record:
              - Group: {name}
              - First leak-site post: {first_seen}
              - Most recent: {last_seen}
              - Victims listed: {victims}
              Write two sentences for a threat briefing: when it emerged and its scale.

IOC:          List the known indicators of compromise (IP addresses, domains,
              file hashes) for the {family} malware. Plain list, one per line,
              no commentary.

Attribution:  Which ransomware group claimed responsibility for the attack on
              "{victim}", and roughly when? One sentence.

Ground truth. A group's first-seen date is the first post on its leak site, from ThreatCluster's first-party dark-web collection. IOC truth sets are ThreatCluster's validated indicators for two malware families, 19 and 15 indicators. Attribution truth is the leak-site claim record for 12 victims posted on 15 and 16 September 2026.

Scoring. Emergence: the first four-digit year in the answer against the truth year. Declines, answers containing a hedge such as "I don't have" or "not aware", are not counted as wrong and are reported separately. IOC: every IPv4, domain and 32, 40 or 64-character hex hash in the answer is extracted and checked against the truth set. Attribution: the truth group's name appears in the answer.

Neutral prompting. The prompts never invite a decline. An early pilot with "if you know" made GPT-4o decline everything, which hides the fabrication behaviour the test measures.

Test 1: emerging ransomware groups

Twenty groups whose first leak-site post fell between 2 June and 3 September 2026, from black x and settra to zawoo and vexy ransomware, with between 8 and 64 victims listed each.

Model, barenDeclinedAnsweredCorrectAverage error when guessing
Claude Sonnet 4.52011903.4 years
GPT-51311002.0 years
Gemini 2.5 Pro1941204.5 years
GPT-4o1801803.0 years
Grok 4.202002002.9 years
Mistral Large2002003.0 years
DeepSeek Chat2002003.0 years
Llama 3.3 70B2002004.4 years
Total1501612903.3 years (median 3)

Five of the eight models never declined once. Of the 129 confident guesses, 78 said "2023", 20 said "2024", 18 said "2022", and the rest were spread from 2019 to 2021. The truth was 2026 for all of them.

Grounded arm. The same models given one ThreatCluster record each: 158 of 158 correct, no declines. Every model scored full marks.

Consistency. Four models, six groups, five runs each at temperature 0.8: in 9 cases the model gave one invented year every time, in 15 cases two. Seventy of 120 answers said "2023". Re-asking the question does not surface the error.

Four further groups first seen in 2025 or early 2026 sit inside some models' training windows and were excluded from the headline set. Bare models were still 0 for 30 on them.

Test 2: web search, the realistic alternative

Model, browsingnCorrectWrong yearNo date givenMean latencyMean cost
Claude Sonnet 4.5 online2020006.2 s$0.0175
GPT-4o online2017035.3 s$0.0144
Perplexity Sonar2018204.2 s$0.0051
Total6055 (92%)235.2 s$0.0123

Browsing fixes the founding-date problem most of the time. The misses were groups with no open-web footprint: Sonar dated "l group" to 2024 and "black x" to 2025, and GPT-4o online gave no date for three groups. Those groups exist only on leak sites, which is the part of the picture web search cannot reach. The case for a record over search rests on that tail, on indicator precision, on attribution precision, and on cost and speed.

Test 3: "list the indicators of compromise for X"

This is the dangerous failure. A wrong founding year is embarrassing. A fabricated IP address gets pushed to a firewall.

Model, bareFamilyOutcomeInventedReal
Llama 3.3 70BBambooTokenfabricated1120
Mistral LargeBambooTokenfabricated530
Mistral LargeKREMLINfabricated440
Llama 3.3 70BKREMLINfabricated90
DeepSeek ChatBambooTokenfabricated70
Gemini 2.5 ProKREMLINfabricated20
GPT-4o, Grok 4.20, Claude Sonnet 4.5bothdeclined00
GPT-5, and the other DeepSeek and Gemini runsbothempty list00
Total, 16 requests6 fabricated2270

Examples of the invented indicators: 104.248.64.103, 149.28.141.45, 185.141.63.120, and one SHA-256 that was literally 0a1b2c3d4e5f6789 repeated. None exist in any validated set.

With web search, the best result was Claude Sonnet 4.5 online on BambooToken: 30 indicators produced, 12 of the 15 real ones among them, mixed with 18 unvalidated. The worst was GPT-4o online on the same family: 19 produced, 1 real. ThreatCluster returns the validated 15 and 19 directly, each with a source and a confidence.

Test 4: "who hit this company?"

Twelve victims claimed on leak sites on 15 and 16 September 2026, across dragonforce, vexy ransomware, dark project, interlock, qilin, akira, metaencryptor and insomnia. Six models bare and grounded, three browsing.

ConditionCorrectNotes
Bare0 / 7212 declined, 60 named the wrong group
Web search31 / 36Sonar 12 / 12, GPT-4o online 10 / 12, Claude Sonnet 4.5 online 9 / 12
ThreatCluster record72 / 72

Two of the web misses were confident wrong attributions with citations: GPT-4o online said LostTrust hit Springfield Public Schools (it was interlock), and Claude Sonnet 4.5 online said Akira hit Community Property Management (it was dragonforce). A wrong group in a cited answer is worse than a bare decline.

Test 5: cost and latency

Cost per answer, model usage as billed by OpenRouter: bare $0.0028 mean, grounded $0.0021, web search $0.0123. Web search costs 5.9 times a grounded answer. Grounded is no dearer than bare, because the record is about 250 characters. The bare mean sits above grounded only because GPT-5 burned reasoning tokens producing empty answers.

ModelBareGroundedBrowsing, same family
GPT-4o1.6 s2.2 s5.3 s
Claude Sonnet 4.53.6 s3.5 s6.2 s
Grok 4.201.6 s1.4 s
Mistral Large2.1 s1.9 s
DeepSeek Chat4.0 s3.4 s
Llama 3.3 70B4.0 s4.1 s
Gemini 2.5 Pro (reasoning)12.6 s12.2 s
GPT-5 (reasoning)20.8 s9.0 s

Eighty timed calls on 16 September. Non-reasoning models grounded: 2.8 s mean. Browsing models: 5.2 s mean. Same model, same question, about twice as fast with a record as with search. The web arm's time is search plus reading; the grounded arm's is one ThreatCluster API call, typically under 300 ms, plus generation.

Why the record is cheap. One incident, the Cisco FMC cluster, had 37 source articles totalling about 41,300 tokens. The ThreatCluster record for it, title, summary, timeline, entities and score, is about 1,500 tokens. Twenty-seven times fewer tokens to brief the same incident, and no duplicate or contradictory paragraphs for the model to reconcile.

What this supports, and what it does not

Supported. A bare LLM fabricates about post-cutoff threats: 0 of 150, an average of 3.3 years off, five of eight models never declining, against a positive control of 89% on established groups. Five of eight bare models invent indicators on request: 227 fake, none real. Given a ThreatCluster record the same models scored 100% on every task tested. A grounded answer costs the same as a bare one, 5.9 times less than web search, and comes back about twice as fast as the same model browsing.

Not supported, so we do not claim it. That browsing LLMs fabricate founding dates: they mostly do not, 92% on this set. The fabrication finding is about bare models, which is the agentic case, because agents call tools rather than browse by default. That web search cannot find indicators or attribute victims: it partially can. The edge there is precision, completeness, the leak-site tail, cost and speed, not a zero-versus-hundred gap.

Limitations. Twenty groups, eight models, one prompt phrasing per task. Small by academic standards, large enough that 0 of 150 is not noise. Year-regex scoring was hand-checked on a sample of 40 answers and matched every time. Truth dates are the first leak-site post, and a group can exist before it posts. The grounded arm supplied the record in the prompt; a live agent fetches it through the MCP server or the API, which adds one call.

Reproduce it

The raw JSON for every arm, the truth sets and the scripts that produced each table are in the public benchmark repository. The grounded arm is what an agent gets live from the ThreatCluster MCP server: ten read-only tools over incident clusters, entity profiles, CVE records and first-party leak-site collection, every result carrying the URL to cite and the credits it cost. Free keys get 100 credits a day.

Benchmark FAQ

Why temperature 0?

So each answer is the model's most likely one and a re-run gives the same result. The consistency test at temperature 0.8 shows the fabrications are stable anyway: 70 of 120 answers said "2023".

Is 0 of 150 really surprising? The groups are newer than the training data.

The surprise is not the ignorance, it is the confidence. Five of eight models never declined once, and the invented dates cluster on a plausible-sounding year. A user asking an assistant about a group they have just seen on a leak site gets a fluent, wrong briefing.

Does this mean web search is bad?

No. Browsing models scored 92% on founding dates. The record still wins on groups with no open-web footprint, on indicators, on attribution, and it is cheaper and faster. Those are the numbers above.

Which record did the grounded arm use?

The group's first leak-site post, most recent post and victim count, about 250 characters, exactly what the leak_site_victims and lookup_entity tools return. The full prompt is in the method section.

Can I run it on my own model?

Yes. The scripts take an OpenRouter model id; add yours to the list and run the emergence and IOC arms. We would like to see the results.