About / How the analysis is made
How the analysis on a ThreatCluster page is made
ThreatCluster is an aggregator. An incident page groups the public reporting on one event and adds a written analysis on top. Most of that analysis is generated by a language model, without a person reading it first. This page says what is generated, what it is generated from, the checks it passes, what still goes wrong, and how to get a record corrected.
What a page is built from
Every incident page starts with articles. Our pipeline reads several hundred security sources every fifteen minutes, keeps the security stories, and groups the articles that cover the same event into one cluster. The articles are listed on the page, with their publisher and a link, and they are the only reporting the analysis is written from.
Three other things feed in, all of them records rather than writing:
- CVE records from MITRE, the NVD, CISA's Known Exploited Vulnerabilities catalogue and public exploit repositories: publication dates, CVSS ratings, KEV status, proof-of-concept dates.
- Leak-site listings we collect directly from ransomware groups' own sites. A listing is a record of the group's claim, not a finding that a breach happened, and the page says so.
- Entities extracted from the articles: threat actors, malware, CVEs, products, companies, countries. These link pages together and drive the related-incident lists.
Nothing on the page comes from the open web at large, and nothing comes from a model's own memory of events. If a fact is not in the listed articles or one of those records, it should not be on the page. When it is, that is an error, and the section on corrections below is how it gets fixed.
What is generated, and from what
Each cluster is passed to a language model with the text of its articles and the CVE records. The model returns the parts of the page listed here. None of them are read by a person before publication.
| On the page | Generated from | Notes |
|---|---|---|
| Headline | The articles | Frozen after first publication so the address does not change. Retitled only by hand. |
| Summary and key points | The articles, plus CVE records for dates and ratings | Up to ten sentences. Regenerated as new articles join the cluster, unless a summary has been pinned after a correction. |
| Timeline | Dates the articles state, plus CVE and leak-site records | Each entry names the article or record it came from. A date that only marks the edge of a range is not an event. |
| Threat score and urgency | The model's rating against a fixed rubric, combined with how many outlets covered the event and how recent it is | A score is a ranking aid for the feed, not an assessment of your exposure. |
| Exploitation status and source confidence | The articles and the KEV record | KEV listing overrides the model's reading of the articles. |
| Questions and answers | The articles | Three questions a defender would ask, answered from the same text. |
Two things on a page are not generated: the list of source articles, and any editor's note. Ask AI, the question box on each page, is a separate feature and answers with citations to these same sources; it is described at ThreatCluster AI.
The checks a page passes
Generated text is checked by code before it is published, and the checks are tightened whenever a correction shows a gap.
- Loaded words need a source. A headline or summary may only call a flaw critical, a zero-day, actively exploited, wormable or unpatched if one of the listed articles uses the word. "Critical" also needs a CVSS rating of 9.0 or higher in the CVE record. A flaw whose fix predates its disclosure is not a zero-day, whatever one article's opening line says. Words that fail the check are removed.
- Dates come from records where records exist. A CVE's publication, KEV and proof-of-concept dates are taken from the record, not from the model's reading of an article.
- Hedges stay hedged. The model is instructed to keep a source's "may have" as "may have", not to infer how long a breach lasted from the dates of the records affected, and not to invent timeline events to fill a date.
- Record feeds do not merge into stories. Templated posts such as leak-site listings, CVE records and exploit listings only join a cluster that names the same CVE, advisory or organisation, so a page about one incident does not fill with unrelated records.
- Pinned summaries are left alone. After a correction the summary is pinned and the pipeline does not regenerate it.
What still goes wrong
The checks catch classes of error, not every error. These are the kinds we have published corrections for, and they can recur:
- A hedge in a source flattened into a statement, or a duration inferred from a date range.
- A severity word a source did not use, or a label one source used that the same source contradicts.
- An early report overtaken by a later one that the cluster has not yet picked up.
- An organisation named in a leak-site listing whose position is not yet on the record.
Every page shows when its analysis was last updated and lists every source, so a claim can be checked against the reporting it came from. Treat the analysis as a starting point and the sources as the record.
Corrections and disputes
If a summary or record about you or your organisation is wrong, write to [email protected]. We check the request against the published sources, change what they do not support and leave what they do. The outcome is published as an editor's note on the page concerned and at /corrections: what was asked, what changed and why. The note names the organisation, never the individual who wrote.
Leak-site listings are not removed on request, because the listing is a fact about the group's behaviour whether or not its claim is true. We add the organisation's position to the record instead.
Complaints about a source article go to its publisher. We correct what we generated; we do not edit what they reported.
The full policy is section 9 of the terms of service. For anything else, the contact page is the place.
More about the platform
- Real-time clustering: how articles become one incident.
- Entity intelligence: the profiles and relationship graph.
- ThreatCluster AI: cited answers on top of the record.
- Corrections: every published editor's note.