Nvidia TensorRT-LLM is a technology platform tracked across 2 threat clusters and 2 intelligence report mentions on ThreatCluster. First observed November 14, 2025; most recent activity November 18, 2025.
Nvidia TensorRT-LLM is Nvidia's AI inference framework designed to run large language models efficiently on GPUs, enabling high-throughput, low-latency deployment. In cybersecurity terms, it represents a high-value attack surface within AI pipelines, with recent disclosures of copy-paste vulnerabilities and remote-code-execution flaws in AI inference frameworks that can enable malware delivery and compromise CI/CD/deployment pipelines.
Researchers have identified critical remote code execution (RCE) vulnerabilities in AI inference engines utilized by Meta, Nvidia, and Microsoft. These flaws could potentially allow attackers to exploit the frameworks,…
Critical remote code execution (RCE) vulnerabilities have been identified in AI inference frameworks from Meta, Nvidia, Microsoft, and open-source projects like vLLM and SGLang. These flaws arise from unsafe code reuse,…