Skip to content
CVE Alert: CVE-2026-94627 – vllm-project

CVE Alert: CVE-2026-94627 – vllm-project

Redpacketsecurity •admin • September 22, 2026

vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when concurrent child requests a single transfer ID in prefill/decode disaggregated deployments. Attackers can trigger GPU memory exhaustion by submitting completion requests with multiple prompts, causing orphaned KV cache blocks to accumulate until process restart and eventually preventing legitimate requests from executing.

## AI Summary Analysis

**Risk verdict:** High availability risk for exposed inference services, but urgency cannot be elevated to active exploitation without KEV, SSVC, PoC or EPSS data.

**Why this matters:** A remote, unauthenticated attacker may be able to consume GPU memory and deny service to legitimate model users without stealing data or modifying results. The practical impact is halted inference, failed customer workloads, expensive GPU capacity loss and possible emergency restarts; repeated requests could create an effective resource-exhaustion attack.

**Most likely attack path:** Because AV:N, AC:L, PR:N and UI:N apply, an attacker only needs network access to the request endpoint and can automate malicious completion requests without user involvement. Scope is unchanged, so direct impact is confined to the affected service, although shared GPU nodes, queues or orchestration platforms may experience knock-on disruption.

**Who is most exposed:** Internet-facing API gateways, multi-tenant hosted inference platforms and deployments using Mooncake with prefill/decode disaggregation are the principal targets. Internal services remain exposed where untrusted users, tenants or compromised workloads can submit requests.

Alert on abnormal prompt concurrency or request bursts from one source or tenant.

Track GPU memory growth, KV-cache occupancy and orphaned allocation rates.

Correlate rising allocation failures, latency spikes and worker restarts.

Review gateway, scheduler and inference logs for repeated multi-prompt requests.

Mitigation and prioritisation:

Upgrade promptly to a vendor-supported release containing the fix; validate Mooncake workflows first.

Restrict inference endpoints with authentication, quotas, rate limits and per-tenant concurrency caps.

Isolate GPU workers and enforce memory watchdogs with rapid, controlled process recycling.

Monitor for service degradation during rollout and retain rollback capacity; no exploitation or EPSS status is provided.

A considerable amount of time and effort goes into maintaining this website, creating backend automation and creating new features and content for you to make actionable intelligence decisions. Everyone that supports the site helps enable new functionality.

If you like the site, please support us on Patreon or Buy Me A Coffee using the buttons below.

Extracted Entities

Attack Types (1)

Platforms (1)