Skip to content
CVE Alert: CVE-2026-93688 – sgl-project

CVE Alert: CVE-2026-93688 – sgl-project

Redpacketsecurity admin September 19, 2026

SGLang through 0.5.19 in prefill/decode disaggregation mode with Mooncake KV transfer backend fails to validate bootstrap_room values, allowing unbounded transfer state allocation. Unauthenticated attackers can reach the decode engine’s POST /generate endpoint and submit arbitrary bootstrap_room values to exhaust prefill process memory until out-of-memory termination.

This is a high-priority availability risk for externally reachable SGLang inference services, but there is insufficient evidence here to classify it as actively exploited or emergency priority.

A remote, unauthenticated attacker could deliberately consume process memory and terminate workloads, disrupting model serving and any applications dependent on the inference endpoint. Likely objectives include denial of service, capacity exhaustion, service instability, and forcing failover to less capable or less trusted infrastructure. Confidentiality and integrity impacts appear limited, although repeated disruption could create material operational and financial consequences.

### Most likely attack path

The network-reachable endpoint can be targeted directly with low-complexity requests requiring no privileges or user interaction. No special environmental conditions are indicated, making automated resource-exhaustion attempts feasible against exposed decode engines. Scope is unchanged, so direct lateral movement is not expected; however, disruption may affect connected prefill services, queues, APIs, and customer-facing applications.

### Who is most exposed

Organisations running SGLang with disaggregated prefill/decode architecture and the specified transfer backend are most at risk, particularly where inference APIs are exposed beyond a trusted service network. Internet-facing research, hosted AI, and multi-tenant GPU platforms warrant immediate review.

Alert on unusual or rapidly increasing `bootstrap_room` values and request rates.

Correlate endpoint access with rising resident memory, allocation failures, OOM kills, and container restarts.

Monitor prefill/decode health checks, queue depth, latency, and worker churn.

Review unauthenticated access logs for repeated malformed or high-volume generation requests.

### Mitigation and prioritisation

Upgrade promptly to the first vendor-supported release containing the validation fix; confirm behaviour in staging.

Restrict the generation endpoint to authenticated, trusted networks and place strict rate and request-size limits at the gateway.

Enforce container memory limits and restart isolation to prevent node-wide impact.

Coordinate changes with GPU platform owners and test failover before production rollout.

KEV, SSVC exploitation status, EPSS, and PoC indicators are not supplied; obtain them before downgrading urgency.

A considerable amount of time and effort goes into maintaining this website, creating backend automation and creating new features and content for you to make actionable intelligence decisions. Everyone that supports the site helps enable new functionality.

If you like the site, please support us on Patreon or Buy Me A Coffee using the buttons below.

Extracted Entities

Attack Types (1)