Back Redpacketsecurity CVE Alert: CVE-2026-92983 – InternLM
InternLM LMDeploy through 0.17.0 in DistServe prefill/decode disaggregation mode fails to release scheduler sessions because the proxy uses user-facing session IDs instead of internal scheduler keys. Unauthenticated attackers can send completion requests to the proxy endpoint that accumulate unreleased scheduler metadata and memory until the prefill worker is out-of-memory killed.
**Risk verdict:** High availability risk for exposed inference services, but urgency cannot be elevated to active-exploitation response because KEV, SSVC and EPSS status are not provided.
**Why this matters:** An unauthenticated remote party can consume worker memory without needing a valid user session, potentially taking an AI inference endpoint offline. The realistic outcome is denial of service, failed requests and disruption to dependent applications; confidentiality and integrity impacts appear limited.
**Most likely attack path:** The attack is network-reachable (AV:N), low-complexity (AC:L), requires no privileges (PR:N) or user interaction (UI:N), and has no stated attack requirements. Repeated requests to the affected proxy can exhaust the prefill worker; Scope is unchanged, so direct lateral movement is not implied, although shared infrastructure may broaden operational impact.
**Who is most exposed:** Internet-facing or partner-facing LMDeploy deployments using DistServe prefill/decode disaggregation are the primary concern, especially shared GPU services and multi-tenant inference gateways.
Alert on unauthenticated request bursts and unusual session-ID patterns.
Track scheduler-session counts and memory growth per worker.
Alert on rapid GPU/host memory exhaustion or prefill-worker restarts.
Correlate proxy traffic spikes with timeout and out-of-memory events.
Mitigation and prioritisation:
Confirm and deploy the vendor’s corrective release; until confirmed, treat upgrades as urgent.
Restrict proxy access through authentication, network allow-lists and rate limits.
Apply per-client quotas, request concurrency caps and worker memory safeguards.
Isolate prefill workers and rehearse rollback before production change.
Obtain KEV, SSVC and EPSS data before final queue ranking; current evidence does not justify “treat as priority 1”.
A considerable amount of time and effort goes into maintaining this website, creating backend automation and creating new features and content for you to make actionable intelligence decisions. Everyone that supports the site helps enable new functionality.
If you like the site, please support us on Patreon or Buy Me A Coffee using the buttons below.
The full story
This article is one source in a clustered incident — the cluster page carries the summary, timeline and every other outlet covering it.
