vLLM before 0.28.0 Denial of Service via negative token ID
vllm-project · vllm
Official source article: GitHub GHSA-25Q3-V2HM-8VPF ↗. Check the applicable product and release in the original source.
8.7HighCVSS 4.0
Recommended action
Patch only the product branches with a verified fix
High technical severity with public exploit material referenced by a structured source; prioritise exposed affected systems while verifying vendor guidance. Verified remediation exists for at least one product or source, but 27 structured product or package states remain unresolved. Apply remediation only to the exact product branch confirmed by its source.
Published severityHigh→Operational priority:Critical, raised one band.upgradedsince 18 Sep 2026
Evidence used
No CISA KEV confirmation is currently recorded.
A structured source references public exploit or proof-of-concept material.
The selected CVSS metric records a network-reachable, unauthenticated path with no user interaction.
EPSS is 0.54% for the current model date.
Compensating controls
Validate the affected product branch and deploy the verified fixed release.
Restrict the affected network interface to trusted sources where business-safe.
Increase monitoring for the attack path and post-exploitation behaviour described in the report.
Verification
Confirm that the asset runs vllm-project vllm and falls inside the recorded affected range.
Verify the installed build against the product-specific fixed version after deployment.
Validate exposure, authentication requirements and compensating controls in the actual environment.
Reopen this reassessment when CVSS, KEV, EPSS, exploit evidence or remediation changes.
Cross-source reconciliation
Remediation availability differs by product scope
Verified remediation exists for at least one product or source, but 27 structured product or package states remain unresolved. Apply remediation only to the exact product branch confirmed by its source.
Red Hat Product Security: 27 affected or under-investigation product states without a fixed product state in the same current advisory
Open-source package ranges1 source-attributed range
These OSV and GitHub advisory ranges apply only to the named package and ecosystem. A listed fixed version is not a universal product patch or proof that an update is installed.
Ecosystem and package
Affected range
First fixed version
Evidence
PyPIvllm
ECOSYSTEM: introduced 0; fixed 0.28.0
0.28.0
OSV record ↗source linked ecosystem record · 29 Sep 2026
Direct vendor intelligence
Authoritative vendor CSAF and VEX advisories
Structured product status and remediation from the issuing vendor. Product-state explanations are always visible; large lists can be searched or downloaded.
1 current
CVE-2026-93592 · CSAF 2.0 · revision 3 · finalRed Hat Product Securityvllm: vLLM: Denial of Service via negative token ID input
27 known affected
The vendor explicitly identifies these products as affected by this CVE.
rhaii/vllm-cpu-rhel9 as a component of Red Hat AI Inference Server
rhaii/vllm-cuda-rhel9 as a component of Red Hat AI Inference Server
rhaii/vllm-gaudi-rhel9 as a component of Red Hat AI Inference Server
rhaii/vllm-neuron-rhel9 as a component of Red Hat AI Inference Server
rhaii/vllm-rocm-rhel9 as a component of Red Hat AI Inference Server
rhaii/vllm-spyre-rhel9 as a component of Red Hat AI Inference Server
rhaii/vllm-tpu-rhel9 as a component of Red Hat AI Inference Server
rhaiis/vllm-cpu-rhel9 as a component of Red Hat AI Inference Server
rhaiis/vllm-cuda-rhel9 as a component of Red Hat AI Inference Server
rhaiis/vllm-neuron-rhel9 as a component of Red Hat AI Inference Server
rhaiis/vllm-rocm-rhel9 as a component of Red Hat AI Inference Server
rhaiis/vllm-spyre-rhel9 as a component of Red Hat AI Inference Server
Summary
A flaw was found in vLLM. An unauthenticated attacker can exploit an input validation vulnerability by submitting negative token IDs to the /v1/embeddings and /pooling endpoints. This action triggers a CUDA device-side assertion, which poisons the GPU context and causes the engine to crash. Consequently, all subsequent requests will fail until the process is restarted, leading to a Denial of Service (DoS).
Remediation
Restrict access to the inference service or filter requests at an upstream reverse proxy or API gateway: 1. Limit network access: Deploy firewall rules or OpenShift NetworkPolicies so that the inference server ports and endpoints are accessible only by authenticated, trusted internal clients. 2. Ingress request validation: Configure an API gateway or reverse proxy placed in front of the model server to inspect request payloads for the `/v1/embeddings` and `/pooling` endpoints, rejecting requests where the `input` array contains negative numbers, or restricting untrusted clients to submitting text prompt strings rather than raw token lists. Caveats: Inspecting request bodies at an intermediary gateway layer may introduce slight processing latency. Modifying reverse proxy rules or applying network policies may require a service configuration reload or restart, which can temporarily disrupt in-flight connections.
Optional official sources
National CERT insights ?CERT means Computer Emergency Response Team; CSIRT is the closely related term Computer Security Incident Response Team.
Choose official national sources for this report. Each advisory shows its original language. Your selection is remembered on this device and included in shared links.
Official European source
ENISA European Vulnerability Database
Official EUVD identifiers, advisory evidence and known-exploited context. Missing fields are not treated as evidence of low risk.
1 current
ENISA EUVD identifier
EUVD-2026-82911
No EUVD known-exploited evidence
vLLM versions before 0.28.0 fail to validate the lower bound of token IDs in the /v1/embeddings and /pooling endpoints, allowing unauthenticated attackers to crash the engine by submitting negative token IDs. A single request with a negative token ID triggers a CUDA device-side assertion that poisons the GPU context, causing all subsequent requests to fail until the process restarts.
EUVD state
Present in the current official mapping
Known exploitation
Not present in the current ENISA EUVD known-exploited dataset. This is not proof of no exploitation.
ENISA score
8.7 · CVSS 4.0
Advisory evidence
No linked advisory details stored yet
Recommended actionPatch only the product branches with a verified fix
High technical severity with public exploit material referenced by a structured source; prioritise exposed affected systems while verifying vendor guidance. Verified remediation exists for at least one product or source, but 27 structured product or package states remain unresolved. Apply remediation only to the exact product branch confirmed by its source.
Fix availability varies by product
01
What, why and how
vLLM versions before 0.28.0 fail to validate the lower bound of token IDs in the /v1/embeddings and /pooling endpoints, allowing unauthenticated attackers to crash the engine by submitting negative token IDs. A single request with a negative token ID triggers a CUDA device-side assertion that poisons the GPU context, causing all subsequent requests to fail until the process restarts.
What
vLLM versions before 0.28.0 fail to validate the lower bound of token IDs in the /v1/embeddings and /pooling endpoints, allowing unauthenticated attackers to crash the engine by submitting negative token IDs. A single request with a negative token ID triggers a CUDA device-side assertion that poisons the GPU context, causing all subsequent requests to fail until the process restarts.
Why
The product uses untrusted input when calculating or using an array index, but the product does not validate or incorrectly validates the index to ensure the index references a valid position within the array.
How
An attacker operating through a network path may attempt exploitation without authentication or user interaction. If successful, the issue may disrupt the affected service.
What
vLLM versions before 0.28.0 fail to validate the lower bound of token IDs in the /v1/embeddings and /pooling endpoints, allowing unauthenticated attackers to crash the engine by submitting negative token IDs. A single request with a negative token ID triggers a CUDA device-side assertion that poisons the GPU context, causing all subsequent requests to fail until the process restarts.
Why
The product uses untrusted input when calculating or using an array index, but the product does not validate or incorrectly validates the index to ensure the index references a valid position within the array.
How
An attacker operating through a network path may attempt exploitation without authentication or user interaction. If successful, the issue may disrupt the affected service.
02
Exploit reality and attack path
CVSS severity, EPSS forecast probability, public exploit material and CISA-confirmed exploitation are separate signals.
Observed exploitation ?Confirmed exploitation and public exploit material are separate signals. Attacks can occur without public proof-of-concept or exploit code.No confirmed evidence
No CISA KEV match was present at the last successful refresh. This means no confirmation from that source, not proof of no exploitation.
Public PoC / exploit material ?Confirmed exploitation and public exploit material are separate signals. Attacks can occur without public proof-of-concept or exploit code.Reference recorded
A structured CVE source labels at least one public reference as exploit material. BlackTree has not independently validated that it is safe, reliable or weaponised.
Likely attack path
a network path → Improper Validation of Array Index → disrupt the affected service
Attack surface
Network
Privileges required
None: unauthenticated exploitation is possible
User interaction
None
Attack complexity
Low: no specialised conditions are recorded
Security boundary
Not a CVSS 4.0 base metric
Weakness ?CWE means Common Weakness Enumeration: a standard category for the underlying weakness.
CWE-129: Improper Validation of Array Index. The product uses untrusted input when calculating or using an array index, but the product does not validate or incorrectly validates the index to ensure the index references a valid position within the array.
CVSS vector ?CVSS means Common Vulnerability Scoring System. The vector records the metric values used to calculate technical severity.
Common Vulnerability Scoring System 4.0: the compact vector below is decoded into plain language.
AVNetworkAttack vector: The vulnerable component can be reached over a network.ACLowAttack complexity: No specialised conditions are required beyond attacker-controlled input.ATNoneAttack requirements: No additional deployment or execution condition is required.PRNonePrivileges required: The attacker does not need an account or existing privileges.UINoneUser interaction: No action by another user is required.VCNoneVulnerable-system confidentiality: No direct loss is represented by this metric.VINoneVulnerable-system integrity: No direct loss is represented by this metric.VAHighVulnerable-system availability: A successful attack can cause a major loss.SCNoneSubsequent-system confidentiality: No direct loss is represented by this metric.SINoneSubsequent-system integrity: No direct loss is represented by this metric.SANoneSubsequent-system availability: No direct loss is represented by this metric.
Post-exploitation / living off the land
Stolen credentials, tokens and legitimate administration functions may provide continued access without deploying a large custom toolset.
NetworkUnauthenticatedDenial of serviceCWE-129Public exploit reference
A
Official authority intelligence
Only matched European and national findings are included. Language selectors and unavailable sources are omitted.
ENISA EUVD · EUVD-2026-82911Official EUVD mapping
vLLM versions before 0.28.0 fail to validate the lower bound of token IDs in the /v1/embeddings and /pooling endpoints, allowing unauthenticated attackers to crash the engine by submitting negative token IDs. A single request with a negative token ID triggers a CUDA device-side assertion that poisons the GPU context, causing all subsequent requests to fail until the process restarts.
Official EUVD record ↗BSI · German · WID-SEC-2026-3187vllm: Schwachstelle ermöglicht Denial of Service
Ein entfernter, anonymer Angreifer kann eine Schwachstelle in vllm ausnutzen, um einen Denial of Service Angriff durchzuführen.
Official advisory ↗JVN iPedia · Japanese · JVNDB-2026-035803vLLMにおける配列インデックスの検証に関する脆弱性
Operational remediation based on structured source evidence.
Status ?Patch availability is based on structured fixed-version fields and authoritative update references. If no fix is verified, check the vendor advisory before making a change.
Use the product-specific evidence above. Patch only products with a verified fixed release, and keep every affected or under-investigation state without a matching fix in the remediation queue.
Workaround
Restrict access to the inference service or filter requests at an upstream reverse proxy or API gateway: 1. Limit network access: Deploy firewall rules or OpenShift NetworkPolicies so that the inference server ports and endpoints are accessible only by authenticated, trusted internal clients. 2. Ingress request validation: Configure an API gateway or reverse proxy placed in front of the model server to inspect request payloads for the `/v1/embeddings` and `/pooling` endpoints, rejecting requests where the `input` array contains negative numbers, or restricting untrusted clients to submitting text prompt strings rather than raw token lists. Caveats: Inspecting request bodies at an intermediary gateway layer may introduce slight processing latency. Modifying reverse proxy rules or applying network policies may require a service configuration reload or restart, which can temporarily disrupt in-flight connections.
Official vendor articles1 stored document
Linked articles are downloaded and versioned as source evidence. A CVE mention or approved update relationship does not, by itself, verify a fix for every product branch.
Technical terms and abbreviations used in this report
CVE
Common Vulnerabilities and Exposures: the public identifier for one disclosed vulnerability.
CVSS
Common Vulnerability Scoring System: a technical severity framework; it is not patching priority by itself.
EPSS
Exploit Prediction Scoring System: FIRST's estimate of the probability that exploitation activity will be observed in the next 30 days; it is a forecast, not confirmation.
CWE
Common Weakness Enumeration: the standard category describing the underlying software or hardware weakness.
CNA
CVE Numbering Authority: an organisation authorised to assign and publish CVE records.
CISA ADP
Cybersecurity and Infrastructure Security Agency Authorized Data Publisher: structured enrichment added to a CVE record.
NVD
National Vulnerability Database: NIST's enrichment service for CVE records.
CERT / CSIRT
A computer security incident response team that publishes warnings or coordinates incident response.
PoC
Proof of concept: public material that demonstrates or helps reproduce exploitation.
CSAF
Common Security Advisory Framework: a machine-readable format for security advisories.
LoTL
Living off the land: abuse of legitimate tools or system functions during an attack.