Evidence used
- No CISA KEV confirmation is currently recorded.
- EPSS is 0.46% for the current model date.
BlackTreeCVE Intelligencevllm-project · vllm
Official source article: GitHub GHSA-MGRM-FGJV-MHV8 ↗. Check the applicable product and release in the original source.
Medium technical severity with no CISA KEV confirmation; remediate through the normal risk-based patch cycle unless local exposure raises the priority.
These OSV and GitHub advisory ranges apply only to the named package and ecosystem. A listed fixed version is not a universal product patch or proof that an update is installed.
| Ecosystem and package | Affected range | First fixed version | Evidence |
|---|---|---|---|
| pipvllm | < 0.8.0 | 0.8.0 | GitHub advisory ↗upstream repository advisory · 8 Jun 2026 |
Medium technical severity with no CISA KEV confirmation; remediate through the normal risk-based patch cycle unless local exposure raises the priority.
Patch availablevLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. The outlines library is one of the backends used by vLLM to support structured output (a.k.a. guided decoding). Outlines provides an optional cache for its compiled grammars on the local filesystem. This cache has been on by default in vLLM. Outlines is also available by default through the OpenAI compatible API server. The affected code in vLLM is vllm/model_executor/guided_decoding/outlines_logits_processors.py, which unconditionally uses the cache from outlines. A malicious user can send a stream of very short decoding requests with unique schemas, resulting in an addition to the cache for each request. This can result in a Denial of Service if the filesystem runs out of space. Note that even if vLLM was configured to use a different backend by default, it is still possible to choose outlines on a per-request basis using the guided_decoding_backend key of the extra_body field of the request. This issue applies only to the V0 engine and is fixed in 0.8.0.
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. The outlines library is one of the backends used by vLLM to support structured output (a.k.a. guided decoding). Outlines provides an optional cache for its compiled grammars on the local filesystem. This cache has been on by default in vLLM. Outlines is also available by default through the OpenAI compatible API server. The affected code in vLLM is vllm/model_executor/guided_decoding/outlines_logits_processors.py, which unconditionally uses the cache from outlines. A malicious user can send a stream of very short decoding requests with unique schemas, resulting in an addition to the cache for each request. This can result in a Denial of Service if the filesystem runs out of space. Note that even if vLLM was configured to use a different backend by default, it is still possible to choose outlines on a per-request basis using the guided_decoding_backend key of the extra_body field of the request. This issue applies only to the V0 engine and is fixed in 0.8.0.
The product allocates a reusable resource or group of resources on behalf of an actor without imposing any intended restrictions on the size or number of resources that can be allocated.
An attacker operating through a network path may attempt exploitation with low privileges. If successful, the issue may disrupt the affected service.
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. The outlines library is one of the backends used by vLLM to support structured output (a.k.a. guided decoding). Outlines provides an optional cache for its compiled grammars on the local filesystem. This cache has been on by default in vLLM. Outlines is also available by default through the OpenAI compatible API server. The affected code in vLLM is vllm/model_executor/guided_decoding/outlines_logits_processors.py, which unconditionally uses the cache from outlines. A malicious user can send a stream of very short decoding requests with unique schemas, resulting in an addition to the cache for each request. This can result in a Denial of Service if the filesystem runs out of space. Note that even if vLLM was configured to use a different backend by default, it is still possible to choose outlines on a per-request basis using the guided_decoding_backend key of the extra_body field of the request. This issue applies only to the V0 engine and is fixed in 0.8.0.
The product allocates a reusable resource or group of resources on behalf of an actor without imposing any intended restrictions on the size or number of resources that can be allocated.
An attacker operating through a network path may attempt exploitation with low privileges. If successful, the issue may disrupt the affected service.
CVSS severity, EPSS forecast probability, public exploit material and CISA-confirmed exploitation are separate signals.
No CISA KEV match was present at the last successful refresh. This means no confirmation from that source, not proof of no exploitation.
No exploit-tagged reference or CISA SSVC proof-of-concept state is currently recorded. Research may still exist outside the structured feeds.
CWE-770: Allocation of Resources Without Limits or Throttling. The product allocates a reusable resource or group of resources on behalf of an actor without imposing any intended restrictions on the size or number of resources that can be allocated.
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:HCommon Vulnerability Scoring System 3.1: the compact vector below is decoded into plain language.
Operational remediation based on structured source evidence.
Published 19 Mar 2025 · Last source change 19 Mar 2025, 20:15 UTC · CWE-770 · Allocation of Resources Without Limits or Throttling
Core structured fields are present and their contributing authorities are shown above.
No material field changes have been recorded since change tracking began. Routine source refreshes and cosmetic edits are intentionally excluded.