Vllm-Project Vllm vulnerabilities
50 known vulnerabilities affecting vllm-project/vllm.
Total CVEs
50
CISA KEV
0
Public exploits
1
Exploited in wild
0
Severity breakdown
CRITICAL7HIGH18MEDIUM23LOW2
Vulnerabilities
Page 3 of 3
CVE-2025-48942P4MEDIUMCVSS 6.5v>= 0.8.0, < 0.9.02025-05-30
CVE-2025-48942 [MEDIUM] CWE-248 CVE-2025-48942: vLLM is an inference and serving engine for large language models (LLMs). In versions 0.8.0 up to bu
vLLM is an inference and serving engine for large language models (LLMs). In versions 0.8.0 up to but excluding 0.9.0, hitting the /v1/completions API with a invalid json_schema as a Guided Param kills the vllm server. This vulnerability is similar GHSA-9hcf-v7m4-6m2j/CVE-2025-48943, but for regex instead of a JSON schema. Version 0.9.0 fixes the is
nvd
CVE-2025-48887P4MEDIUMCVSS 6.5v>= 0.6.4, < 0.9.02025-05-30
CVE-2025-48887 [MEDIUM] CWE-1333 CVE-2025-48887: vLLM, an inference and serving engine for large language models (LLMs), has a Regular Expression Den
vLLM, an inference and serving engine for large language models (LLMs), has a Regular Expression Denial of Service (ReDoS) vulnerability in the file `vllm/entrypoints/openai/tool_parsers/pythonic_tool_parser.py` of versions 0.6.4 up to but excluding 0.9.0. The root cause is the use of a highly complex and nested regular expression for tool call det
nvd
CVE-2026-9540P4MEDIUMCVSS 5.3v0.19.02026-05-26
CVE-2026-9540 [MEDIUM] CWE-404 CVE-2026-9540: A vulnerability was identified in vllm-project vllm 0.19.0. This issue affects some unknown processi
A vulnerability was identified in vllm-project vllm 0.19.0. This issue affects some unknown processing of the component OpenAI-compatible Serving Path. Such manipulation leads to denial of service. It is possible to launch the attack remotely. The exploit is publicly available and might be used. The pull request to fix this issue awaits acceptance.
nvd
CVE-2026-78684P4MEDIUMCVSS 5.3fixed in 0.27.02026-08-25
CVE-2026-78684 [MEDIUM] CWE-400 CVE-2026-78684: vLLM before 0.27.0 fails to properly classify DeepStream as a GPU backend and omits pixel-limit enfo
vLLM before 0.27.0 fails to properly classify DeepStream as a GPU backend and omits pixel-limit enforcement in its decode path. Unauthenticated attackers can activate DeepStream at request time to initialize the process-wide GPU decode pool and submit video that bypasses resource controls, causing partial denial of service for concurrent requests.
nvd
CVE-2026-73555P4MEDIUMCVSS 5.3fixed in 0.26.02026-08-13
CVE-2026-73555 [MEDIUM] CWE-209 CVE-2026-73555: vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the validation_e
vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the validation_exception_handler in vllm/entrypoints/openai/server_utils.py converts FastAPI RequestValidationError objects with str(exc), and sanitize_message in vllm/entrypoints/utils.py does not remove traceback-style file paths, allowing unauthenticated malformed
nvd
CVE-2026-73558P4MEDIUMCVSS 5.3fixed in 0.27.02026-08-13
CVE-2026-73558 [MEDIUM] CWE-190 CVE-2026-73558: vLLM is an inference and serving engine for large language models. Prior to 0.27.0, an integer overf
vLLM is an inference and serving engine for large language models. Prior to 0.27.0, an integer overflow in blockIdx.x * 2 * d in activation_kernels.cu can cause act_and_mul_kernel to consume another batched user's input, allowing a request processed in the same inference batch to receive a partial or complete copy of another user's inference result.
nvd
CVE-2026-12491P4MEDIUMCVSS 4.8≥ 0.11.0, < 0.24.02026-06-17
CVE-2026-12491 [MEDIUM] CWE-115 CVE-2026-12491: A flaw was found in vLLM, an open-source library for large language model inference. This vulnerabil
A flaw was found in vLLM, an open-source library for large language model inference. This vulnerability arises from improper handling of image metadata, specifically EXIF orientation and PNG transparency (tRNS) data, during image processing. When images are converted to RGB, transparency information may be implicitly discarded or remapped, leading t
nvd
CVE-2026-71486P4MEDIUMCVSS 4.3fixed in 0.26.02026-08-17
CVE-2026-71486 [MEDIUM] CWE-400 CVE-2026-71486: vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the /v1/completi
vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the /v1/completions/derender and /v1/chat/completions/derender endpoints accept caller-supplied GenerateResponse objects whose generate_responses, choices, token_ids, prompt_logprobs, logprobs.content, top_logprobs, and routed_experts structures are processed by Onli
nvd
CVE-2025-25183P4LOWCVSS 2.6fixed in 0.7.22025-02-07
CVE-2025-25183 [LOW] CWE-354 CVE-2025-25183: vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Maliciously co
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Maliciously constructed statements can lead to hash collisions, resulting in cache reuse, which can interfere with subsequent responses and cause unintended behavior. Prefix caching makes use of Python's built-in hash() function. As of Python 3.12, the behavior of has
nvd
CVE-2025-46570P4LOWCVSS 2.6fixed in 0.9.02025-05-29
CVE-2025-46570 [LOW] CWE-208 CVE-2025-46570: vLLM is an inference and serving engine for large language models (LLMs). Prior to version 0.9.0, wh
vLLM is an inference and serving engine for large language models (LLMs). Prior to version 0.9.0, when a new prompt is processed, if the PageAttention mechanism finds a matching prefix chunk, the prefill process speeds up, which is reflected in the TTFT (Time to First Token). These timing differences caused by matching chunks are significant enough to
nvd
← Previous3 / 3