<?xml version='1.0' encoding='UTF-8'?>
<?xml-stylesheet href="/static/style.xsl" type="text/xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <id>https://cve.radiocsirt.org/rss/recent/all/10</id>
  <title>Most recent entries from all</title>
  <updated>2026-10-02T08:55:04.195828+00:00</updated>
  <author>
    <name>Vulnerability-Lookup</name>
    <email>csirt@opendfir.org</email>
  </author>
  <link href="https://cve.radiocsirt.org" rel="alternate"/>
  <generator uri="https://lkiesow.github.io/python-feedgen" version="1.0.0">python-feedgen</generator>
  <subtitle>Contains only the most 10 recent entries.</subtitle>
  <entry>
    <id>https://cve.radiocsirt.org/vuln/cve-2026-69147</id>
    <title>CVE-2026-69147 — vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation</title>
    <updated>2026-10-02T08:55:04.197561+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> vllm-project vllm</p>
<p>vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set media_io_kwargs.video.video_backend to pynvvideocodec, and MediaConnector.fetch_video forwards that choice to VideoMediaIO even when startup configuration selected a software decoder. The engine's _reserve_mm_ipc_gpu_memory logic budgets decoder memory only from static configuration, so the request-selected VIDEO_LOADER_REGISTRY backend can create a CUDA context, decoder surfaces, and decoded-frame allocations that were not removed from the engine's KV-cache budget. An attacker able to submit video requests to a video-capable GPU deployment with PyNvVideoCodec installed can exhaust shared GPU memory, causing request failures, worker crashes, or denial of service. The first release containing the fix is version 0.28.0.</p></div>
    </content>
    <link href="https://cve.radiocsirt.org/vuln/cve-2026-69147"/>
  </entry>
  <entry>
    <id>https://cve.radiocsirt.org/vuln/ghsa-8pw2-6jv3-mj5j</id>
    <title>GHSA-8pw2-6jv3-mj5j — vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation</title>
    <updated>2026-10-02T08:55:04.197618+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> PyPI: vllm</p>
<p>## Summary</p>
<p>Current vLLM `main` lets an inference request choose the PyNvVideoCodec GPU video decoder through `media_io_kwargs.video.video_backend`, but engine GPU memory reservation is computed only from static startup configuration and `VLLM_VIDEO_LOADER_BACKEND`. If the server starts with the default OpenCV/software backend and no `--mm-ipc-gpu-memory-gb` budget, a client can still route a video request into the PyNvVideoCodec path after startup, causing frontend CUDA-context, decoder-surface, and decoded-frame GPU allocations that were not carved out of the engine KV-cache budget.</p>
<p>## Technical Details</p>
<p>The vulnerable boundary is the split between request-time media decoding choices in the API server and startup-time memory budgeting in the engine worker. Request bodies for Chat Completions and Responses expose `media_io_kwargs`, and those values are forwarded to the shared media connector. For video inputs, `MediaConnector.fetch_video()` copies `self.media_io_kwargs["video"]` into `video_io_kwargs`, only setting a model-derived backend when `video_backend` is absent. `VideoMediaIO.__init__()` then consumes `video_backend` from those kwargs and loads that backend from `VIDEO_LOADER_REGISTRY`.</p>
<p>The relevant request-side source path is:</p>
<p>```python
video_io_kwargs = dict(self.media_io_kwargs.get("video", {}))
if "video_backend" not in video_io_kwargs and (
    video_backend := get_video_loader_backend_for_processor(video_processor)
):
    video_io_kwargs["video_backend"] =…</p></div>
    </content>
    <link href="https://cve.radiocsirt.org/vuln/ghsa-8pw2-6jv3-mj5j"/>
  </entry>
</feed>
