<?xml version='1.0' encoding='UTF-8'?>
<?xml-stylesheet href="/static/style.xsl" type="text/xsl"?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" version="2.0">
  <channel>
    <title>Most recent entries from all</title>
    <link>https://cve.radiocsirt.org</link>
    <description>Contains only the most 10 recent entries.</description>
    <docs>http://www.rssboard.org/rss-specification</docs>
    <generator>python-feedgen</generator>
    <language>en</language>
    <lastBuildDate>Fri, 09 Oct 2026 10:49:46 +0000</lastBuildDate>
    <item>
      <title>CVE-2025-46570 — vLLM’s Chunk-Based Prefix Caching Vulnerable to Potential Timing Side-Channel</title>
      <link>https://cve.radiocsirt.org/vuln/cve-2025-46570</link>
      <description>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; vllm-project vllm&lt;/p&gt;
&lt;p&gt;vLLM is an inference and serving engine for large language models (LLMs). Prior to version 0.9.0, when a new prompt is processed, if the PageAttention mechanism finds a matching prefix chunk, the prefill process speeds up, which is reflected in the TTFT (Time to First Token). These timing differences caused by matching chunks are significant enough to be recognized and exploited. This issue has been patched in version 0.9.0.&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; vllm-project vllm&lt;/p&gt;
&lt;p&gt;vLLM is an inference and serving engine for large language models (LLMs). Prior to version 0.9.0, when a new prompt is processed, if the PageAttention mechanism finds a matching prefix chunk, the prefill process speeds up, which is reflected in the TTFT (Time to First Token). These timing differences caused by matching chunks are significant enough to be recognized and exploited. This issue has been patched in version 0.9.0.&lt;/p&gt;</content:encoded>
      <guid isPermaLink="false">https://cve.radiocsirt.org/vuln/cve-2025-46570</guid>
    </item>
    <item>
      <title>GHSA-4qjh-9fv9-r85r — Potential Timing Side-Channel Vulnerability in vLLM’s Chunk-Based Prefix Caching</title>
      <link>https://cve.radiocsirt.org/vuln/ghsa-4qjh-9fv9-r85r</link>
      <description>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; PyPI: vllm&lt;/p&gt;
&lt;p&gt;This issue arises from the prefix caching mechanism, which may expose the system to a timing side-channel attack.&lt;/p&gt;
&lt;p&gt;## Description
When a new prompt is processed, if the PageAttention mechanism finds a matching prefix chunk, the prefill process speeds up, which is reflected in the TTFT (Time to First Token). Our tests revealed that the timing differences caused by matching chunks are significant enough to be recognized and exploited.&lt;/p&gt;
&lt;p&gt;For instance, if the victim has submitted a sensitive prompt or if a valuable system prompt has been cached, an attacker sharing the same backend could attempt to guess the victim&amp;#39;s input. By measuring the TTFT based on prefix matches, the attacker could verify if their guess is correct, leading to potential leakage of private information.&lt;/p&gt;
&lt;p&gt;Unlike token-by-token sharing mechanisms, vLLM’s chunk-based approach (PageAttention) processes tokens in larger units (chunks). In our tests, with chunk_size=2, the timing differences became noticeable enough to allow attackers to infer whether portions of their input match the victim&amp;#39;s prompt at the chunk level.&lt;/p&gt;
&lt;p&gt;## Environment&lt;/p&gt;
&lt;p&gt;- GPU: NVIDIA A100 (40G)
- CUDA: 11.8
- PyTorch: 2.3.1
- OS: Ubuntu 18.04
- vLLM: v0.5.1
Configuration: We launched vLLM using the default settings and adjusted chunk_size=2 to evaluate the TTFT.&lt;/p&gt;
&lt;p&gt;## Leakage
We conducted our tests using LLaMA2-70B-GPTQ on a single device. We analyzed the timing differences when prompts shared prefixes of 2 chunks, and plotted the corresponding ROC c…&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; PyPI: vllm&lt;/p&gt;
&lt;p&gt;This issue arises from the prefix caching mechanism, which may expose the system to a timing side-channel attack.&lt;/p&gt;
&lt;p&gt;## Description
When a new prompt is processed, if the PageAttention mechanism finds a matching prefix chunk, the prefill process speeds up, which is reflected in the TTFT (Time to First Token). Our tests revealed that the timing differences caused by matching chunks are significant enough to be recognized and exploited.&lt;/p&gt;
&lt;p&gt;For instance, if the victim has submitted a sensitive prompt or if a valuable system prompt has been cached, an attacker sharing the same backend could attempt to guess the victim&amp;#39;s input. By measuring the TTFT based on prefix matches, the attacker could verify if their guess is correct, leading to potential leakage of private information.&lt;/p&gt;
&lt;p&gt;Unlike token-by-token sharing mechanisms, vLLM’s chunk-based approach (PageAttention) processes tokens in larger units (chunks). In our tests, with chunk_size=2, the timing differences became noticeable enough to allow attackers to infer whether portions of their input match the victim&amp;#39;s prompt at the chunk level.&lt;/p&gt;
&lt;p&gt;## Environment&lt;/p&gt;
&lt;p&gt;- GPU: NVIDIA A100 (40G)
- CUDA: 11.8
- PyTorch: 2.3.1
- OS: Ubuntu 18.04
- vLLM: v0.5.1
Configuration: We launched vLLM using the default settings and adjusted chunk_size=2 to evaluate the TTFT.&lt;/p&gt;
&lt;p&gt;## Leakage
We conducted our tests using LLaMA2-70B-GPTQ on a single device. We analyzed the timing differences when prompts shared prefixes of 2 chunks, and plotted the corresponding ROC c…&lt;/p&gt;</content:encoded>
      <guid isPermaLink="false">https://cve.radiocsirt.org/vuln/ghsa-4qjh-9fv9-r85r</guid>
    </item>
  </channel>
</rss>
