<?xml version='1.0' encoding='UTF-8'?>
<?xml-stylesheet href="/static/style.xsl" type="text/xsl"?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" version="2.0">
  <channel>
    <title>Most recent entries from all</title>
    <link>https://cve.radiocsirt.org</link>
    <description>Contains only the most 10 recent entries.</description>
    <docs>http://www.rssboard.org/rss-specification</docs>
    <generator>python-feedgen</generator>
    <language>en</language>
    <lastBuildDate>Sat, 03 Oct 2026 20:35:36 +0000</lastBuildDate>
    <item>
      <title>CVE-2026-54234 — vLLM: Remote DoS in vLLM via Invalid Recovered Token Reinjection</title>
      <link>https://cve.radiocsirt.org/vuln/cve-2026-54234</link>
      <description>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; vllm-project vllm&lt;/p&gt;
&lt;p&gt;vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.0, a frontend-legal multi-request speculative decoding workload can cause the rejection sampler to produce a recovered token equal to the model vocabulary size boundary value, which is then converted to negative one when the engine selects the next live token for a request and is written back into the drafter&amp;#39;s input ids; that out-of-vocabulary value is later consumed by the model&amp;#39;s embedding and attention path and crashes the engine worker with a GPU device-side assertion. The same triggering request sequence is reachable through the public gRPC Generate and Abort endpoints, so a remote client that can send generation requests can crash the shared engine worker, aborting concurrent requests and causing a service-wide denial of service for other clients of the deployment until the worker is restarted. This issue is fixed in version 0.24.0.&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; vllm-project vllm&lt;/p&gt;
&lt;p&gt;vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.0, a frontend-legal multi-request speculative decoding workload can cause the rejection sampler to produce a recovered token equal to the model vocabulary size boundary value, which is then converted to negative one when the engine selects the next live token for a request and is written back into the drafter&amp;#39;s input ids; that out-of-vocabulary value is later consumed by the model&amp;#39;s embedding and attention path and crashes the engine worker with a GPU device-side assertion. The same triggering request sequence is reachable through the public gRPC Generate and Abort endpoints, so a remote client that can send generation requests can crash the shared engine worker, aborting concurrent requests and causing a service-wide denial of service for other clients of the deployment until the worker is restarted. This issue is fixed in version 0.24.0.&lt;/p&gt;</content:encoded>
      <guid isPermaLink="false">https://cve.radiocsirt.org/vuln/cve-2026-54234</guid>
    </item>
    <item>
      <title>GHSA-8wr5-jm2h-8r4f — vLLM has Remote DoS via Invalid Recovered Token Reinjection</title>
      <link>https://cve.radiocsirt.org/vuln/ghsa-8wr5-jm2h-8r4f</link>
      <description>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; PyPI: vllm&lt;/p&gt;
&lt;p&gt;## Summary&lt;/p&gt;
&lt;p&gt;A frontend-legal multi-request speculative workload can make vLLM produce an out-of-vocabulary recovered token equal to `vocab_size`, convert that value to `-1` when choosing the next live token for a request, and then feed that `-1` back into the next drafter input ids. On Qwen3 GPTQ this reaches the worker-side drafting / attention path and crashes the engine with a GPU `device-side assert`.&lt;/p&gt;
&lt;p&gt;The same issue is reachable through the public gRPC request surface by sending a specific overlapping `Generate` / `Abort` sequence.&lt;/p&gt;
&lt;p&gt;## Impact&lt;/p&gt;
&lt;p&gt;- A remote client that can send public gRPC generation requests can crash the
  shared vLLM engine worker
- The triggering request sequence aborts concurrent requests and prevents later
  requests from completing until the worker is restarted
- In shared deployments, this is a service-wide denial of service for other
  clients, not just a failure isolated to the attacking requests
- The failure is reproducible, so repeated request sequences can sustain the
  outage&lt;/p&gt;
&lt;p&gt;## Affected version&lt;/p&gt;
&lt;p&gt;- Confirmed on vLLM `0.17.1`
- Earlier and later versions have not been checked yet in this report&lt;/p&gt;
&lt;p&gt;## Repro model&lt;/p&gt;
&lt;p&gt;- Official Hugging Face repo:
  - [`Qwen/Qwen3-0.6B-GPTQ-Int8`](https://huggingface.co/Qwen/Qwen3-0.6B-GPTQ-Int8)
- Anyone wants to reproduce the bug with my PoC scripts should download `Qwen3-0.6B-GPTQ-Int8` first&lt;/p&gt;
&lt;p&gt;## Trigger chain&lt;/p&gt;
&lt;p&gt;1. A legal multi-request speculative workload keeps structured-output state,
   speculative decoding,…&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; PyPI: vllm&lt;/p&gt;
&lt;p&gt;## Summary&lt;/p&gt;
&lt;p&gt;A frontend-legal multi-request speculative workload can make vLLM produce an out-of-vocabulary recovered token equal to `vocab_size`, convert that value to `-1` when choosing the next live token for a request, and then feed that `-1` back into the next drafter input ids. On Qwen3 GPTQ this reaches the worker-side drafting / attention path and crashes the engine with a GPU `device-side assert`.&lt;/p&gt;
&lt;p&gt;The same issue is reachable through the public gRPC request surface by sending a specific overlapping `Generate` / `Abort` sequence.&lt;/p&gt;
&lt;p&gt;## Impact&lt;/p&gt;
&lt;p&gt;- A remote client that can send public gRPC generation requests can crash the
  shared vLLM engine worker
- The triggering request sequence aborts concurrent requests and prevents later
  requests from completing until the worker is restarted
- In shared deployments, this is a service-wide denial of service for other
  clients, not just a failure isolated to the attacking requests
- The failure is reproducible, so repeated request sequences can sustain the
  outage&lt;/p&gt;
&lt;p&gt;## Affected version&lt;/p&gt;
&lt;p&gt;- Confirmed on vLLM `0.17.1`
- Earlier and later versions have not been checked yet in this report&lt;/p&gt;
&lt;p&gt;## Repro model&lt;/p&gt;
&lt;p&gt;- Official Hugging Face repo:
  - [`Qwen/Qwen3-0.6B-GPTQ-Int8`](https://huggingface.co/Qwen/Qwen3-0.6B-GPTQ-Int8)
- Anyone wants to reproduce the bug with my PoC scripts should download `Qwen3-0.6B-GPTQ-Int8` first&lt;/p&gt;
&lt;p&gt;## Trigger chain&lt;/p&gt;
&lt;p&gt;1. A legal multi-request speculative workload keeps structured-output state,
   speculative decoding,…&lt;/p&gt;</content:encoded>
      <guid isPermaLink="false">https://cve.radiocsirt.org/vuln/ghsa-8wr5-jm2h-8r4f</guid>
    </item>
  </channel>
</rss>
