CWE-704
Allowed-with-ReviewIncorrect Type Conversion or Cast
Abstraction: Class · Status: Incomplete
The product does not correctly convert an object, resource, or structure from one type to a different type.
358 vulnerabilities reference this CWE, most recent first.
GHSA-PF36-R9C6-H97J
Vulnerability from github – Published: 2022-11-21 22:18 – Updated: 2022-11-21 22:18Impact
When printing a tensor, we get it's data as a const char* array (since that's the underlying storage) and then we typecast it to the element type. However, conversions from char to bool are undefined if the char is not 0 or 1, so sanitizers/fuzzers will crash.
Patches
We have patched the issue in GitHub commit 1be743703279782a357adbf9b77dcb994fe8b508.
The fix will be included in TensorFlow 2.11.0. We will also cherrypick this commit on TensorFlow 2.10.1, TensorFlow 2.9.3, and TensorFlow 2.8.4, as these are also affected and still in supported range.
For more information
Please consult our security guide for more information regarding the security model and how to contact us with issues and questions.
Attribution
This vulnerability was discovered via internal fuzzing.
{
"affected": [
{
"package": {
"ecosystem": "PyPI",
"name": "tensorflow"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"fixed": "2.8.4"
}
],
"type": "ECOSYSTEM"
}
]
},
{
"package": {
"ecosystem": "PyPI",
"name": "tensorflow-cpu"
},
"ranges": [
{
"events": [
{
"introduced": "2.9.0"
},
{
"fixed": "2.9.3"
}
],
"type": "ECOSYSTEM"
}
]
},
{
"package": {
"ecosystem": "PyPI",
"name": "tensorflow-gpu"
},
"ranges": [
{
"events": [
{
"introduced": "2.10.0"
},
{
"fixed": "2.10.1"
}
],
"type": "ECOSYSTEM"
}
]
},
{
"package": {
"ecosystem": "PyPI",
"name": "tensorflow"
},
"ranges": [
{
"events": [
{
"introduced": "2.9.0"
},
{
"fixed": "2.9.3"
}
],
"type": "ECOSYSTEM"
}
]
},
{
"package": {
"ecosystem": "PyPI",
"name": "tensorflow"
},
"ranges": [
{
"events": [
{
"introduced": "2.10.0"
},
{
"fixed": "2.10.1"
}
],
"type": "ECOSYSTEM"
}
]
},
{
"package": {
"ecosystem": "PyPI",
"name": "tensorflow-cpu"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"fixed": "2.8.4"
}
],
"type": "ECOSYSTEM"
}
]
},
{
"package": {
"ecosystem": "PyPI",
"name": "tensorflow-gpu"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"fixed": "2.8.4"
}
],
"type": "ECOSYSTEM"
}
]
},
{
"package": {
"ecosystem": "PyPI",
"name": "tensorflow-gpu"
},
"ranges": [
{
"events": [
{
"introduced": "2.9.0"
},
{
"fixed": "2.9.3"
}
],
"type": "ECOSYSTEM"
}
]
},
{
"package": {
"ecosystem": "PyPI",
"name": "tensorflow-cpu"
},
"ranges": [
{
"events": [
{
"introduced": "2.10.0"
},
{
"fixed": "2.10.1"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2022-41911"
],
"database_specific": {
"cwe_ids": [
"CWE-704"
],
"github_reviewed": true,
"github_reviewed_at": "2022-11-21T22:18:11Z",
"nvd_published_at": "2022-11-18T22:15:00Z",
"severity": "MODERATE"
},
"details": "### Impact\nWhen [printing a tensor](https://github.com/tensorflow/tensorflow/blob/807cae8a807960fd7ac2313cde73a11fc15e7942/tensorflow/core/framework/tensor.cc#L1200-L1227), we get it\u0027s data as a `const char*` array (since that\u0027s the underlying storage) and then we typecast it to the element type. However, conversions from `char` to `bool` are undefined if the `char` is not `0` or `1`, so sanitizers/fuzzers will crash.\n\n### Patches\nWe have patched the issue in GitHub commit [1be743703279782a357adbf9b77dcb994fe8b508](https://github.com/tensorflow/tensorflow/commit/1be743703279782a357adbf9b77dcb994fe8b508).\n\nThe fix will be included in TensorFlow 2.11.0. We will also cherrypick this commit on TensorFlow 2.10.1, TensorFlow 2.9.3, and TensorFlow 2.8.4, as these are also affected and still in supported range.\n\n### For more information\nPlease consult [our security guide](https://github.com/tensorflow/tensorflow/blob/master/SECURITY.md) for more information regarding the security model and how to contact us with issues and questions.\n\n### Attribution\nThis vulnerability was discovered via internal fuzzing.\n",
"id": "GHSA-pf36-r9c6-h97j",
"modified": "2022-11-21T22:18:11Z",
"published": "2022-11-21T22:18:11Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/tensorflow/tensorflow/security/advisories/GHSA-pf36-r9c6-h97j"
},
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2022-41911"
},
{
"type": "WEB",
"url": "https://github.com/tensorflow/tensorflow/commit/1be743703279782a357adbf9b77dcb994fe8b508"
},
{
"type": "PACKAGE",
"url": "https://github.com/tensorflow/tensorflow"
},
{
"type": "WEB",
"url": "https://github.com/tensorflow/tensorflow/blob/807cae8a807960fd7ac2313cde73a11fc15e7942/tensorflow/core/framework/tensor.cc#L1200-L1227"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:H/PR:L/UI:R/S:U/C:N/I:N/A:H",
"type": "CVSS_V3"
}
],
"summary": "Invalid char to bool conversion when printing a tensor"
}
GHSA-PFCW-X62W-QQW9
Vulnerability from github – Published: 2022-05-13 01:31 – Updated: 2022-05-13 01:31This vulnerability allows remote attackers to execute arbitrary code on vulnerable installations of Foxit Reader 9.0.0.29935. User interaction is required to exploit this vulnerability in that the target must visit a malicious page or open a malicious file. The specific flaw exists within the parsing of subform elements. The issue results from the lack of proper validation of user-supplied data, which can result in a type confusion condition. An attacker can leverage this vulnerability to execute code under the context of the current process. Was ZDI-CAN-5371.
{
"affected": [],
"aliases": [
"CVE-2018-9937"
],
"database_specific": {
"cwe_ids": [
"CWE-704"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2018-05-17T15:29:00Z",
"severity": "HIGH"
},
"details": "This vulnerability allows remote attackers to execute arbitrary code on vulnerable installations of Foxit Reader 9.0.0.29935. User interaction is required to exploit this vulnerability in that the target must visit a malicious page or open a malicious file. The specific flaw exists within the parsing of subform elements. The issue results from the lack of proper validation of user-supplied data, which can result in a type confusion condition. An attacker can leverage this vulnerability to execute code under the context of the current process. Was ZDI-CAN-5371.",
"id": "GHSA-pfcw-x62w-qqw9",
"modified": "2022-05-13T01:31:41Z",
"published": "2022-05-13T01:31:41Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2018-9937"
},
{
"type": "WEB",
"url": "https://www.foxitsoftware.com/support/security-bulletins.php"
},
{
"type": "WEB",
"url": "https://zerodayinitiative.com/advisories/ZDI-18-321"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.0/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H",
"type": "CVSS_V3"
}
]
}
GHSA-PFGR-7F76-8GWH
Vulnerability from github – Published: 2022-05-13 01:34 – Updated: 2022-05-13 01:34This vulnerability allows remote attackers to execute arbitrary code on vulnerable installations of Foxit Reader 9.0.1.1049. User interaction is required to exploit this vulnerability in that the target must visit a malicious page or open a malicious file. The specific flaw exists within the removeDataObject method. By performing actions in JavaScript, an attacker can trigger a type confusion condition. An attacker can leverage this vulnerability to execute code under the context of the current process. Was ZDI-CAN-6033.
{
"affected": [],
"aliases": [
"CVE-2018-14270"
],
"database_specific": {
"cwe_ids": [
"CWE-704"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2018-07-31T20:29:00Z",
"severity": "HIGH"
},
"details": "This vulnerability allows remote attackers to execute arbitrary code on vulnerable installations of Foxit Reader 9.0.1.1049. User interaction is required to exploit this vulnerability in that the target must visit a malicious page or open a malicious file. The specific flaw exists within the removeDataObject method. By performing actions in JavaScript, an attacker can trigger a type confusion condition. An attacker can leverage this vulnerability to execute code under the context of the current process. Was ZDI-CAN-6033.",
"id": "GHSA-pfgr-7f76-8gwh",
"modified": "2022-05-13T01:34:38Z",
"published": "2022-05-13T01:34:38Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2018-14270"
},
{
"type": "WEB",
"url": "https://www.foxitsoftware.com/support/security-bulletins.php"
},
{
"type": "WEB",
"url": "https://zerodayinitiative.com/advisories/ZDI-18-730"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.0/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H",
"type": "CVSS_V3"
}
]
}
GHSA-PFX2-HM33-3J69
Vulnerability from github – Published: 2022-05-13 01:34 – Updated: 2022-05-13 01:34This vulnerability allows remote attackers to execute arbitrary code on vulnerable installations of Foxit Reader 9.0.1.1049. User interaction is required to exploit this vulnerability in that the target must visit a malicious page or open a malicious file. The specific flaw exists within the print method. By performing actions in JavaScript, an attacker can trigger a type confusion condition. An attacker can leverage this vulnerability to execute code under the context of the current process. Was ZDI-CAN-6032.
{
"affected": [],
"aliases": [
"CVE-2018-14269"
],
"database_specific": {
"cwe_ids": [
"CWE-704"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2018-07-31T20:29:00Z",
"severity": "HIGH"
},
"details": "This vulnerability allows remote attackers to execute arbitrary code on vulnerable installations of Foxit Reader 9.0.1.1049. User interaction is required to exploit this vulnerability in that the target must visit a malicious page or open a malicious file. The specific flaw exists within the print method. By performing actions in JavaScript, an attacker can trigger a type confusion condition. An attacker can leverage this vulnerability to execute code under the context of the current process. Was ZDI-CAN-6032.",
"id": "GHSA-pfx2-hm33-3j69",
"modified": "2022-05-13T01:34:38Z",
"published": "2022-05-13T01:34:38Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2018-14269"
},
{
"type": "WEB",
"url": "https://www.foxitsoftware.com/support/security-bulletins.php"
},
{
"type": "WEB",
"url": "https://zerodayinitiative.com/advisories/ZDI-18-729"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.0/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H",
"type": "CVSS_V3"
}
]
}
GHSA-PGM3-H7VV-WJCF
Vulnerability from github – Published: 2022-05-14 01:35 – Updated: 2022-05-14 01:35A Type Confusion (CWE-843) vulnerability exists in Eurotherm by Schneider Electric GUIcon V2.0 (Gold Build 683.0) on c3core.dll which could cause remote code to be executed when parsing a GD1 file
{
"affected": [],
"aliases": [
"CVE-2018-7815"
],
"database_specific": {
"cwe_ids": [
"CWE-704"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2019-02-06T23:29:00Z",
"severity": "HIGH"
},
"details": "A Type Confusion (CWE-843) vulnerability exists in Eurotherm by Schneider Electric GUIcon V2.0 (Gold Build 683.0) on c3core.dll which could cause remote code to be executed when parsing a GD1 file",
"id": "GHSA-pgm3-h7vv-wjcf",
"modified": "2022-05-14T01:35:22Z",
"published": "2022-05-14T01:35:22Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2018-7815"
},
{
"type": "WEB",
"url": "https://www.schneider-electric.com/ww/en/download/document/SEVD-2018-338-01"
},
{
"type": "WEB",
"url": "http://www.securityfocus.com/bid/106218"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.0/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H",
"type": "CVSS_V3"
}
]
}
GHSA-PH72-CQR5-QPP7
Vulnerability from github – Published: 2026-10-05 23:42 – Updated: 2026-10-05 23:42Affected
- Ecosystem / package: pip /
vllm - Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit
752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the scale-out transport path reaches.
Summary
vLLM's disaggregated scale-out transport splits a multimodal request into a trusted render step (POST /v1/chat/completions/render) and a separate generate step (POST /inference/v1/generate). The generate route decodes a caller-supplied features object — serialized encoder tensors (kwargs_data), multimodal hashes (mm_hashes), placeholder ranges (mm_placeholders), and the internal field-processor selection — and forwards it into the engine as if it had come from the trusted renderer, with no rebinding to (or validation against) the active model's renderer contract. Because the two routes are ordinary auth-guarded HTTP endpoints (the /inference prefix is registered by default on generate-capable servers), any authenticated caller can submit an otherwise-valid render body with a single forged field.
Depending on which field is forged, this produces:
- an engine-fatal crash of the shared EngineCore process (denial of service), reproduced as a CUDA illegal-memory-access, a post-admission rank-mismatch
ValueError, and a hardassert— three independent forged fields (sites 1, 2, 3); - silent cross-request encoder-cache poisoning / disclosure of shared encoder state when the cache-key hash is not bound to the payload (site 4);
- transport-level integrity loss when sparse placeholder masks are dropped during render-to-generate replay (site 5).
All five share one root cause and one fix shape: the reconstructed multimodal state on the scale-out path is trusted without being rebound to, and validated against, the active model's renderer output before it reaches the engine.
These sites are distinct from prior multimodal hardening. Site 1 survives GHSA-wv77-2vpf-vmmg (that fix validates full tensor shape in MultiModalDataParser/get_input_embeddings on the prompt-embeds path), because our request forges image_grid_thw metadata with the pixel bytes intact and reaches the Qwen2 vision RoPE/cu_seqlens and image_embeds.split sink, which the shape-check fix does not rebind. Site 4 is distinct from GHSA-c65p-x677-fgj6 (which folds metadata into MultiModalHasher.serialize_item to stop hash collisions), because the scale-out generate path trusts a caller-supplied mm_hash as the cache key with no origin binding, so that fix does not stop a caller from submitting a victim's hash or a kwargs_data=None cache read.
Affected code
Links pinned to the confirmed commit 752a3a504485 (v0.25.1).
Shared entry point and control surface for all five sites:
POST /inference/v1/generateroute:vllm/entrypoints/scale_out/token_in_token_out/api_router.py#L46-L75.- Public
featuresschema (kwargs_data,mm_hashes,mm_placeholders):vllm/entrypoints/scale_out/token_in_token_out/protocol.py#L42-L63. ServingTokens.serve_tokens()copies the decoded geometry and hashes into engine structures without rebinding to the renderer schema:vllm/entrypoints/scale_out/token_in_token_out/serving.py#L145-L172.- The scale-out routers are registered by default on generate-capable servers and
/inferenceis treated as an ordinary auth-guarded prefix:vllm/entrypoints/openai/api_server.py#L217-L219.
Site 1 — forged Qwen grid geometry (engine-fatal DoS). The decoded image_grid_thw is never rebound to the rendered pixel-tensor element count.
- Reconstruction with no model-specific geometry invariant:
serving.py#L147-L172. Qwen2VisionTransformer.prepare_encoder_metadata()derives RoPE tables,cu_seqlens, and the FlashAttentionmax_seqlenfrom the caller-supplied grid:qwen2_vl.py#L658-L712, consumed inforward()(#L713-L751)._process_image_input()computesimage_embeds.split(sizes)from the same untrusted metadata:qwen2_vl.py#L1348-L1369; M-RoPE positions atqwen2_vl.py#L1223.
The only shape check is assert grid_thw.ndim == 2; the split sizes and the vision-encoder call are then derived directly from the caller-supplied grid, with no cross-check against the pixel-tensor row count:
# vllm/model_executor/models/qwen2_vl.py Lines 1348-1369
def _process_image_input(
self, image_input: Qwen2VLImageInputs
) -> tuple[torch.Tensor, ...]:
grid_thw = image_input["image_grid_thw"]
assert grid_thw.ndim == 2
if image_input["type"] == "image_embeds":
image_embeds = image_input["image_embeds"]
else:
pixel_values = image_input["pixel_values"]
if self.use_data_parallel:
return run_dp_sharded_mrope_vision_model(
self.visual, pixel_values, grid_thw.tolist(), rope_type="rope_3d"
)
else:
image_embeds = self.visual(pixel_values, grid_thw=grid_thw)
# Split concatenated embeddings for each image item.
merge_size = self.visual.spatial_merge_size
sizes = (grid_thw.prod(-1) // merge_size // merge_size).tolist()
return image_embeds.split(sizes)
Site 2 — wire-selected field-processor type confusion (engine-fatal DoS). MsgpackDecoder._decode_mm_field_elem() trusts a wire-selected field-factory name and constructs the internal field processor directly from caller data.
vllm/v1/serial_utils.py#L440-L454readsfactory_meth_name, factory_kw = obj["field"]and callsgetattr(MultiModalFieldConfig, factory_meth_name).- Qwen2-VL field/schema contract:
qwen2_vl.py#L763-L791andQwen2VLImagePixelInputs(#L119-L144); parse/validate atqwen2_vl.py#L1300. - Rank mismatch raised post-admission:
vllm/utils/tensor_schema.py#L155-L171, reached fromTensorSchema.__init__→validate()(#L63).
Site 3 — non-positive placeholder length (engine-fatal DoS via reachable assert). PlaceholderRangeInfo{offset,length} is accepted as unconstrained integers and copied verbatim into the engine's PlaceholderRange.
- Schema:
protocol.py#L28-L35. - Copied into
PlaceholderRange:serving.py#L147-L153. - The only guard is an upper bound on embed count (
num_embeds > mm_encoder_cache_size), with no non-positive check:vllm/v1/engine/input_processor.py#L459-L464; raw length returned byvllm/multimodal/inputs.py#L152-L154; window selection assumes non-empty ranges atvllm/multimodal/utils.py#L114-L134. - Sink:
assert start_idx < end_idxatvllm/v1/worker/gpu/mm/encoder_runner.py#L114(duplicated atvllm/v1/worker/gpu_model_runner.py#L3192); turned into a fatal shutdown by EngineCore's uncaught-exception path atvllm/v1/engine/core.py#L1229-L1233.
A length of 0 makes num_encoder_tokens == 0, so end_idx collapses to 0 and the bare assert fires inside the engine worker:
# vllm/v1/worker/gpu/mm/encoder_runner.py Lines 108-114
pos_info = mm_feature.mm_position
start_pos = pos_info.offset
num_encoder_tokens = pos_info.length
start_idx = max(cur_query_start - start_pos, 0)
end_idx = min(cur_query_end - start_pos, num_encoder_tokens)
assert start_idx < end_idx
Site 4 — cache hash not bound to payload (integrity / disclosure). features.mm_hashes (the cache key) and kwargs_data (the tensor) are independent fields with no origin or integrity binding.
- Schema exposing the caller-controlled hash and the
None= resolve-from-cache semantics:protocol.py#L42-L63. - Handler forwards the caller's hashes unchanged into
mm_input(...):serving.py#L164-L170. - Input processing copies the caller hash into the feature identifier with only a string-type check:
vllm/v1/engine/input_processor.py#L165-L181. - Sink — the receiver cache returns the cached tensor solely by that key (
cache_key = feature.mm_hash or feature.identifier):vllm/multimodal/cache.py#L602-L607.
The cache key is the caller-supplied hash with no verification against the tensor bytes, so a forged mm_hash both stores under and reads back another request's slot:
# vllm/multimodal/cache.py Lines 601-607
for feature in mm_features:
cache_key = feature.mm_hash or feature.identifier
self.touch_receiver_cache_item(cache_key, feature.data)
for feature in mm_features:
cache_key = feature.mm_hash or feature.identifier
feature.data = self.get_and_update_item(feature.data, cache_key)
return mm_features
Site 5 — dropped sparse placeholder mask (transport integrity loss). The render path serializes placeholders as only offset/length, so models relying on sparse is_embed masks lose the mask during render-to-generate replay.
ServingRender._extract_mm_features()builds eachPlaceholderRangeInfo(offset=p.offset, length=p.length), discardingis_embed:vllm/entrypoints/scale_out/render/serving.py#L212-L229.- Transport schema carries no field for the mask:
protocol.py#L28. ServingTokens.serve_tokens()reconstructs a densePlaceholderRangeregardless of the original:serving.py#L148-L152.
Impact
A single authenticated request to a scale-out multimodal deployment can:
- Crash the shared EngineCore process (sites 1, 2, 3), taking the served model down for every tenant (
/health→ 503). Availability-only; no code execution or data disclosure demonstrated for these sites. - Silently poison or read back another request's shared encoder-cache state (site 4) — an integrity/disclosure primitive. Attack complexity is High because the attacker must know or induce the victim's content hash; there is no availability impact for this site.
- Corrupt backend-visible placeholder semantics across the render-to-generate boundary (site 5) for models that depend on sparse
is_embedmasks.
The forged multimodal payload is small; only the trust in its self-declared geometry/identity is the defect.
Suggested Fix
On the scale-out path, do not trust caller-supplied multimodal state as renderer-produced. After decoding features, rebind and validate the reconstructed MultiModalKwargsItem against the active model's renderer contract at the HTTP boundary:
- Reject any request whose decoded grid geometry is inconsistent with the rendered pixel-tensor element count and declared placeholder span, before it reaches
prepare_encoder_metadata()(site 1). - Rebind each field's processor type to the schema the active model's renderer declares (or reject if it differs), instead of reconstructing internal field processors from wire-selected factory names (site 2).
- Reject any
PlaceholderRangeInfowithlength <= 0(or out-of-rangeoffset) with a request-scoped 4xx, and convert the encoder-runner invariant into a checked, request-scoped error rather than a process-fatalassert(site 3). - Recompute or verify the content hash for submitted
kwargs_databefore using it as a cache key, and namespace receiver-cache keys to a server-generated or principal scope, refusing cache-reads for hashes the caller did not legitimately produce (site 4). - Serialize
is_embedinPlaceholderRangeInfo, validate its length against the placeholder span, and reconstruct it on replay (site 5).
Site 1 — validate grid geometry before the vision encoder. Replace the bare assert grid_thw.ndim == 2 in _process_image_input()/_process_video_input() with a shared helper that recomputes the split sizes and rejects a grid whose patch-row count does not match the pixel tensor (and rejects non-positive / non-merge-divisible dims), so the mismatch never reaches image_embeds.split():
# vllm/model_executor/models/qwen2_vl.py — _process_image_input()
- grid_thw = image_input["image_grid_thw"]
- assert grid_thw.ndim == 2
+ grid_thw = image_input["image_grid_thw"]
+ input_type = image_input["type"]
+ input_tensor = (
+ image_input["image_embeds"]
+ if input_type == "image_embeds"
+ else image_input["pixel_values"]
+ )
+ sizes = _validate_qwen2_vl_input_geometry(
+ modality="image",
+ input_type=input_type,
+ input_tensor=input_tensor,
+ grid_thw=grid_thw,
+ spatial_merge_size=self.visual.spatial_merge_size,
+ )
...
- # Split concatenated embeddings for each image item.
- merge_size = self.visual.spatial_merge_size
- sizes = (grid_thw.prod(-1) // merge_size // merge_size).tolist()
return image_embeds.split(sizes)
where the helper raises before the encoder runs:
# vllm/model_executor/models/qwen2_vl.py — new _validate_qwen2_vl_input_geometry()
if t <= 0 or h <= 0 or w <= 0:
raise ValueError(f"{modality} grid_thw row {index} must be positive ...")
if h % spatial_merge_size != 0 or w % spatial_merge_size != 0:
raise ValueError(f"{modality} grid_thw row {index} must be divisible ...")
...
if actual_rows != expected_rows:
raise ValueError(
f"{modality} {row_kind} do not match grid_thw: "
f"expected {expected_rows}, got {actual_rows}."
)
Site 3 — constrain the placeholder schema. Make PlaceholderRangeInfo reject non-positive lengths and negative offsets at the Pydantic boundary (plus parallel-length, non-overlapping, and within-prompt validators), turning the process-fatal assert into a request-scoped 422:
# vllm/entrypoints/scale_out/token_in_token_out/protocol.py — PlaceholderRangeInfo
- offset: int
- length: int
+ offset: int = Field(ge=0)
+ length: int = Field(gt=0)
Site 4 — bind the cache key to the payload. Derive each cache key by hashing the submitted serialized tensor (ignoring the caller's mm_hashes) and refuse cache-only reads, so a forged hash can neither poison nor read a victim's slot:
# vllm/entrypoints/scale_out/token_in_token_out/serving.py — serve_tokens()
+ mm_hashes = _bind_mm_hashes_to_kwargs_data(
+ features.mm_hashes, features.kwargs_data,
+ )
engine_input = mm_input(
prompt_token_ids=request.token_ids,
mm_kwargs=MultiModalKwargsItems(mm_kwargs),
- mm_hashes=features.mm_hashes,
+ mm_hashes=mm_hashes,
mm_placeholders=mm_placeholders,
cache_salt=request.cache_salt,
)
where _bind_mm_hashes_to_kwargs_data() raises on kwargs_data is None (cache-only read) and derives sha256(modality || "\0" || serialized_item) per item. Site 2 applies the same rebind-to-declared-schema pattern in mm_serde.py (passing modality + mm_processor into decode_mm_kwargs_item), and site 5 adds an is_embed field to PlaceholderRangeInfo with from_placeholder_range/to_placeholder_range helpers so the sparse mask survives render-to-generate replay. Each site fix ships with a regression test. This packet groups the five sites because they share one entry point (/inference/v1/generate + /render) and one root cause; we are happy to split it into per-component advisories (for example, engine-fatal input-validation vs. cache-key binding vs. transport-schema integrity) if the vLLM team prefers.
Credit
Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)
These vulnerabilities were discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51898
{
"affected": [
{
"package": {
"ecosystem": "PyPI",
"name": "vllm"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"fixed": "0.30.0"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2026-105754"
],
"database_specific": {
"cwe_ids": [
"CWE-1284",
"CWE-20",
"CWE-617",
"CWE-639",
"CWE-668",
"CWE-704"
],
"github_reviewed": true,
"github_reviewed_at": "2026-10-05T23:42:50Z",
"nvd_published_at": null,
"severity": "MODERATE"
},
"details": "## Affected\n\n- **Ecosystem / package:** pip / `vllm`\n- **Affected versions:** vLLM \u2264 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the scale-out transport path reaches.\n\n## Summary\n\nvLLM\u0027s disaggregated **scale-out** transport splits a multimodal request into a trusted render step (`POST /v1/chat/completions/render`) and a separate generate step (`POST /inference/v1/generate`). The generate route decodes a caller-supplied `features` object \u2014 serialized encoder tensors (`kwargs_data`), multimodal hashes (`mm_hashes`), placeholder ranges (`mm_placeholders`), and the internal field-processor selection \u2014 and forwards it into the engine **as if it had come from the trusted renderer**, with no rebinding to (or validation against) the active model\u0027s renderer contract. Because the two routes are ordinary auth-guarded HTTP endpoints (the `/inference` prefix is registered by default on generate-capable servers), any authenticated caller can submit an otherwise-valid render body with a single forged field.\n\nDepending on which field is forged, this produces:\n\n- an **engine-fatal crash** of the shared EngineCore process (denial of service), reproduced as a CUDA illegal-memory-access, a post-admission rank-mismatch `ValueError`, and a hard `assert` \u2014 three independent forged fields (sites 1, 2, 3);\n- **silent cross-request encoder-cache poisoning / disclosure** of shared encoder state when the cache-key hash is not bound to the payload (site 4);\n- **transport-level integrity loss** when sparse placeholder masks are dropped during render-to-generate replay (site 5).\n\nAll five share one root cause and one fix shape: the reconstructed multimodal state on the scale-out path is trusted without being rebound to, and validated against, the active model\u0027s renderer output before it reaches the engine.\n\nThese sites are distinct from prior multimodal hardening. Site 1 survives [GHSA-wv77-2vpf-vmmg](https://github.com/vllm-project/vllm/security/advisories/GHSA-wv77-2vpf-vmmg) (that fix validates full tensor shape in `MultiModalDataParser`/`get_input_embeddings` on the prompt-embeds path), because our request forges `image_grid_thw` metadata with the pixel bytes intact and reaches the Qwen2 vision RoPE/`cu_seqlens` and `image_embeds.split` sink, which the shape-check fix does not rebind. Site 4 is distinct from [GHSA-c65p-x677-fgj6](https://github.com/vllm-project/vllm/security/advisories/GHSA-c65p-x677-fgj6) (which folds metadata into `MultiModalHasher.serialize_item` to stop hash collisions), because the scale-out generate path trusts a caller-supplied `mm_hash` as the cache key with no origin binding, so that fix does not stop a caller from submitting a victim\u0027s hash or a `kwargs_data=None` cache read.\n\n## Affected code\n\nLinks pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34) (v0.25.1).\n\nShared entry point and control surface for all five sites:\n\n- `POST /inference/v1/generate` route: [`vllm/entrypoints/scale_out/token_in_token_out/api_router.py#L46-L75`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_out/token_in_token_out/api_router.py#L46-L75).\n- Public `features` schema (`kwargs_data`, `mm_hashes`, `mm_placeholders`): [`vllm/entrypoints/scale_out/token_in_token_out/protocol.py#L42-L63`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_out/token_in_token_out/protocol.py#L42-L63).\n- `ServingTokens.serve_tokens()` copies the decoded geometry and hashes into engine structures without rebinding to the renderer schema: [`vllm/entrypoints/scale_out/token_in_token_out/serving.py#L145-L172`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_out/token_in_token_out/serving.py#L145-L172).\n- The scale-out routers are registered by default on generate-capable servers and `/inference` is treated as an ordinary auth-guarded prefix: [`vllm/entrypoints/openai/api_server.py#L217-L219`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/api_server.py#L217-L219).\n\n**Site 1 \u2014 forged Qwen grid geometry (engine-fatal DoS).** The decoded `image_grid_thw` is never rebound to the rendered pixel-tensor element count.\n\n- Reconstruction with no model-specific geometry invariant: [`serving.py#L147-L172`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_out/token_in_token_out/serving.py#L147-L172).\n- `Qwen2VisionTransformer.prepare_encoder_metadata()` derives RoPE tables, `cu_seqlens`, and the FlashAttention `max_seqlen` from the caller-supplied grid: [`qwen2_vl.py#L658-L712`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/model_executor/models/qwen2_vl.py#L658-L712), consumed in [`forward()` (`#L713-L751`)](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/model_executor/models/qwen2_vl.py#L713-L751).\n- `_process_image_input()` computes `image_embeds.split(sizes)` from the same untrusted metadata: [`qwen2_vl.py#L1348-L1369`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/model_executor/models/qwen2_vl.py#L1348-L1369); M-RoPE positions at [`qwen2_vl.py#L1223`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/model_executor/models/qwen2_vl.py#L1223).\n\nThe only shape check is `assert grid_thw.ndim == 2`; the split sizes and the vision-encoder call are then derived directly from the caller-supplied grid, with no cross-check against the pixel-tensor row count:\n\n```python\n# vllm/model_executor/models/qwen2_vl.py Lines 1348-1369\n def _process_image_input(\n self, image_input: Qwen2VLImageInputs\n ) -\u003e tuple[torch.Tensor, ...]:\n grid_thw = image_input[\"image_grid_thw\"]\n assert grid_thw.ndim == 2\n\n if image_input[\"type\"] == \"image_embeds\":\n image_embeds = image_input[\"image_embeds\"]\n else:\n pixel_values = image_input[\"pixel_values\"]\n\n if self.use_data_parallel:\n return run_dp_sharded_mrope_vision_model(\n self.visual, pixel_values, grid_thw.tolist(), rope_type=\"rope_3d\"\n )\n else:\n image_embeds = self.visual(pixel_values, grid_thw=grid_thw)\n\n # Split concatenated embeddings for each image item.\n merge_size = self.visual.spatial_merge_size\n sizes = (grid_thw.prod(-1) // merge_size // merge_size).tolist()\n return image_embeds.split(sizes)\n```\n\n**Site 2 \u2014 wire-selected field-processor type confusion (engine-fatal DoS).** `MsgpackDecoder._decode_mm_field_elem()` trusts a wire-selected field-factory name and constructs the internal field processor directly from caller data.\n\n- [`vllm/v1/serial_utils.py#L440-L454`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/serial_utils.py#L440-L454) reads `factory_meth_name, factory_kw = obj[\"field\"]` and calls `getattr(MultiModalFieldConfig, factory_meth_name)`.\n- Qwen2-VL field/schema contract: [`qwen2_vl.py#L763-L791`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/model_executor/models/qwen2_vl.py#L763-L791) and [`Qwen2VLImagePixelInputs` (`#L119-L144`)](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/model_executor/models/qwen2_vl.py#L119-L144); parse/validate at [`qwen2_vl.py#L1300`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/model_executor/models/qwen2_vl.py#L1300).\n- Rank mismatch raised post-admission: [`vllm/utils/tensor_schema.py#L155-L171`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/utils/tensor_schema.py#L155-L171), reached from [`TensorSchema.__init__` \u2192 `validate()` (`#L63`)](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/utils/tensor_schema.py#L63).\n\n**Site 3 \u2014 non-positive placeholder length (engine-fatal DoS via reachable assert).** `PlaceholderRangeInfo{offset,length}` is accepted as unconstrained integers and copied verbatim into the engine\u0027s `PlaceholderRange`.\n\n- Schema: [`protocol.py#L28-L35`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_out/token_in_token_out/protocol.py#L28-L35).\n- Copied into `PlaceholderRange`: [`serving.py#L147-L153`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_out/token_in_token_out/serving.py#L147-L153).\n- The only guard is an upper bound on embed count (`num_embeds \u003e mm_encoder_cache_size`), with no non-positive check: [`vllm/v1/engine/input_processor.py#L459-L464`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/engine/input_processor.py#L459-L464); raw length returned by [`vllm/multimodal/inputs.py#L152-L154`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/inputs.py#L152-L154); window selection assumes non-empty ranges at [`vllm/multimodal/utils.py#L114-L134`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/utils.py#L114-L134).\n- Sink: `assert start_idx \u003c end_idx` at [`vllm/v1/worker/gpu/mm/encoder_runner.py#L114`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/worker/gpu/mm/encoder_runner.py#L114) (duplicated at [`vllm/v1/worker/gpu_model_runner.py#L3192`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/worker/gpu_model_runner.py#L3192)); turned into a fatal shutdown by EngineCore\u0027s uncaught-exception path at [`vllm/v1/engine/core.py#L1229-L1233`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/engine/core.py#L1229-L1233).\n\nA `length` of `0` makes `num_encoder_tokens == 0`, so `end_idx` collapses to `0` and the bare `assert` fires inside the engine worker:\n\n```python\n# vllm/v1/worker/gpu/mm/encoder_runner.py Lines 108-114\n pos_info = mm_feature.mm_position\n start_pos = pos_info.offset\n num_encoder_tokens = pos_info.length\n\n start_idx = max(cur_query_start - start_pos, 0)\n end_idx = min(cur_query_end - start_pos, num_encoder_tokens)\n assert start_idx \u003c end_idx\n```\n\n**Site 4 \u2014 cache hash not bound to payload (integrity / disclosure).** `features.mm_hashes` (the cache key) and `kwargs_data` (the tensor) are independent fields with no origin or integrity binding.\n\n- Schema exposing the caller-controlled hash and the `None` = resolve-from-cache semantics: [`protocol.py#L42-L63`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_out/token_in_token_out/protocol.py#L42-L63).\n- Handler forwards the caller\u0027s hashes unchanged into `mm_input(...)`: [`serving.py#L164-L170`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_out/token_in_token_out/serving.py#L164-L170).\n- Input processing copies the caller hash into the feature identifier with only a string-type check: [`vllm/v1/engine/input_processor.py#L165-L181`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/engine/input_processor.py#L165-L181).\n- Sink \u2014 the receiver cache returns the cached tensor solely by that key (`cache_key = feature.mm_hash or feature.identifier`): [`vllm/multimodal/cache.py#L602-L607`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L602-L607).\n\nThe cache key is the caller-supplied hash with no verification against the tensor bytes, so a forged `mm_hash` both stores under and reads back another request\u0027s slot:\n\n```python\n# vllm/multimodal/cache.py Lines 601-607\n for feature in mm_features:\n cache_key = feature.mm_hash or feature.identifier\n self.touch_receiver_cache_item(cache_key, feature.data)\n\n for feature in mm_features:\n cache_key = feature.mm_hash or feature.identifier\n feature.data = self.get_and_update_item(feature.data, cache_key)\n return mm_features\n```\n\n**Site 5 \u2014 dropped sparse placeholder mask (transport integrity loss).** The render path serializes placeholders as only `offset`/`length`, so models relying on sparse `is_embed` masks lose the mask during render-to-generate replay.\n\n- `ServingRender._extract_mm_features()` builds each `PlaceholderRangeInfo(offset=p.offset, length=p.length)`, discarding `is_embed`: [`vllm/entrypoints/scale_out/render/serving.py#L212-L229`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_out/render/serving.py#L212-L229).\n- Transport schema carries no field for the mask: [`protocol.py#L28`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_out/token_in_token_out/protocol.py#L28).\n- `ServingTokens.serve_tokens()` reconstructs a dense `PlaceholderRange` regardless of the original: [`serving.py#L148-L152`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/scale_out/token_in_token_out/serving.py#L148-L152).\n\n## Impact\n\nA single authenticated request to a scale-out multimodal deployment can:\n\n- **Crash the shared EngineCore process** (sites 1, 2, 3), taking the served model down for every tenant (`/health` \u2192 503). Availability-only; no code execution or data disclosure demonstrated for these sites.\n- **Silently poison or read back another request\u0027s shared encoder-cache state** (site 4) \u2014 an integrity/disclosure primitive. Attack complexity is High because the attacker must know or induce the victim\u0027s content hash; there is no availability impact for this site.\n- **Corrupt backend-visible placeholder semantics across the render-to-generate boundary** (site 5) for models that depend on sparse `is_embed` masks.\n\nThe forged multimodal payload is small; only the trust in its self-declared geometry/identity is the defect.\n\n\n## Suggested Fix\n\nOn the scale-out path, do not trust caller-supplied multimodal state as renderer-produced. After decoding `features`, **rebind and validate** the reconstructed `MultiModalKwargsItem` against the active model\u0027s renderer contract at the HTTP boundary:\n\n1. Reject any request whose decoded grid geometry is inconsistent with the rendered pixel-tensor element count and declared placeholder span, before it reaches `prepare_encoder_metadata()` (site 1).\n2. Rebind each field\u0027s processor type to the schema the active model\u0027s renderer declares (or reject if it differs), instead of reconstructing internal field processors from wire-selected factory names (site 2).\n3. Reject any `PlaceholderRangeInfo` with `length \u003c= 0` (or out-of-range `offset`) with a request-scoped 4xx, and convert the encoder-runner invariant into a checked, request-scoped error rather than a process-fatal `assert` (site 3).\n4. Recompute or verify the content hash for submitted `kwargs_data` before using it as a cache key, and namespace receiver-cache keys to a server-generated or principal scope, refusing cache-reads for hashes the caller did not legitimately produce (site 4).\n5. Serialize `is_embed` in `PlaceholderRangeInfo`, validate its length against the placeholder span, and reconstruct it on replay (site 5).\n\n**Site 1 \u2014 validate grid geometry before the vision encoder.** Replace the bare `assert grid_thw.ndim == 2` in `_process_image_input()`/`_process_video_input()` with a shared helper that recomputes the split sizes and rejects a grid whose patch-row count does not match the pixel tensor (and rejects non-positive / non-merge-divisible dims), so the mismatch never reaches `image_embeds.split()`:\n\n```python\n# vllm/model_executor/models/qwen2_vl.py \u2014 _process_image_input()\n- grid_thw = image_input[\"image_grid_thw\"]\n- assert grid_thw.ndim == 2\n+ grid_thw = image_input[\"image_grid_thw\"]\n+ input_type = image_input[\"type\"]\n+ input_tensor = (\n+ image_input[\"image_embeds\"]\n+ if input_type == \"image_embeds\"\n+ else image_input[\"pixel_values\"]\n+ )\n+ sizes = _validate_qwen2_vl_input_geometry(\n+ modality=\"image\",\n+ input_type=input_type,\n+ input_tensor=input_tensor,\n+ grid_thw=grid_thw,\n+ spatial_merge_size=self.visual.spatial_merge_size,\n+ )\n...\n- # Split concatenated embeddings for each image item.\n- merge_size = self.visual.spatial_merge_size\n- sizes = (grid_thw.prod(-1) // merge_size // merge_size).tolist()\n return image_embeds.split(sizes)\n```\n\nwhere the helper raises before the encoder runs:\n\n```python\n# vllm/model_executor/models/qwen2_vl.py \u2014 new _validate_qwen2_vl_input_geometry()\n if t \u003c= 0 or h \u003c= 0 or w \u003c= 0:\n raise ValueError(f\"{modality} grid_thw row {index} must be positive ...\")\n if h % spatial_merge_size != 0 or w % spatial_merge_size != 0:\n raise ValueError(f\"{modality} grid_thw row {index} must be divisible ...\")\n ...\n if actual_rows != expected_rows:\n raise ValueError(\n f\"{modality} {row_kind} do not match grid_thw: \"\n f\"expected {expected_rows}, got {actual_rows}.\"\n )\n```\n\n**Site 3 \u2014 constrain the placeholder schema.** Make `PlaceholderRangeInfo` reject non-positive lengths and negative offsets at the Pydantic boundary (plus parallel-length, non-overlapping, and within-prompt validators), turning the process-fatal `assert` into a request-scoped 422:\n\n```python\n# vllm/entrypoints/scale_out/token_in_token_out/protocol.py \u2014 PlaceholderRangeInfo\n- offset: int\n- length: int\n+ offset: int = Field(ge=0)\n+ length: int = Field(gt=0)\n```\n\n**Site 4 \u2014 bind the cache key to the payload.** Derive each cache key by hashing the submitted serialized tensor (ignoring the caller\u0027s `mm_hashes`) and refuse cache-only reads, so a forged hash can neither poison nor read a victim\u0027s slot:\n\n```python\n# vllm/entrypoints/scale_out/token_in_token_out/serving.py \u2014 serve_tokens()\n+ mm_hashes = _bind_mm_hashes_to_kwargs_data(\n+ features.mm_hashes, features.kwargs_data,\n+ )\n engine_input = mm_input(\n prompt_token_ids=request.token_ids,\n mm_kwargs=MultiModalKwargsItems(mm_kwargs),\n- mm_hashes=features.mm_hashes,\n+ mm_hashes=mm_hashes,\n mm_placeholders=mm_placeholders,\n cache_salt=request.cache_salt,\n )\n```\n\nwhere `_bind_mm_hashes_to_kwargs_data()` raises on `kwargs_data is None` (cache-only read) and derives `sha256(modality || \"\\0\" || serialized_item)` per item. Site 2 applies the same rebind-to-declared-schema pattern in `mm_serde.py` (passing `modality` + `mm_processor` into `decode_mm_kwargs_item`), and site 5 adds an `is_embed` field to `PlaceholderRangeInfo` with `from_placeholder_range`/`to_placeholder_range` helpers so the sparse mask survives render-to-generate replay. Each site fix ships with a regression test. This packet groups the five sites because they share one entry point (`/inference/v1/generate` + `/render`) and one root cause; we are happy to split it into per-component advisories (for example, engine-fatal input-validation vs. cache-key binding vs. transport-schema integrity) if the vLLM team prefers.\n\n## Credit\n\n**Reported by:** Patch the Planet (Trail of Bits + OpenAI collaboration)\n\nThese vulnerabilities were discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.\n\n---\n\n**Proposed fix:** a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51898",
"id": "GHSA-ph72-cqr5-qpp7",
"modified": "2026-10-05T23:42:50Z",
"published": "2026-10-05T23:42:50Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/vllm-project/vllm/security/advisories/GHSA-ph72-cqr5-qpp7"
},
{
"type": "WEB",
"url": "https://github.com/vllm-project/vllm/pull/51898"
},
{
"type": "WEB",
"url": "https://github.com/vllm-project/vllm/commit/1970f3ed4be7fa8620e4ddc4a12c36a8384cfc27"
},
{
"type": "PACKAGE",
"url": "https://github.com/vllm-project/vllm"
},
{
"type": "WEB",
"url": "https://github.com/vllm-project/vllm/releases/tag/v0.30.0"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H",
"type": "CVSS_V3"
}
],
"summary": "vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features"
}
GHSA-PH76-MMF2-83HF
Vulnerability from github – Published: 2022-05-17 00:16 – Updated: 2022-05-17 00:16Some Huawei smartphones with software AGS-L09C233B019,AGS-W09C233B019,KOB-L09C233B017,KOB-W09C233B012 have a type confusion vulnerability. The program initializes a variable using one type, but it later accesses that variable using a type that is different with the original type when do certain register operation. Successful exploit could result in buffer overflow then may cause malicious code execution.
{
"affected": [],
"aliases": [
"CVE-2017-8159"
],
"database_specific": {
"cwe_ids": [
"CWE-704"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2017-11-22T19:29:00Z",
"severity": "HIGH"
},
"details": "Some Huawei smartphones with software AGS-L09C233B019,AGS-W09C233B019,KOB-L09C233B017,KOB-W09C233B012 have a type confusion vulnerability. The program initializes a variable using one type, but it later accesses that variable using a type that is different with the original type when do certain register operation. Successful exploit could result in buffer overflow then may cause malicious code execution.",
"id": "GHSA-ph76-mmf2-83hf",
"modified": "2022-05-17T00:16:17Z",
"published": "2022-05-17T00:16:17Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2017-8159"
},
{
"type": "WEB",
"url": "http://www.huawei.com/en/psirt/security-advisories/2017/huawei-sa-20171018-02-smartphone-en"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.0/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H",
"type": "CVSS_V3"
}
]
}
GHSA-PMHC-XXM7-J8VV
Vulnerability from github – Published: 2022-05-13 01:34 – Updated: 2022-05-13 01:34This vulnerability allows remote attackers to execute arbitrary code on vulnerable installations of Foxit Reader 9.0.1.1049. User interaction is required to exploit this vulnerability in that the target must visit a malicious page or open a malicious file. The specific flaw exists within the addPageOpenJSMessage method. By performing actions in JavaScript, an attacker can trigger a type confusion condition. The attacker can leverage this vulnerability to execute code under the context of the current process. Was ZDI-CAN-6006.
{
"affected": [],
"aliases": [
"CVE-2018-14243"
],
"database_specific": {
"cwe_ids": [
"CWE-704"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2018-07-31T20:29:00Z",
"severity": "HIGH"
},
"details": "This vulnerability allows remote attackers to execute arbitrary code on vulnerable installations of Foxit Reader 9.0.1.1049. User interaction is required to exploit this vulnerability in that the target must visit a malicious page or open a malicious file. The specific flaw exists within the addPageOpenJSMessage method. By performing actions in JavaScript, an attacker can trigger a type confusion condition. The attacker can leverage this vulnerability to execute code under the context of the current process. Was ZDI-CAN-6006.",
"id": "GHSA-pmhc-xxm7-j8vv",
"modified": "2022-05-13T01:34:41Z",
"published": "2022-05-13T01:34:41Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2018-14243"
},
{
"type": "WEB",
"url": "https://www.foxitsoftware.com/support/security-bulletins.php"
},
{
"type": "WEB",
"url": "https://zerodayinitiative.com/advisories/ZDI-18-703"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.0/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H",
"type": "CVSS_V3"
}
]
}
GHSA-PMQ5-6QCV-88Q4
Vulnerability from github – Published: 2022-05-17 03:04 – Updated: 2022-05-17 03:04Adobe Acrobat Reader versions 15.020.20042 and earlier, 15.006.30244 and earlier, 11.0.18 and earlier have an exploitable type confusion vulnerability in the XSLT engine related to localization functionality. Successful exploitation could lead to arbitrary code execution.
{
"affected": [],
"aliases": [
"CVE-2017-2962"
],
"database_specific": {
"cwe_ids": [
"CWE-704"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2017-01-11T04:59:00Z",
"severity": "HIGH"
},
"details": "Adobe Acrobat Reader versions 15.020.20042 and earlier, 15.006.30244 and earlier, 11.0.18 and earlier have an exploitable type confusion vulnerability in the XSLT engine related to localization functionality. Successful exploitation could lead to arbitrary code execution.",
"id": "GHSA-pmq5-6qcv-88q4",
"modified": "2022-05-17T03:04:01Z",
"published": "2022-05-17T03:04:01Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2017-2962"
},
{
"type": "WEB",
"url": "https://helpx.adobe.com/security/products/acrobat/apsb17-01.html"
},
{
"type": "WEB",
"url": "http://www.securityfocus.com/bid/95340"
},
{
"type": "WEB",
"url": "http://www.securitytracker.com/id/1037574"
},
{
"type": "WEB",
"url": "http://www.zerodayinitiative.com/advisories/ZDI-17-026"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.0/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H",
"type": "CVSS_V3"
}
]
}
GHSA-PPGQ-34JH-MHCX
Vulnerability from github – Published: 2022-05-13 01:33 – Updated: 2022-05-13 01:33This vulnerability allows remote attackers to execute arbitrary code on vulnerable installations of Foxit Reader 9.2.0.9297. User interaction is required to exploit this vulnerability in that the target must visit a malicious page or open a malicious file. The specific flaw exists within the handling of PDF files. The issue results from the lack of proper validation of user-supplied data, which can result in a type confusion condition. An attacker can leverage this vulnerability to execute code in the context of the current process. Was ZDI-CAN-6819.
{
"affected": [],
"aliases": [
"CVE-2018-17685"
],
"database_specific": {
"cwe_ids": [
"CWE-704"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2019-01-24T04:29:00Z",
"severity": "HIGH"
},
"details": "This vulnerability allows remote attackers to execute arbitrary code on vulnerable installations of Foxit Reader 9.2.0.9297. User interaction is required to exploit this vulnerability in that the target must visit a malicious page or open a malicious file. The specific flaw exists within the handling of PDF files. The issue results from the lack of proper validation of user-supplied data, which can result in a type confusion condition. An attacker can leverage this vulnerability to execute code in the context of the current process. Was ZDI-CAN-6819.",
"id": "GHSA-ppgq-34jh-mhcx",
"modified": "2022-05-13T01:33:51Z",
"published": "2022-05-13T01:33:51Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2018-17685"
},
{
"type": "WEB",
"url": "https://www.foxitsoftware.com/support/security-bulletins.php"
},
{
"type": "WEB",
"url": "https://www.zerodayinitiative.com/advisories/ZDI-18-1204"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.0/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H",
"type": "CVSS_V3"
}
]
}
No mitigation information available for this CWE.
No CAPEC attack patterns related to this CWE.