<?xml version='1.0' encoding='UTF-8'?>
<?xml-stylesheet href="/static/style.xsl" type="text/xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <id>https://cve.radiocsirt.org/rss/recent/all/10</id>
  <title>Most recent entries from all</title>
  <updated>2026-10-06T09:35:13.687636+00:00</updated>
  <author>
    <name>Vulnerability-Lookup</name>
    <email>csirt@opendfir.org</email>
  </author>
  <link href="https://cve.radiocsirt.org" rel="alternate"/>
  <generator uri="https://lkiesow.github.io/python-feedgen" version="1.0.0">python-feedgen</generator>
  <subtitle>Contains only the most 10 recent entries.</subtitle>
  <entry>
    <id>https://cve.radiocsirt.org/vuln/brew-scrapy-cve-2026-55520</id>
    <title>BREW-scrapy-CVE-2026-55520 — Protego has exponential backtracking ReDoS in robots.txt URL wildcard matching</title>
    <updated>2026-10-06T09:35:13.689956+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> Homebrew: scrapy</p>
<p>### Problem description</p>
<p>Protego constructs regular expressions to match URLs against `robots.txt` `Allow:` and `Disallow:` directives, see `protego._urlpattern._URLPattern._prepare_pattern_for_regex()`. Every `*` in the directive value is translated into a lazy `.*?` regex piece, thus a specially crafted directive value with many asterisks may produce a regex that freezes the parser due to exponential backtracking.</p>
<p>### Impact</p>
<p>Parsing a specially crafted `robots.txt` with `protego.Protego.parse()` and then trying to match an URL with `protego.Protego.can_fetch()` results in the latter call not returning for a period dependent on the length of the URL.</p>
<p>### Proof of concept</p>
<p>```python
from protego import Protego</p>
<p>robotstxt = f"""
User-agent: *
Disallow: /{"*1" * 12}*Z
"""
rp = Protego.parse(robotstxt)
url = "/" + "1" * 60
rp.can_fetch(url, "mybot")  # freezes
```</p></div>
    </content>
    <link href="https://cve.radiocsirt.org/vuln/brew-scrapy-cve-2026-55520"/>
  </entry>
  <entry>
    <id>https://cve.radiocsirt.org/vuln/cve-2026-55520</id>
    <title>CVE-2026-55520 — Protego: Exponential backtracking ReDoS in robots.txt URL wildcard matching</title>
    <updated>2026-10-06T09:35:13.690012+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> scrapy protego</p>
<p>Protego is a pure-Python robots.txt parser with support for modern conventions. Prior to 0.6.2, protego._urlpattern._URLPattern._prepare_pattern_for_regex translates every asterisk in an Allow or Disallow directive into a lazy regular-expression wildcard, so a directive containing many asterisks creates exponential backtracking. After protego.Protego.parse processes a crafted robots.txt file, protego.Protego.can_fetch can spend an attacker-controlled period matching a near-miss URL and deny service to the crawler. The vulnerable path is src/protego/_urlpattern.py in the _URLPattern match logic. This issue is fixed in version 0.6.2.</p></div>
    </content>
    <link href="https://cve.radiocsirt.org/vuln/cve-2026-55520"/>
  </entry>
  <entry>
    <id>https://cve.radiocsirt.org/vuln/ghsa-wjmf-p669-5m5p</id>
    <title>GHSA-wjmf-p669-5m5p — Protego has exponential backtracking ReDoS in robots.txt URL wildcard matching</title>
    <updated>2026-10-06T09:35:13.690043+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> PyPI: Protego</p>
<p>### Problem description</p>
<p>Protego constructs regular expressions to match URLs against `robots.txt` `Allow:` and `Disallow:` directives, see `protego._urlpattern._URLPattern._prepare_pattern_for_regex()`. Every `*` in the directive value is translated into a lazy `.*?` regex piece, thus a specially crafted directive value with many asterisks may produce a regex that freezes the parser due to exponential backtracking.</p>
<p>### Impact</p>
<p>Parsing a specially crafted `robots.txt` with `protego.Protego.parse()` and then trying to match an URL with `protego.Protego.can_fetch()` results in the latter call not returning for a period dependent on the length of the URL.</p>
<p>### Proof of concept</p>
<p>```python
from protego import Protego</p>
<p>robotstxt = f"""
User-agent: *
Disallow: /{"*1" * 12}*Z
"""
rp = Protego.parse(robotstxt)
url = "/" + "1" * 60
rp.can_fetch(url, "mybot")  # freezes
```</p></div>
    </content>
    <link href="https://cve.radiocsirt.org/vuln/ghsa-wjmf-p669-5m5p"/>
  </entry>
</feed>
