<?xml version='1.0' encoding='UTF-8'?>
<?xml-stylesheet href="/static/style.xsl" type="text/xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <id>https://cve.radiocsirt.org/rss/recent/all/10</id>
  <title>Most recent entries from all</title>
  <updated>2026-10-02T19:37:57.066079+00:00</updated>
  <author>
    <name>Vulnerability-Lookup</name>
    <email>csirt@opendfir.org</email>
  </author>
  <link href="https://cve.radiocsirt.org" rel="alternate"/>
  <generator uri="https://lkiesow.github.io/python-feedgen" version="1.0.0">python-feedgen</generator>
  <subtitle>Contains only the most 10 recent entries.</subtitle>
  <entry>
    <id>https://cve.radiocsirt.org/vuln/cve-2026-46497</id>
    <title>CVE-2026-46497 — SSRF via sitemap-derived URLs in Crawlee for Python</title>
    <updated>2026-10-02T19:37:57.098102+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> apify crawlee-python</p>
<p>Crawlee is a web scraping and browser automation library. From version 1.0.0 to before version 1.7.0, Crawlee is vulnerable to SSRF via sitemap-derived URLs. This issue has been patched in version 1.7.0.</p></div>
    </content>
    <link href="https://cve.radiocsirt.org/vuln/cve-2026-46497"/>
  </entry>
  <entry>
    <id>https://cve.radiocsirt.org/vuln/ghsa-3r75-xc34-5f44</id>
    <title>GHSA-3r75-xc34-5f44 — Crawlee for Python: SSRF via sitemap-derived URLs</title>
    <updated>2026-10-02T19:37:57.098161+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> PyPI: crawlee</p>
<p>## Overview</p>
<p>- **Vulnerability type:** Blind SSRF
- **Affected components:** `src/crawlee/_utils/sitemap.py`, `src/crawlee/_utils/robots.py`, `src/crawlee/request_loaders/_sitemap_request_loader.py`, and all built-in HTTP clients.
- **Trigger:** an attacker-controlled sitemap or `robots.txt` containing a URL that points to an internal host (layer 1) or uses a non-http scheme (layer 2).</p>
<p>Two-layer SSRF via sitemap-derived URLs:</p>
<p>### 1) Cross-host HTTP SSRF</p>
<p>Base case, affects every HTTP client.** Sitemap entries and `robots.txt` `Sitemap:` directives were accepted regardless of the host they pointed to. A sitemap on `example.com` could push `http://internal.corp/admin` into the crawler's queue, and the configured HTTP client would dispatch the request.</p>
<p>### 2) Non-HTTP scheme SSRF</p>
<p>Escalation, only `CurlImpersonateHttpClient`.** Nested-sitemap fetching dispatches the URL straight to the HTTP client, bypassing the `Request` construction step where Pydantic enforces `http(s)`. Combined with the libcurl-backed `CurlImpersonateHttpClient`, this lets `gopher://`, `file://`, `dict://`, `ftp://`, etc., through.</p>
<p>## Root cause</p>
<p>Crawlee already validates URL schemes through Pydantic's `AnyHttpUrl` (via `validate_http_url` in `src/crawlee/_utils/urls.py`) wherever a crawl target is materialised as a `Request`: the `Request.url` field is declared as `Annotated[str, BeforeValidator(validate_http_url), Field(frozen=True)]`. Anything that becomes a `Request` is therefore guaranteed to be…</p></div>
    </content>
    <link href="https://cve.radiocsirt.org/vuln/ghsa-3r75-xc34-5f44"/>
  </entry>
</feed>
