OPENSUSE-SU-2026:21900-1
Vulnerability from csaf_opensuse - Published: 2026-09-21 12:46 - Updated: 2026-09-24 17:48Summary
Security update for jsoup, re2j
Severity
Important
Notes
Title of the patch: Security update for jsoup, re2j
Description of the patch: This update for jsoup, re2j fixes the following issue:
- CVE-2026-75140: The builder copies the entire inherited namespace map on every start element, causing quadratic time
and memory complexity, which attackers can exploit to trigger an OutOfMemoryError and terminate the application
(bsc#1275912).
Changes for jsoup:
- Upgrade to upstream version 1.23.2
+ Improvement: Improved consecutive StreamParser.selectFirst()
calls during progressive parsing, so later matches are
returned with their parsed contents when earlier selections
had left them as parser lookahead. E.g., given
<title>One</title><p id=hit>Full</p><p>Next</p>, selecting
title and then #hit now advances the partial lookahead and
returns <p id="hit">Full</p>, rather than returning an empty
<p id="hit"></p> before its content is parsed. The updated
readiness tracking follows StreamParser's normal emission
order across implicit HTML structure and parser recovery.
+ Improvement: Improved XML parser performance and memory use
for documents with many nested namespace declarations by
recording namespace changes within each element scope
(bsc#1275912, CVE-2026-75140).
+ Improvement: Improved W3CDom conversion performance for
documents with many nested namespace declarations. The W3C
converter now uses the same optimized namespace tracking as
the XML parser.
+ Improvement: Improved W3CDom XML conversion to retain
processing instructions, comments outside the root element,
and CDATA sections, which were previously dropped or converted
to text.
+ Improvement: DOM mutation methods, including child insertion
and replacement, now reject operations that would create a
cycle, such as making a node its own child or moving an
ancestor beneath a descendant.
+ Improvement: Added Elements#before(Node), after(Node),
prepend(Node), and append(Node) to match the existing HTML
string methods.
+ Improvement: Large file-backed uploads through
Connection.requestBodyStream(InputStream) now stream directly
with the JDK HttpClient on Java 11+, rather than being loaded
fully into memory first.
+ Improvement: Extended Java 11+ HTTP client reuse from requests
sharing a Jsoup.newSession() to ordinary Jsoup.connect()
calls, reducing transport thread and connection setup churn
under sustained request loads. Sessions with custom
authentication or SSL contexts continue to use their own
client.
+ Change: Aligned the XML parser stack depth and lookups to the
configured maximum, which now defaults to 512 for both HTML
and XML. Use Parser#setMaxDepth(int) to configure.
+ Bugfix: Fixed W3CDom namespace conversion in several cases:
- Namespace declarations and prefixed attributes now carry the
correct namespace URI, so namespace-aware DOM lookups work
as expected.
- Attributes added after parsing, or included through subtree
conversion, now use inherited prefix declarations.
- Namespace declarations now apply regardless of attribute
order, and an empty declaration shadows an inherited binding
only within its scope.
- With namespace awareness disabled, inherited and undeclared
prefixes now receive the declarations needed for XML
serialization.
- Valid HTML names that are not XML QNames, such as a:b:c, are
normalized. Attributes that still cannot be represented are
skipped, and unrepresentable elements no longer change the
surrounding tree.
+ Bugfix: Fixed W3CDom conversion of programmatically created or
renamed elements whose names can be represented in a jsoup
HTML DOM but are not valid XML names, such as 1abc. These
names are now normalized (e.g. _1abc) instead of causing a
NullPointerException.
+ Bugfix: Fixed XML doctype serialization when a system
identifier contains a double quote, which could otherwise
produce invalid XML.
+ Bugfix: XML serialization now repairs element and attribute
names that start with an invalid character, rather than
outputting null elements or dropping attributes. For example,
an attribute named 1a is written as _1a. Additional leading
underscores keep repaired attribute names unique if they
conflict with another attribute.
+ Bugfix: Supplementary Unicode characters are now escaped
correctly when serializing with non-UTF, non-ASCII output
charsets such as ISO-8859-1. Previously, characters could be
emitted unescaped when their low 16-bit value was
representable by the configured charset, causing replacement
or corruption when the output was encoded.
+ Bugfix: Fixed the JDK HttpClient implementation to accept
responses missing a Content-Type header, matching the
HttpURLConnection implementation.
+ Bugfix: Fixed HTTP response content-type matching to handle
media types case-insensitively and recognize structured +xml
suffixes, including vendor-specific media types.
+ Bugfix: HTTP request URL normalization now percent-encodes
ASCII control characters, DEL, and embedded fragment
delimiters, keeping normalized URLs valid for HTTP requests
while preserving existing escapes.
+ Bugfix: Corrected multipart form encoding to percent-escape CR
and LF in field names and filenames, matching the HTML form
submission specification. Multipart file content-types
containing CR or LF are now rejected with a
ValidationException.
+ Bugfix: Aligned trailing comment placement with the HTML
specification: comments after </body> remain children of the
html element, while comments after </html> remain children of
the document.
+ Bugfix: When using the optional re2j regular expression
engine, memory allocation errors caused by complex selector
patterns at match time are now normalized to a
ValidationException with a Pattern complexity error message.
+ Bugfix: Fixed parsing of malformed SVG and MathML content so
that breakout HTML tags are placed according to the HTML
specification.
+ Bugfix: Fixed deeply nested malformed HTML parsing that could
lose the document body because stack lookups did not align to
the configured maximum parser depth.
+ Bugfix: Aligned RCDATA, RAWTEXT, and script-data parsing with
the HTML specification: malformed end tags no longer consume
following markup, unclosed title/textarea content stays text
through EOF, and custom text tags match exact names.
+ Bugfix: Improved URL validation during HTTP/HTTPS URL
resolution and cleaning; resolved URLs without a host are now
rejected instead of being accepted based only on their scheme
prefix, aligning to RFC 9110. Valid relative links and
non-HTTP(S) schemes are unchanged.
+ Bugfix: Redirects with malformed single-slash HTTP locations
now use standard URL resolution to align with browsers.
+ Bugfix: Template fragment parsing now handles unmatched
</template> tags without throwing a ValidationException.
+ Bugfix: Improved source tracking for adopted formatting
elements and malformed markup ending at EOF.
* Changes of 1.23.1
+ Improvement: Reduced retained memory when parsing with source
position tracking enabled (Parser#setTrackPosition(true)).
Source ranges are now stored in compact parser-owned span
records instead of node and attribute user data, and Position
objects are created lazily when source ranges are read. This
cuts tracked DOM retained size by about 50-60% on
representative benchmark documents, while keeping
Node#sourceRange(), Element#endSourceRange(), and
Attribute#sourceRange() behavior intact.
+ Improvement: Added Element#classList(), an immutable snapshot
of an element's class names in attribute order. Use hasClass()
when you just need to test for one class, classList() when you
want to read or iterate classes without needing a mutable
result, and classNames() when you want the existing mutable,
deduplicated set that can be written back with
classNames(Set). The class APIs now share an HTML-whitespace
scanner, which also makes classNames() faster and lighter on
allocation, especially when walking many elements without
class names.
+ Improvement: Aligned HTML parser scope classification with the
current HTML spec for select, foreignObject, and template.
+ Improvement: Simplified the HTML tree builder's scope,
implied-end-tag, and special-element checks by caching
parser-only options on Tag. That improves HTML parser
throughput by about 10% on small inputs and up to about 30% on
larger inputs in the benchmark fixtures.
+ Improvement: Improved HTML parser throughput stability by
making hot tokeniser scan paths compile more predictably.
+ Improvement: <noscript> fallback markup is now parsed into an
inspectable DOM subtree in both the document head and body.
The fallback acts as a contained parsing island, so malformed
markup cannot disrupt the surrounding document structure,
while normal HTML tokenization still applies within it. This
also improves round-trip serialization.
+ Improvement: Improved redirect credential handling as a
defense-in-depth measure: explicit authorization headers and
request cookies are no longer forwarded across origins,
reducing exposure through open redirects and aligning with
HTTP guidance. Cookies managed by a CookieStore continue to
follow their configured scope.
+ Improvement: Elements can now append their outer HTML,
including their own tags, directly to an Appendable with
Node#outerHtml(Appendable), without first creating a String.
This complements Element#html(Appendable), which appends inner
HTML only.
+ Improvement: Aligned CDATA tokenization with the HTML spec:
CDATA syntax in HTML content is parsed as a bogus comment,
while it remains supported in SVG, MathML, and XML. Also
improved namespace-aware fragment parsing so SVG and MathML
contexts, HTML integration points, and context-sensitive
tokenizer states are handled correctly.
+ Improvement: When using the optional re2j regular expression
engine, stack overflows caused by complex selector patterns
are now normalized to a ValidationException with a Pattern
complexity error message.
+ Bugfix: Fixed HTML parsing of mixed-case RCDATA end tags after
tag-shaped text. For example, <title><p>Foo</TiTLE> and
<textarea><img src=x></TeXtArEa> now keep the tag-shaped
content as text instead of promoting it to markup.
+ Bugfix: Fixed W3CDom XML conversion so plain XML elements
don't serialize with the reserved XML namespace as the default
namespace. Explicit XML namespaces and xml:* attributes are
still preserved.
+ Bugfix: Preserve control characters in parsed tag names.
+ Bugfix: Updated HTTP redirects to follow the specification:
307 and 308 preserve the request method and content, 301 and
302 only change POST to GET, and Location is followed only for
301, 302, 303, 307, and 308 responses. Streamed request bodies
are not buffered; if an automatic redirect requires replaying
one, execution fails, so the caller can resend with a fresh
stream.
+ Bugfix: Corrected the Cleaner's same-site link detection to
compare hostnames rather than URL prefixes when applying
rel=nofollow.
+ Build change: Cleaned up the Maven build for the multi-release
JAR so Java 8 and Java 11+ sources compile as separate source
sets. This avoids spurious Java 8 compiler warnings from
newer-language overlay sources, keeps long-running parser
checks behind an explicit profile, and preserves the same
published artifacts and runtime behavior.
+ Build change: Improved parallelism and tuned timing in our
integration tests, so that a full mvn clean verify drops from
~ 1m18s to ~ 21 seconds.
* Changes of 1.22.2
+ Improvement: Expanded and clarified NodeTraversor support for
in-place DOM rewrites during NodeVisitor.head(). Current-node
edits such as remove, replace, and unwrap now recover more
predictably, while traversal stays within the original root
subtree. This makes single-pass tree cleanup and normalization
visitors easier to write, for example when unwrapping
presentational elements or replacing text nodes as you walk
the DOM.
+ Documentation: clarified that a configured Cleaner may be
reused across concurrent threads, and that shared Safelist
instances should not be mutated while in use.
+ Improvement: Updated the default HTML TagSet for current HTML
elements: added dialog, search, picture, and slot; made ins,
del, button, audio, video, and canvas inline by default
(Tag#isInline(), aligned to phrasing content in the spec); and
added readable Element.text() boundaries for controls and
embedded objects via the new Tag.TextBoundary option. This
improves pretty-printing and keeps normalized text from
running adjacent words together.
+ Android (R8/ProGuard): added a rule to ignore the optional
re2j dependency when not present.
+ Bugfix: Fixed a NodeTraversor regression in 1.21.2 where
removing or replacing the current node during head() could
revisit the replacement node and loop indefinitely. The
traversal docs now also clarify which inserted nodes are
visited in the current pass.
+ Bugfix: Parsing during charset sniffing no longer fails if an
advisory available() call throws IOException, as seen on JDK 8
HttpURLConnection.
+ Bugfix: Cleaner no longer makes relative URL attributes in the
input document absolute when cleaning or validating a
Document. URL normalization now applies only to the cleaned
output, and Safelist.isSafeAttribute() is side effect free.
+ Bugfix: Cleaner no longer duplicates enforced attributes when
the input Document preserves attribute case. A case-variant
source attribute is now replaced by the enforced attribute in
the cleaned output.
+ Bugfix: If a per-request SOCKS proxy is configured, jsoup now
avoids using the JDK HttpClient, because the JDK would
silently ignore that proxy and attempt to connect directly.
Those requests now fall back to the legacy HttpURLConnection
transport instead, which does support SOCKS.
+ Bugfix: Connection.Response.streamParser() and
DataUtil.streamParser(Path, ...) could fail on small inputs
without a declared charset, if the initial 5 KB charset sniff
fully consumed the input and closed it before the stream parse
began.
+ Bugfix: In XML mode, doctypes with an internal subset, such as
<!DOCTYPE root [<!ENTITY name "value">]>, now round-trip
correctly. The subset is preserved as raw text only; entities
are not expanded and external DTDs are not loaded.
+ Build change: Migrated the integration test server from Jetty
to Netty, which actively maintains support for our minimum JDK
target (8).
* Changes of 1.22.1
+ Improvement: Added support for using the re2j regular
expression engine for regex-based CSS selectors (e.g.
[attr~=regex], :matches(regex)), which ensures linear-time
performance for regex evaluation. This allows safer handling
of arbitrary user-supplied query regexes. To enable, add the
com.google.re2j dependency to your classpath, e.g.:
<dependency>
<groupId>com.google.re2j</groupId>
<artifactId>re2j</artifactId>
<version>1.8</version>
</dependency>
(If you already have that dependency in your classpath, but
you want to keep using the Java regex engine, you can disable
re2j via System.setProperty("jsoup.useRe2j", "false").) You
can confirm that the re2j engine has been enabled correctly by
calling Regex.usingRe2j().
+ Improvement: Added an instance method Parser#unescape(String,
boolean) that unescapes HTML entities using the parser's
configuration (e.g. to support error tracking), complementing
the existing static utility Parser.unescapeEntities(String,
boolean).
+ Improvement: Added a configurable maximum parser depth (to
limit the number of open elements on stack) to both HTML and
XML parsers. The HTML parser now defaults to a depth of 512 to
match browser behavior, and protect against unbounded stack
growth, while the XML parser keeps unlimited depth by default,
but can opt into a limit via Parser.setMaxDepth().
+ Build: added CI coverage for JDK 25.
+ Build: added a CI fuzzer for contextual fragment parsing (in
addition to existing full body HTML and XML fuzzers).
+ Change: Set a removal schedule of jsoup 1.24.1 for previously
deprecated APIs.
+ Bugfix: Previously cached child Elements of an Element were
not correctly invalidated in Node#replaceWith(Node), which
could lead to incorrect results when subsequently calling
Element#children().
+ Bugfix: Attribute selector values are now compared literally
without trimming. Previously, jsoup trimmed whitespace from
selector values and from element attribute values, which could
cause mismatches with browser behavior (e.g. [attr=" foo "]).
Now matches align with the CSS specification and browser
engines.
+ Bugfix: When using the JDK HttpClient, any system default
proxy (ProxySelector.getDefault()) was ignored. Now, the
system proxy is used if a per-request proxy is not set.
+ Bugfix: A ValidationException could be thrown in the adoption
agency algorithm with particularly broken input. Now logged as
a parse error.
+ Bugfix: Null characters in the HTML body were not consistently
removed; and in foreign content were not correctly replaced.
+ Bugfix: An IndexOutOfBoundsException could be thrown when
parsing a body fragment with crafted input. Now logged as a
parse error.
+ Bugfix: When using StructuralEvaluators (e.g., a parent child
selector) across many retained threads, their memoized results
could also be retained, increasing memory use. These results
are now cleared immediately after use, reducing overall memory
consumption.
+ Bugfix: Cloning a Parser now preserves any custom TagSet
applied to the parser.
+ Bugfix: Custom tags marked as Tag.Void now parse and serialize
like the built-in void elements: they no longer consume
following content, and the XML serializer emits the expected
self-closing form.
+ Bugfix: The <br> element is once again classified as an inline
tag (Tag.isBlock() == false), matching common developer
expectations and its role as phrasing content in HTML, while
pretty-printing and text extraction continue to treat it as a
line break in the rendered output.
+ Bugfix: Fixed an intermittent truncation issue when fetching
and parsing remote documents via Jsoup.connect(url).get(). On
responses without a charset header, the initial charset sniff
could sometimes (depending on buffering / available()
behavior) be mistaken for end-of-stream and a partial parse
reused, dropping trailing content.
+ Bugfix: TagSet copies no longer mutate their template during
lazy lookups, preventing cross-thread
ConcurrentModificationException when parsing with shared
sessions.
+ Bugfix: Fixed parsing of <svg> foreignObject content nested
within a <p>, which could incorrectly move the HTML subtree
outside the SVG.
+ Change: Deprecated internal helper
org.jsoup.internal.Functions (for removal in v1.23.1). This
was previously used to support older Android API levels
without full java.util.function coverage; jsoup now requires
core library desugaring so this indirection is no longer
necessary.
* Changes of 1.21.2
+ Change: Deprecated internal (yet visible) methods
Normalizer#normalize(String, bool) and
Attribute#shouldCollapseAttribute(Document.OutputSettings).
These will be removed in a future version.
+ Change: Deprecated
Connection#sslSocketFactory(SSLSocketFactory) in favor of the
new Connection#sslContext(SSLContext). Using sslSocketFactory
will force the use of the legacy HttpUrlConnection
implementation, which does not support HTTP/2.
+ Improvement: When pretty-printing, if there are consecutive
text nodes (via DOM manipulation), the non-significant
whitespace between them will be collapsed.
+ Improvement: Updated Connection.Response#statusMessage() to
return a simple loggable string message (e.g. "OK") when using
the HttpClient implementation, which doesn't otherwise return
any server-set status message.
+ Improvement: Attributes#size() and Attributes#isEmpty() now
exclude any internal attributes (such as user data) from their
count. This aligns with the attributes' serialized output and
iterator.
+ Improvement: Added Connection#sslContext(SSLContext) to
provide a custom SSL (TLS) context to requests, supporting
both the HttpClient and the legacy HttUrlConnection
implementations.
+ Improvement: Performance optimizations for DOM manipulation
methods including when repeatedly removing an element's first
child (element.child(0).remove()), and when using
Parser#parseBodyFragement() to parse a large number of direct
children.
+ Bugfix: When parsing from an InputStream and a multibyte
character happened to straddle a buffer boundary, the stream
would not be completely read.
+ Bugfix: In NodeTraversor, if a last child element was removed
during the head() call, the parent would be visited twice.
+ Bugfix: Cloning an Element that has an Attributes object would
add an empty internal user-data attribute to that clone, which
would cause unexpected results for Attributes#size() and
Attributes#isEmpty().
+ Bugfix: In a multithreaded application where multiple threads
are calling Element#children() on the same element
concurrently, a race condition could happen when the method
was generating the internal child element cache (a filtered
view of its child nodes). Since concurrent reads of DOM
objects should be threadsafe without external synchronization,
this method has been updated to execute atomically.
+ Bugfix: When parsing HTML with svg:script elements in SVG
elements, don't enter the Text insertion mode, but continue to
parse as foreign content. Otherwise, misnested HTML could then
cause an IndexOutOfBoundsException.
+ Bugfix: Malformed HTML could throw an
IndexOutOfBoundsException during the adoption agency.
* Changes of 1.21.1
+ Change: Removed previously deprecated methods.
+ Change: Deprecated the :matchText pseduo-selector due to its
side effects on the DOM; use the new ::textnode selector and
the Element#selectNodes(String css, Class<T> type) method
instead.
+ Change: Deprecated Connection.Response#bufferUp() in lieu of
Connection.Response#readFully() which can throw a checked
IOException.
+ Change: Deprecated internal methods
Validate#ensureNotNull(Object) (replaced by typed
Validate#expectNotNull(T)); protected HTML appenders from
Attribute and Node.
+ Change: If you happen to be using any of the deprecated
methods, please take the opportunity now to migrate away from
them, as they will be removed in a future release.
+ Improvement: Enhanced the Selector to support direct matching
against nodes such as comments and text nodes. For example,
you can now find an element that follows a specific comment:
::comment:contains(prices) + p will select p elements
immediately after a <!-- prices: --> comment. Supported types
include ::node, ::leafnode, ::comment, ::text, ::data, and
::cdata. Node contextual selectors like ::node:contains(text),
:matches(regex), and :blank are also supported. Introduced
Element#selectNodes(String css) and Element#selectNodes(String
css, Class<T> nodeType) for direct node selection.
+ Improvement: Added TagSet#onNewTag(Consumer<Tag> customizer):
register a callback that's invoked for each new or cloned Tag
when it's inserted into the set. Enables dynamic tweaks of tag
options (for example, marking all custom tags as self-closing,
or everything in a given namespace as preserving whitespace).
+ Improvement: Made TokenQueue and CharacterReader
autocloseable, to ensure that they will release their buffers
back to the buffer pool, for later reuse.
+ Improvement: Added Selector#evaluatorOf(String css), as a
clearer way to obtain an Evaluator from a CSS query. An alias
of QueryParser.parse(String css).
+ Improvement: Custom tags (defined via the TagSet) in a foreign
namespace (e.g. SVG) can be configured to parse as data tags.
+ Improvement: Added NodeVisitor#traverse(Node) to simplify node
traversal calls (vs. importing NodeTraversor).
+ Improvement: Updated the default user-agent string to improve
compatibility.
+ Improvement: The HTML parser now allows the specific text-data
type (Data, RcData) to be customized for known tags.
(Previously, that was only supported on custom tags.)
+ Improvement: Added Connection.Response#readFully() as a
replacement for Connection.Response#bufferUp() with an
explicit IOException. Similarly, added
Connection.Response#readBody() over
Connection.Response#body(). Deprecated
Connection.Response#bufferUp().
+ Improvement: When serializing HTML, the < and > characters are
now escaped in attributes. This helps prevent a class of
mutation XSS attacks.
+ Improvement: Changed Connection to prefer using the JDK's
HttpClient over HttpUrlConnection, if available, to enable
HTTP/2 support by default. Users can disable via
-Djsoup.useHttpClient=false.
+ Bugfix: The contents of a script in a svg foreign context
should be parsed as script data, not text.
+ Bugfix: Tag#isFormSubmittable() was updating the Tag's
options.
+ Bugfix: The HTML pretty-printer would incorrectly trim
whitespace when text followed an inline element in a block
element.
+ Bugfix: Custom tags with hyphens or other non-letter
characters in their names now work correctly as Data or RcData
tags. Their closing tags are now tokenized properly.
+ Bugfix: When cloning an Element, the clone would retain the
source's cached child Element list (if any), which could lead
to incorrect results when modifying the clone's child
elements.
* Changes of 1.20.1
+ Change: To better follow the HTML5 spec and current browsers,
the HTML parser no longer allows self-closing tags (<foo />)
to close HTML elements by default. Foreign content (SVG,
MathML), and content parsed with the XML parser, still
supports self-closing tags. If you need specific HTML tags to
support self-closing, you can register a custom tag via the
TagSet configured in Parser.tagSet(), using
Tag#set(Tag.SelfClose). Standard void tags (such as <img>,
<br>, etc.) continue to behave as usual and are not affected
by this change.
+ Change: The following internal components have been
deprecated. If you do happen to be using any of these, please
take the opportunity now to migrate away from them, as they
will be removed in jsoup 1.21.1.
- ChangeNotifyingArrayList,
Document.updateMetaCharsetElement(),
Document.updateMetaCharsetElement(boolean),
HtmlTreeBuilder.isContentForTagData(String),
Parser.isContentForTagData(String),
Parser.setTreeBuilder(TreeBuilder), Tag.formatAsBlock(),
Tag.isFormListed(), TokenQueue.addFirst(String),
TokenQueue.chompTo(String),
TokenQueue.chompToIgnoreCase(String),
TokenQueue.consumeToIgnoreCase(String),
TokenQueue.consumeWord(), TokenQueue.matchesAny(String...)
+ Improvement: Rebuilt the HTML pretty-printer, to simplify and
consolidate the implementation, improve consistency, support
custom Tags, and provide a cleaner path for ongoing
improvements. The specific HTML produced by the pretty-printer
may be different from previous versions.
+ Improvement: Added the ability to define custom tags, and to
modify properties of known tags, via the TagSet tag
collection. Their properties can impact both the parse and how
content is serialized (output as HTML or XML).
+ Improvement: Element.cssSelector() will prefer to return
shorter selectors by using ancestor IDs when available and
unique. E.g. #id > div > p instead of html > body > div > div
> p.
+ Improvement: Added Elements.deselect(int index),
Elements.deselect(Object o), and Elements.deselectAll()
methods to remove elements from the Elements list without
removing them from the underlying DOM. Also added
Elements.asList() method to get a modifiable list of elements
without affecting the DOM. (Individual Elements remain linked
to the DOM.)
+ Improvement: Added support for sending a request body from an
InputStream with Connection.requestBodyStream(InputStream
stream).
+ Improvement: The XML parser now supports scoped xmlns: prefix
namespace declarations, and applies the correct namespace to
Tags and Attributes. Also, added Tag#prefix(),
Tag#localName(), Attribute#prefix(), Attribute#localName(),
and Attribute#namespace() to retrieve these.
+ Improvement: CSS identifiers are now escaped and unescaped
correctly to the CSS spec. Element#cssSelector() will emit
appropriately escaped selectors, and the QueryParser supports
those. Added Selector.escapeCssIdentifier() and `
Selector.unescapeCssIdentifier().
+ Improvement: Refactored the CSS QueryParser into a clearer
recursive descent parser.
+ Improvement: CSS selectors with consecutive combinators (e.g.
div >> p) will throw an explicit parse exception.
+ Performance: reduced the shallow size of an Element from 40 to
32 bytes, and the NodeList from 32 to 24.
+ Performance: reduced GC load of new StringBuilders when
tokenizing input HTML.
+ Improvement: Made Parser instances threadsafe, so that
inadvertent use of the same instance across threads will not
lead to errors. For actual concurrency, use
Parser#newInstance() per thread.
+ Bugfix: Element names containing characters invalid in XML are
now normalized to valid XML names when serializing.
+ Bugfix: When serializing to XML, characters that are invalid
in XML 1.0 should be removed (not encoded).
+ Bugfix: When converting a Document to the W3C DOM in W3CDom,
elements with an attribute in an undeclared namespace now get
a declaration of xmlns:prefix="undefined". This allows
subsequent serialization to XML via W3CDom.asString() to
succeed.
+ Bugfix: The StreamParser could emit the final elements of a
document twice, due to how onNodeCompleted was fired when
closing out the stack.
+ Bugfix: When parsing with the XML parser and error tracking
enabled, the trailing ? in <?xml version="1.0"?> would
incorrectly emit an error.
+ Bugfix: Calling Element#cssSelector() on an element with
combining characters in the class or ID now produces the
correct output.
* Changes of 1.19.1
+ Change: Added support for http/2 requests in Jsoup.connect(),
when running on Java 11+, via the Java HttpClient
implementation.
- In this version of jsoup, the default is to make requests
via the HttpUrlConnection implementation: use
System.setProperty("jsoup.useHttpClient", "true"); to enable
making requests via the HttpClient (if available), which
will enable http/2 support. This will become the default in
a later version of jsoup, so now is a good time to validate
it.
- If you are repackaging the jsoup jar in your deployment
(i.e. creating a shaded- or a fat-jar), make sure to specify
that as a Multi-Release JAR.
- If the HttpClient impl is not available in your JRE,
requests will continue to be made via HttpURLConnection (in
http/1.1 mode).
+ Change: Updated the minimum Android API Level validation from
10 to 21. As with previous jsoup versions, Android developers
need to enable core library desugaring. The minimum Java
version remains Java 8.
+ Change: Removed previously deprecated class:
org.jsoup.UncheckedIOException (replace with
java.io.UncheckedIOException); moved previously deprecated
method Element Element#forEach(Consumer) to void
Element#forEach(Consumer()).
+ Change: Deprecated the methods
Document#updateMetaCharsetElement(boolean) and
Document#updateMetaCharsetElement(), as the setting had no
effect. When Document#charset(Charset) is called, the
document's meta charset or XML encoding instruction is always
set.
+ Improvement: When cleaning HTML with a Safelist that preserves
relative links, the isValid() method will now consider these
links valid. Additionally, the enforced attribute rel=nofollow
will only be added to external links when configured in the
safelist.
+ Improvement: Added Element#selectStream(String query) and
Element#selectStream(Evaluator) methods, that return a Stream
of matching elements. Elements are evaluated and returned as
they are found, and the stream can be terminated early.
+ Improvement: Element objects now implement Iterable, enabling
them to be used in enhanced for loops.
+ Improvement: Added support for fragment parsing from a Reader
via Parser#parseFragmentInput(Reader, Element, String).
+ Improvement: Reintroduced CLI executable examples, in
jsoup-examples.jar.
+ Improvement: Optimized performance of selectors like #id
.class (and other similar descendant queries) by around 4.6x,
by better balancing the Ancestor evaluator's cost function in
the query planner.
+ Improvement: Removed the legacy parsing rules for <isindex>
tags, which would autovivify a form element with labels. This
is no longer in the spec.
+ Improvement: Added Elements.selectFirst(String cssQuery) and
Elements.expectFirst(String cssQuery), to select the first
matching element from an Elements list.
+ Improvement: When parsing with the XML parser, XML
Declarations and Processing Instructions are directly handled,
vs bouncing through the HTML parser's bogus comment handler.
Serialization for non-doctype declarations no longer end with
a spurious !.
+ Improvement: When converting parsed HTML to XML or the W3C
DOM, element names containing < are normalized to _ to ensure
valid XML. For example, <foo<bar> becomes <foo_bar>, as XML
does not allow < in element names, but HTML5 does.
+ Improvement: Reimplemented the HTML5 Adoption Agency Algorithm
to the current spec. This handles mis-nested formating /
structural elements.
+ Bugfix: If an element has an ; in an attribute name, it could
not be converted to a W3C DOM element, and so subsequent XPath
queries could miss that element. Now, the attribute name is
more completely normalized.
+ Bugfix: For backwards compatibility, reverted the internal
attribute key for doctype names to "name".
+ Bugfix: In Connection, skip cookies that have no name, rather
than throwing a validation exception.
+ Bugfix: When running on JDK 1.8, the error
java.lang.NoSuchMethodError:
java.nio.ByteBuffer.flip()Ljava/nio/ByteBuffer; could be
thrown when calling Response#body() after parsing from a URL
and the buffer size was exceeded.
+ Bugfix: For backwards compatibility, allow null InputStream
inputs to Jsoup.parse(InputStream stream, ...), by returning
an empty Document.
+ Bugfix: A template tag containing an li within an open li
would be parsed incorrectly, as it was not recognized as a
"special" tag (which have additional processing rules). Also,
added the SVG and MathML namespace tags to the list of special
tags.
+ Bugfix: A template tag containing a button within an open
button would be parsed incorrectly, as the "in button scope"
check was not aware of the template element. Corrected other
instances including MathML and SVG elements, also.
+ Bugfix: An :nth-child selector with a negative digit-less
step, such as :nth-child(-n+2), would be parsed incorrectly as
a positive step, and so would not match as expected.
+ Bugfix: Calling doc.charset(charset) on an empty XML document
would throw an IndexOutOfBoundsException.
+ Bugfix: Fixed a memory leak when reusing a nested
StructuralEvaluator (e.g., a selector ancestor chain like A B
C) by ensuring cache reset calls cascade to inner members.
+ Bugfix: Concurrent calls to doc.clone().append(html) were not
supported. When a document was cloned, its Parser was not
cloned but was a shallow copy of the original parser.
* Changes of 1.18.3
+ Bugfix: When serializing to XML, attribute names containing -,
., or digits were incorrectly marked as invalid and removed.
* Changes of 1.18.2
+ Improvement: Optimized the throughput and memory use
throughout the input read and parse flows, with heap
allocations and GC down between -6% and -89%, and throughput
improved up to +143% for small inputs. Most inputs sizes will
see throughput increases of ~ 20%. These performance
improvements come through recycling the backing byte[] and
char[] arrays used to read and parse the input.
+ Improvement: Speed optimized html() and Entities.escape() when
the input contains UTF characters in a supplementary plane, by
around 49%.
+ Improvement: The form associated elements returned by
FormElement.elements() now reflect changes made to the DOM,
subsequently to the original parse.
+ Improvement: In the TreeBuilder, the onNodeInserted() and
onNodeClosed() events are now also fired for the outermost /
root Document node. This enables source position tracking on
the Document node (which was previously unset). And it also
enables the node traversor to see the outer Document node.
+ Improvement: Selected Elements can now be position swapped
inline using Elements#set().
+ Bugfix: Element.cssSelector() would fail if the element's
class contained a * character.
+ Bugfix: When tracking source ranges, a text node following an
invalid self-closing element may be left untracked.
+ Bugfix: When a document has no doctype, or a doctype not named
html, it should be parsed in Quirks Mode.
+ Bugfix: With a selector like div:has(span + a), the has()
component was not working correctly, as the inner combining
query caused the evaluator to match those against the outer's
siblings, not children.
+ Bugfix: A selector query that included multiple :has()
components in a nested :has() might incorrectly execute.
+ Bugfix: When cookie names in a response are duplicated, the
simple view of cookies available via
Connection.Response#cookies() will provide the last one set.
Generally it is better to use the Jsoup.newSession method to
maintain a cookie jar, as that applies appropriate path
selection on cookies when making requests.
+ Bugfix: When parsing named HTML entities, base entities should
resolve if they are a prefix of the input token (and not in an
attribute).
+ Bugfix: Fixed incorrect tracking of source ranges for
attributes merged from late-occurring elements that were
implicitly created (html or body).
+ Bugfix: Follow the current HTML specification in the tokenizer
to allow < as part of a tag name, instead of emitting it as a
character node.
+ Bugfix: Similarly, allow a < as the start of an attribute
name, vs creating a new element. The previous behavior was
intended to parse closer to what we anticipated the author's
intent to be, but that does not align to the spec or to how
browsers behave.
* Changes of 1.18.1
+ Improvement: Stream Parser: A StreamParser provides a
progressive parse of its input. For URL requests, available
via Connection.Response.streamParser(). As each Element is
completed, it is emitted via a Stream or Iterator interface.
Elements returned will be complete with all their children,
and an (empty) next sibling, if applicable. Elements (or their
children) may be removed from the DOM during the parse, for
e.g. to conserve memory, providing a mechanism to parse an
input document that would otherwise be too large to fit into
memory, yet still providing a DOM interface to the document
and its elements. Additionally, the parser provides a
selectFirst(String query) / selectNext(String query), which
will run the parser until a hit is found, at which point the
parse is suspended. It can be resumed via another select()
call, or via the stream() or iterator() methods.
+ Improvement: Download Progress: added a Response Progress
event interface, which reports progress and URLs are
downloaded (and parsed). Set via
Connection.onResponseProgress(). Supported on both a session
and a single connection level.
+ Improvement: Added Path accepting parse methods:
Jsoup.parse(Path), Jsoup.parse(path, charsetName, baseUri,
parser), etc.
+ Improvement: Updated the button tag configuration to include a
space between multiple button elements in the Element.text()
method.
+ Improvement: Added support for the ns|* all elements in
namespace Selector.
+ Improvement: When normalising attribute names during
serialization, invalid characters are now replaced with _, vs
being stripped. This should make the process clearer, and
generally prevent an invalid attribute name being coerced
unexpectedly.
+ Change: Removed previously deprecated internal classes and
methods.
+ Build change: the built jar's OSGi manifest no longer imports
itself.
+ Bugfix: When tracking source positions, if the first node was
a TextNode, its position was incorrectly set to -1.
+ Bugfix: When connecting (or redirecting) to URLs with
characters such as {, } in the path, a Malformed URL exception
would be thrown (if in development), or the URL might
otherwise not be escaped correctly (if in production). The URL
encoding process has been improved to handle these characters
correctly.
+ Bugfix: When using W3CDom with a custom output Document, a
Null Pointer Exception would be thrown.
+ Bugfix: The :has() selector did not match correctly when using
sibling combinators (like e.g.: h1:has(+h2)).
+ Bugfix: The :empty selector incorrectly matched elements that
started with a blank text node and were followed by non-empty
nodes, due to an incorrect short-circuit.
+ Bugfix: Element.cssSelector() would fail with "Did not find
balanced marker" when building a selector for elements that
had a ( or [ in their class names. And selectors with those
characters escaped would not match as expected.
+ Bugfix: Updated Entities.escape(string) to make the escaped
text suitable for both text nodes and attributes (previously
was only for text nodes). This does not impact the output of
Element.html() which correctly applies a minimal escape
depending on if the use will be for text data or in a quoted
attribute.
+ Fuzz: a Stack Overflow exception could occur when resolving a
crafted <base href> URL, in the normalizing regex.
* Changes of 1.17.2
+ Improvement: Attribute object accessors: Added
Element.attribute(String) and Attributes.attribute(String) to
more simply obtain an Attribute object.
+ Improvement: Attribute source tracking: If source tracking is
on, and an Attribute's key is changed (via
Attribute.setKey(String)), the source range is now still
tracked in Attribute.sourceRange().
+ Improvement: Wildcard attribute selector: Added support for
the [*] element with any attribute selector. And also restored
support for selecting by an empty attribute name prefix ([^]).
+ Bugfix: Mixed-cased source position: When tracking the source
position of attributes, if the source attribute name was
mix-cased but the parser was lower-case normalizing attribute
names, the source position for that attribute was not tracked
+ Bugfix: Source position NPE: When tracking the source position
of a body fragment parse, a null pointer exception was thrown.
+ Bugfix: Multi-point emoji entity: A multi-point encoded emoji
entity may be incorrectly decoded to the replacement character
+ Bugfix: Selector sub-expressions: (Regression) in a selector
like parent [attr=va], other, the , OR was binding to
[attr=va] instead of parent [attr=va], causing incorrect
selections. The fix includes a EvaluatorDebug class that
generates a sexpr to represent the query, allowing simpler and
more thorough query parse tests.
+ Bugfix: XML CData output: When generating XML-syntax output
from parsed HTML, script nodes containing (pseudo) CData
sections would have an extraneous CData section added, causing
script execution errors. Now, the data content is emitted in a
HTML/XML/XHTML polyglot format, if the data is not already
within a CData section.
+ Bugfix: Thread safety: The :has evaluator held a
non-thread-safe Iterator, and so if an Evaluator object was
shared across multiple concurrent threads, a NoSuchElement
exception may be thrown, and the selected results may be
incorrect. Now, the iterator object is a thread-local.
* Changes of 1.17.1
+ Improvement: in Jsoup.connect(), added support for
request-level authentication, supporting authentication to
proxies and to servers.
+ Improvement: in the Elements list, added direct support for
`#set(index, element)`, `#remove(index)`, `#remove(object)`,
`#clear()`, `#removeAll(collection)`,
`#retainAll(collection)`, `#removeIf(filter)`,
`#replaceAll(operator)`. These methods update the original
DOM, as well as the Elements list.
+ Improvement: added the NodeIterator class, to efficiently
traverse a node tree using the Iterator interface. And
added Stream Element#stream() and Node#nodeStream() methods,
to enable fluent composable stream pipelines of node
traversals.
+ Improvement: when changing the OutputSettings syntax to XML,
the xhtml EscapeMode is automatically set by default.
+ Improvement: added the `:is(selector list)` pseudo-selector,
which finds elements that match any of the selectors in the
selector list. Useful for making large ORed selectors more
readable.
+ Improvement: repackaged the library with native (vs automatic)
JPMS module support.
+ Improvement: better fidelity of source positions when tracking
is enabled. And implicitly created or closed elements are
tracked and detectable via Range.isImplicit().
+ Improvement: when source tracking is enabled, the source
position for attribute names and values is now available.
Attribute#sourceRange() provides the ranges.
+ Improvement: when running concurrently under Java 21+ Virtual
Threads, virtual threads could be pinned to their carrier
platform thread when parsing an input stream. To improve
performance, particularly when parsing fetched URLs, the
internal ConstrainableInputStream has been replaced by
ControllableInputStream, which avoids the locking which caused
that pinning.
+ Improvement: in Jsoup.Connect, allow any XML mimetype as a
supported mimetype. Was previously limited to
`{application|text}/xml`. This enables for e.g. fetching SVGs
with a image/svg+xml mimetype, without having to disable
mimetype validation.
+ Bugfix: when outputting with XML syntax, HTML elements that
were parsed as data nodes (<script> and <style>) should be
emitted as CDATA nodes, so that they can be parsed correctly
by an XML parser.
+ Bugfix: the Immediate Parent selector `>` could match elements
above the root context element, causing incorrect elements to
be returned when used on elements other than the root document
+ Bugfix: in a sub-query such as `p:has(> span, > i)`,
combinators following the `,` Or combinator would be
incorrectly skipped, such that the sub-query was parsed as `i`
instead of `> i`.
+ Bugfix: in W3CDom, if the jsoup input document contained an
empty doctype, the conversion would fail with a DOMException.
Now, said doctype is discarded, and the conversion continues.
+ Bugfix: when cleaning a document containing SVG elements (or
other foreign elements that have preserved case names),
the cleaned output would be incorrectly nested if the safelist
had a different case than the input document.
+ Bugfix: when cleaning a document, the output style of unknown
self-closing tags from the input was not preserved in the
output. (So a <foo /> in the input, if safe-listed, would be
output as <foo></foo>.)
+ Build Improvement: added a local test proxy implementation,
for proxy integration tests.
+ Build Improvement: added tests for HTTPS request support,
using a local self-signed cert. Includes proxy tests.
+ Change: the InputStream returned in Connection.Response
.bodyStream() is no longer a ConstrainedInputStream, and so is
not subject to settings such as timeout or maximum size. It is
now a plain BufferedInputStream around the response stream.
Whilst this behaviour was not documented, you may have been
inadvertently relying on those constraints. The constraints
are still applied to other methods such as .parse() and
.bufferUp(). So if you do want a constrained
BufferedInputStream, you may do Connection.Response.bufferUp()
.bodyStream().
* Changes of 1.16.2
+ Improvement: optimized the performance of complex CSS
selectors, by adding a cost-based query planner. Evaluators
are sorted by their relative execution cost, and executed in
order of lower to higher cost. This speeds the matching
process by ensuring that simpler evaluations (such as a tag
name match) are conducted prior to more complex evaluations
(such as an attribute regex, or a deep child scan with a
:has).
+ Improvement: added support for <svg> and <math> tags (and
their children). This includes tag namespaces and case
preservation on applicable tags and attributes.
+ Improvement: when converting jsoup Documents to W3C Documents
in W3CDom, HTML documents will be placed in the
`http://www.w3.org/1999/xhtml` namespace by default, per the
HTML5 spec. This can be controlled by setting
`W3CDom#namespaceAware(false)`.
+ Improvement: speed optimized the Structural Evaluators by
memoizing previous evaluations. Particularly the `~` (any
preceding sibling) and `:nth-of-type` selectors are improved.
+ Improvement: tweaked the performance of the Element
nextElementSibling, previousElementSibling,
firstElementSibling, lastElementSibling, firstElementChild,
and lastElementChild. They now inplace filter/skip in the
child-node list, vs having to allocate and scan a complete
Element filtered list.
+ Improvement: optimized internal methods that previously called
Element.children() to use filter/skip child-node list
accessors instead, reducing new Element List allocations.
+ Improvement: tweaked the performance of parsing :pseudo
selectors.
+ Improvement: when using the `:empty` pseudo-selector, blank
textnodes are now considered empty. Previously, an element
containing any whitespace was not considered empty.
+ Improvement: in forms, <input type="image"> should be excluded
from formData() (and hence from form submissions).
+ Improvement: in Safelist, made isSafeTag and isSafeAttribute
public methods, for extensibility.
+ Bugfix: `form` elements and empty elements (such as `img`) did
not have their attributes de-duplicated.
+ Bugfix: if Document.OutputSettings was cloned from a clone, an
NPE would be thrown when used.
+ Bugfix: in Jsoup.connect(url), URL paths containing a %2B were
incorrectly recoded to a '+', or a '+' was recoded to a ' '.
Fixed by reverting to the previous behavior of not encoding
supplied paths, other than normalizing to ASCII.
+ Bugfix: in Jsoup.connect(url), strings containing supplemental
characters (e.g. emoji) were not URL escaped correctly.
+ Bugfix: in Jsoup.connect(url), the ConstrainableInputStream
would clear Thread interrupts when reading the body. This
precluded callers from spawning a thread, running a number of
requests for a length of time, then joining that thread after
interrupting it.
+ Bugfix: when tracking HTML source positions, the closing tags
for H1...H6 elements were not tracked correctly.
+ Bugfix: in Jsoup.connect(), a DELETE method request did not
support a request body.
+ Bugfix: when calling Element.cssSelector() on an extremely
deeply nested element, a StackOverflowError could occur.
Further, a StackOverflowError may occur when running the query
+ Bugfix: appending a node back to its original Element after
empty() would throw an Index out of bounds exception. Also,
now the child nodes that were removed have their parent node
cleared, fully detaching them from the original parent.
+ Bugfix: in Jsoup.Connection when adding headers, the value may
have been assumed to be an incorrectly decoded ISO_8859_1
string, and re-encoded as UTF-8. The value is now left as-is.
+ Change: removed previously deprecated methods
Document#normalise, Element#forEach(org.jsoup.helper
.Consumer<>), Node#forEach(org.jsoup.helper.Consumer<>), and
the org.jsoup.helper.Consumer interface; the latter being a
previously required compatibility shim prior to Android's
de-sugaring support.
+ Change: the previous compatibility shim
org.jsoup.UncheckedIOException is deprecated in favor of the
now supported java.io.UncheckedIOException. If you are
catching the former, modify your code to catch the latter
+ Change: blocked noscript tags from being added to Safelists,
due to incompatibilities between parsers with and without
script-mode enabled.
* Changes of 1.16.1
+ Improvement: in Jsoup.connect(url), natively support URLs with
Unicode characters in the path or query string, without having
to be escaped by the caller.
+ Improvement: Calling Node.remove() on a node with no parent is
now a no-op, vs a validation error.
+ Bugfix: aligned the HTML Tree Builder processing steps for
AfterBody and AfterAfterBody to the updated WHATWG standard,
to not pop the stack to close <body> or <html> elements. This
prevents an errant </html> closing preceding structure. Also
added appropriate error message outputs in this case.
+ Bugfix: Corrected support for ruby elements (<ruby>, <rp>,
<rt>, and <rtc>) to current spec.
+ Bugfix: When using Node.before(node) or Node.after(node), if
the incoming node was a sibling of the context node, the
incoming node may be inserted into the wrong relative location
+ Bugfix: In Jsoup.connect(url), if the input URL had components
that were already % escaped, they would be escaped again,
causing errors when fetched.
+ Bugfix: when tracking input source positions, text in tables
that was fostered had invalid positions.
+ Bugfix: If the Document.OutputSettings class was initialized,
and then Entities.escape(String) called, an NPE may be thrown
due to a class loading circular dependency.
+ Bugfix: when pretty-printing, the first inline Element or
Comment in a block would not be wrap-indented if it were
preceded by a blank text node.
+ Bugfix: when pretty-printing a <pre> containing block tags,
those tags were incorrectly indented.
+ Bugfix: when pretty-printing nested inlineable blocks (such as
a <p> in a <td>), the inner element should be indented.
+ Bugfix: <br> tags should be wrap-indented when in block tags
(and not when in inline tags).
+ Bugfix: the contents of a sufficiently large <textarea> with
un-escaped HTML closing tags may be incorrectly parsed to an
empty node.
* Changes of 1.15.4
+ Improvement: added the ability to escape CSS selectors (tags,
IDs, classes) to match elements that don't follow regular CSS
syntax. For example, to match by classname
<p class="one.two">, use document.select("p.one\\.two");
+ Improvement: when pretty-printing, wrap text that follows a
<br> tag.
+ Improvement: when pretty-printing, normalize newlines that
follow self-closing tags in custom tags.
+ Improvement: when pretty-printing, collapse non-significant
whitespace between a block and an inline tag.
+ Improvement: in Element#forEach and Node#forEachNode, use
java.util.function.Consumer instead of the previous Android
compatibility shim org.jsoup.helper.Consumer. Subsequently,
the latter has been deprecated.
+ Improvement: added a new method Document#forms(), to
conveniently retrieve a List<FormElement> containing the
<form> elements in a document.
+ Improvement: added a new method Document#expectForm(query),
to find the first matching FormElement, or blow up trying.
+ Bugfix: URLs containing characters such as [ and ] were not
escaped correctly, and would throw a MalformedURLException
when fetched.
+ Bugfix: Element.cssSelector would create invalid selectors for
elements where the tag name, ID, or classnames needed to be
escaped (e.g. if a class name contained a ':' or '.').
+ Bugfix: element.text() should have a space between a block and
an inline element.
+ Bugfix: if a Node or an Element was replaced with itself, that
node would incorrectly be orphaned.
+ Bugfix: form data on a previous request was copied to a new
request in newRequest(), resulting in an accumulation of form
data when executing multi-step form submissions, or data sent
to later requests incorrectly. Now, newRequest() only copies
session related settings (cookies, proxy settings, user-agent,
etc) but not the request data nor the body.
+ Bugfix: fixed an issue in Safelist.removeAttributes which
could throw a ConcurrentModificationException when using the
":all" pseudo-attribute.
+ Bugfix: given extremely deeply nested HTML, a number of
methods in Element could throw a StackOverflowError due to
excessive recursion. Namely: #data(), #hasText(), #parents(),
and #wrap(html).
+ Change: deprecated the unused Document#normalise() method.
Normalization occurs during the HTML tree construction, and no
longer as a distinct phase.
Patchnames: openSUSE-Leap-16.0-1728
Terms of use: CSAF 2.0 data is provided by SUSE under the Creative Commons License 4.0 with Attribution (CC-BY-4.0).
7.5 (High)
Affected products
Recommended
4 products
| Product | Identifier | Version | Remediation |
|---|---|---|---|
| Unresolved product id: openSUSE Leap 16.0:jsoup-0:1.23.2-160000.1.1.noarch | — |
Vendor Fix
|
|
| Unresolved product id: openSUSE Leap 16.0:jsoup-javadoc-0:1.23.2-160000.1.1.noarch | — |
Vendor Fix
|
|
| Unresolved product id: openSUSE Leap 16.0:re2j-0:1.8-160000.1.1.noarch | — |
Vendor Fix
|
|
| Unresolved product id: openSUSE Leap 16.0:re2j-javadoc-0:1.8-160000.1.1.noarch | — |
Vendor Fix
|
Threats
Impact
important
References
6 references
{
"document": {
"aggregate_severity": {
"namespace": "https://www.suse.com/support/security/rating/",
"text": "important"
},
"category": "csaf_security_advisory",
"csaf_version": "2.0",
"distribution": {
"text": "Copyright 2024 SUSE LLC. All rights reserved.",
"tlp": {
"label": "WHITE",
"url": "https://www.first.org/tlp/"
}
},
"lang": "en",
"notes": [
{
"category": "summary",
"text": "Security update for jsoup, re2j",
"title": "Title of the patch"
},
{
"category": "description",
"text": "This update for jsoup, re2j fixes the following issue:\n\n- CVE-2026-75140: The builder copies the entire inherited namespace map on every start element, causing quadratic time\n and memory complexity, which attackers can exploit to trigger an OutOfMemoryError and terminate the application\n (bsc#1275912).\n\nChanges for jsoup:\n\n- Upgrade to upstream version 1.23.2\n + Improvement: Improved consecutive StreamParser.selectFirst()\n calls during progressive parsing, so later matches are\n returned with their parsed contents when earlier selections\n had left them as parser lookahead. E.g., given\n \u003ctitle\u003eOne\u003c/title\u003e\u003cp id=hit\u003eFull\u003c/p\u003e\u003cp\u003eNext\u003c/p\u003e, selecting\n title and then #hit now advances the partial lookahead and\n returns \u003cp id=\"hit\"\u003eFull\u003c/p\u003e, rather than returning an empty\n \u003cp id=\"hit\"\u003e\u003c/p\u003e before its content is parsed. The updated\n readiness tracking follows StreamParser\u0027s normal emission\n order across implicit HTML structure and parser recovery.\n + Improvement: Improved XML parser performance and memory use\n for documents with many nested namespace declarations by\n recording namespace changes within each element scope\n (bsc#1275912, CVE-2026-75140).\n + Improvement: Improved W3CDom conversion performance for\n documents with many nested namespace declarations. The W3C\n converter now uses the same optimized namespace tracking as\n the XML parser.\n + Improvement: Improved W3CDom XML conversion to retain\n processing instructions, comments outside the root element,\n and CDATA sections, which were previously dropped or converted\n to text.\n + Improvement: DOM mutation methods, including child insertion\n and replacement, now reject operations that would create a\n cycle, such as making a node its own child or moving an\n ancestor beneath a descendant.\n + Improvement: Added Elements#before(Node), after(Node),\n prepend(Node), and append(Node) to match the existing HTML\n string methods.\n + Improvement: Large file-backed uploads through\n Connection.requestBodyStream(InputStream) now stream directly\n with the JDK HttpClient on Java 11+, rather than being loaded\n fully into memory first.\n + Improvement: Extended Java 11+ HTTP client reuse from requests\n sharing a Jsoup.newSession() to ordinary Jsoup.connect()\n calls, reducing transport thread and connection setup churn\n under sustained request loads. Sessions with custom\n authentication or SSL contexts continue to use their own\n client.\n + Change: Aligned the XML parser stack depth and lookups to the\n configured maximum, which now defaults to 512 for both HTML\n and XML. Use Parser#setMaxDepth(int) to configure.\n + Bugfix: Fixed W3CDom namespace conversion in several cases:\n - Namespace declarations and prefixed attributes now carry the\n correct namespace URI, so namespace-aware DOM lookups work\n as expected.\n - Attributes added after parsing, or included through subtree\n conversion, now use inherited prefix declarations.\n - Namespace declarations now apply regardless of attribute\n order, and an empty declaration shadows an inherited binding\n only within its scope.\n - With namespace awareness disabled, inherited and undeclared\n prefixes now receive the declarations needed for XML\n serialization.\n - Valid HTML names that are not XML QNames, such as a:b:c, are\n normalized. Attributes that still cannot be represented are\n skipped, and unrepresentable elements no longer change the\n surrounding tree.\n + Bugfix: Fixed W3CDom conversion of programmatically created or\n renamed elements whose names can be represented in a jsoup\n HTML DOM but are not valid XML names, such as 1abc. These\n names are now normalized (e.g. _1abc) instead of causing a\n NullPointerException.\n + Bugfix: Fixed XML doctype serialization when a system\n identifier contains a double quote, which could otherwise\n produce invalid XML.\n + Bugfix: XML serialization now repairs element and attribute\n names that start with an invalid character, rather than\n outputting null elements or dropping attributes. For example,\n an attribute named 1a is written as _1a. Additional leading\n underscores keep repaired attribute names unique if they\n conflict with another attribute.\n + Bugfix: Supplementary Unicode characters are now escaped\n correctly when serializing with non-UTF, non-ASCII output\n charsets such as ISO-8859-1. Previously, characters could be\n emitted unescaped when their low 16-bit value was\n representable by the configured charset, causing replacement\n or corruption when the output was encoded.\n + Bugfix: Fixed the JDK HttpClient implementation to accept\n responses missing a Content-Type header, matching the\n HttpURLConnection implementation.\n + Bugfix: Fixed HTTP response content-type matching to handle\n media types case-insensitively and recognize structured +xml\n suffixes, including vendor-specific media types.\n + Bugfix: HTTP request URL normalization now percent-encodes\n ASCII control characters, DEL, and embedded fragment\n delimiters, keeping normalized URLs valid for HTTP requests\n while preserving existing escapes.\n + Bugfix: Corrected multipart form encoding to percent-escape CR\n and LF in field names and filenames, matching the HTML form\n submission specification. Multipart file content-types\n containing CR or LF are now rejected with a\n ValidationException.\n + Bugfix: Aligned trailing comment placement with the HTML\n specification: comments after \u003c/body\u003e remain children of the\n html element, while comments after \u003c/html\u003e remain children of\n the document.\n + Bugfix: When using the optional re2j regular expression\n engine, memory allocation errors caused by complex selector\n patterns at match time are now normalized to a\n ValidationException with a Pattern complexity error message.\n + Bugfix: Fixed parsing of malformed SVG and MathML content so\n that breakout HTML tags are placed according to the HTML\n specification.\n + Bugfix: Fixed deeply nested malformed HTML parsing that could\n lose the document body because stack lookups did not align to\n the configured maximum parser depth.\n + Bugfix: Aligned RCDATA, RAWTEXT, and script-data parsing with\n the HTML specification: malformed end tags no longer consume\n following markup, unclosed title/textarea content stays text\n through EOF, and custom text tags match exact names.\n + Bugfix: Improved URL validation during HTTP/HTTPS URL\n resolution and cleaning; resolved URLs without a host are now\n rejected instead of being accepted based only on their scheme\n prefix, aligning to RFC 9110. Valid relative links and\n non-HTTP(S) schemes are unchanged.\n + Bugfix: Redirects with malformed single-slash HTTP locations\n now use standard URL resolution to align with browsers.\n + Bugfix: Template fragment parsing now handles unmatched\n \u003c/template\u003e tags without throwing a ValidationException.\n + Bugfix: Improved source tracking for adopted formatting\n elements and malformed markup ending at EOF.\n * Changes of 1.23.1\n + Improvement: Reduced retained memory when parsing with source\n position tracking enabled (Parser#setTrackPosition(true)).\n Source ranges are now stored in compact parser-owned span\n records instead of node and attribute user data, and Position\n objects are created lazily when source ranges are read. This\n cuts tracked DOM retained size by about 50-60% on\n representative benchmark documents, while keeping\n Node#sourceRange(), Element#endSourceRange(), and\n Attribute#sourceRange() behavior intact.\n + Improvement: Added Element#classList(), an immutable snapshot\n of an element\u0027s class names in attribute order. Use hasClass()\n when you just need to test for one class, classList() when you\n want to read or iterate classes without needing a mutable\n result, and classNames() when you want the existing mutable,\n deduplicated set that can be written back with\n classNames(Set). The class APIs now share an HTML-whitespace\n scanner, which also makes classNames() faster and lighter on\n allocation, especially when walking many elements without\n class names.\n + Improvement: Aligned HTML parser scope classification with the\n current HTML spec for select, foreignObject, and template.\n + Improvement: Simplified the HTML tree builder\u0027s scope,\n implied-end-tag, and special-element checks by caching\n parser-only options on Tag. That improves HTML parser\n throughput by about 10% on small inputs and up to about 30% on\n larger inputs in the benchmark fixtures.\n + Improvement: Improved HTML parser throughput stability by\n making hot tokeniser scan paths compile more predictably.\n + Improvement: \u003cnoscript\u003e fallback markup is now parsed into an\n inspectable DOM subtree in both the document head and body.\n The fallback acts as a contained parsing island, so malformed\n markup cannot disrupt the surrounding document structure,\n while normal HTML tokenization still applies within it. This\n also improves round-trip serialization.\n + Improvement: Improved redirect credential handling as a\n defense-in-depth measure: explicit authorization headers and\n request cookies are no longer forwarded across origins,\n reducing exposure through open redirects and aligning with\n HTTP guidance. Cookies managed by a CookieStore continue to\n follow their configured scope.\n + Improvement: Elements can now append their outer HTML,\n including their own tags, directly to an Appendable with\n Node#outerHtml(Appendable), without first creating a String.\n This complements Element#html(Appendable), which appends inner\n HTML only.\n + Improvement: Aligned CDATA tokenization with the HTML spec:\n CDATA syntax in HTML content is parsed as a bogus comment,\n while it remains supported in SVG, MathML, and XML. Also\n improved namespace-aware fragment parsing so SVG and MathML\n contexts, HTML integration points, and context-sensitive\n tokenizer states are handled correctly.\n + Improvement: When using the optional re2j regular expression\n engine, stack overflows caused by complex selector patterns\n are now normalized to a ValidationException with a Pattern\n complexity error message.\n + Bugfix: Fixed HTML parsing of mixed-case RCDATA end tags after\n tag-shaped text. For example, \u003ctitle\u003e\u003cp\u003eFoo\u003c/TiTLE\u003e and\n \u003ctextarea\u003e\u003cimg src=x\u003e\u003c/TeXtArEa\u003e now keep the tag-shaped\n content as text instead of promoting it to markup.\n + Bugfix: Fixed W3CDom XML conversion so plain XML elements\n don\u0027t serialize with the reserved XML namespace as the default\n namespace. Explicit XML namespaces and xml:* attributes are\n still preserved.\n + Bugfix: Preserve control characters in parsed tag names.\n + Bugfix: Updated HTTP redirects to follow the specification:\n 307 and 308 preserve the request method and content, 301 and\n 302 only change POST to GET, and Location is followed only for\n 301, 302, 303, 307, and 308 responses. Streamed request bodies\n are not buffered; if an automatic redirect requires replaying\n one, execution fails, so the caller can resend with a fresh\n stream.\n + Bugfix: Corrected the Cleaner\u0027s same-site link detection to\n compare hostnames rather than URL prefixes when applying\n rel=nofollow.\n + Build change: Cleaned up the Maven build for the multi-release\n JAR so Java 8 and Java 11+ sources compile as separate source\n sets. This avoids spurious Java 8 compiler warnings from\n newer-language overlay sources, keeps long-running parser\n checks behind an explicit profile, and preserves the same\n published artifacts and runtime behavior.\n + Build change: Improved parallelism and tuned timing in our\n integration tests, so that a full mvn clean verify drops from\n ~ 1m18s to ~ 21 seconds.\n * Changes of 1.22.2\n + Improvement: Expanded and clarified NodeTraversor support for\n in-place DOM rewrites during NodeVisitor.head(). Current-node\n edits such as remove, replace, and unwrap now recover more\n predictably, while traversal stays within the original root\n subtree. This makes single-pass tree cleanup and normalization\n visitors easier to write, for example when unwrapping\n presentational elements or replacing text nodes as you walk\n the DOM.\n + Documentation: clarified that a configured Cleaner may be\n reused across concurrent threads, and that shared Safelist\n instances should not be mutated while in use.\n + Improvement: Updated the default HTML TagSet for current HTML\n elements: added dialog, search, picture, and slot; made ins,\n del, button, audio, video, and canvas inline by default\n (Tag#isInline(), aligned to phrasing content in the spec); and\n added readable Element.text() boundaries for controls and\n embedded objects via the new Tag.TextBoundary option. This\n improves pretty-printing and keeps normalized text from\n running adjacent words together.\n + Android (R8/ProGuard): added a rule to ignore the optional\n re2j dependency when not present.\n + Bugfix: Fixed a NodeTraversor regression in 1.21.2 where\n removing or replacing the current node during head() could\n revisit the replacement node and loop indefinitely. The\n traversal docs now also clarify which inserted nodes are\n visited in the current pass.\n + Bugfix: Parsing during charset sniffing no longer fails if an\n advisory available() call throws IOException, as seen on JDK 8\n HttpURLConnection.\n + Bugfix: Cleaner no longer makes relative URL attributes in the\n input document absolute when cleaning or validating a\n Document. URL normalization now applies only to the cleaned\n output, and Safelist.isSafeAttribute() is side effect free.\n + Bugfix: Cleaner no longer duplicates enforced attributes when\n the input Document preserves attribute case. A case-variant\n source attribute is now replaced by the enforced attribute in\n the cleaned output.\n + Bugfix: If a per-request SOCKS proxy is configured, jsoup now\n avoids using the JDK HttpClient, because the JDK would\n silently ignore that proxy and attempt to connect directly.\n Those requests now fall back to the legacy HttpURLConnection\n transport instead, which does support SOCKS.\n + Bugfix: Connection.Response.streamParser() and\n DataUtil.streamParser(Path, ...) could fail on small inputs\n without a declared charset, if the initial 5 KB charset sniff\n fully consumed the input and closed it before the stream parse\n began.\n + Bugfix: In XML mode, doctypes with an internal subset, such as\n \u003c!DOCTYPE root [\u003c!ENTITY name \"value\"\u003e]\u003e, now round-trip\n correctly. The subset is preserved as raw text only; entities\n are not expanded and external DTDs are not loaded.\n + Build change: Migrated the integration test server from Jetty\n to Netty, which actively maintains support for our minimum JDK\n target (8).\n * Changes of 1.22.1\n + Improvement: Added support for using the re2j regular\n expression engine for regex-based CSS selectors (e.g.\n [attr~=regex], :matches(regex)), which ensures linear-time\n performance for regex evaluation. This allows safer handling\n of arbitrary user-supplied query regexes. To enable, add the\n com.google.re2j dependency to your classpath, e.g.:\n \u003cdependency\u003e\n \u003cgroupId\u003ecom.google.re2j\u003c/groupId\u003e\n \u003cartifactId\u003ere2j\u003c/artifactId\u003e\n \u003cversion\u003e1.8\u003c/version\u003e\n \u003c/dependency\u003e\n (If you already have that dependency in your classpath, but\n you want to keep using the Java regex engine, you can disable\n re2j via System.setProperty(\"jsoup.useRe2j\", \"false\").) You\n can confirm that the re2j engine has been enabled correctly by\n calling Regex.usingRe2j().\n + Improvement: Added an instance method Parser#unescape(String,\n boolean) that unescapes HTML entities using the parser\u0027s\n configuration (e.g. to support error tracking), complementing\n the existing static utility Parser.unescapeEntities(String,\n boolean).\n + Improvement: Added a configurable maximum parser depth (to\n limit the number of open elements on stack) to both HTML and\n XML parsers. The HTML parser now defaults to a depth of 512 to\n match browser behavior, and protect against unbounded stack\n growth, while the XML parser keeps unlimited depth by default,\n but can opt into a limit via Parser.setMaxDepth().\n + Build: added CI coverage for JDK 25.\n + Build: added a CI fuzzer for contextual fragment parsing (in\n addition to existing full body HTML and XML fuzzers).\n + Change: Set a removal schedule of jsoup 1.24.1 for previously\n deprecated APIs.\n + Bugfix: Previously cached child Elements of an Element were\n not correctly invalidated in Node#replaceWith(Node), which\n could lead to incorrect results when subsequently calling\n Element#children().\n + Bugfix: Attribute selector values are now compared literally\n without trimming. Previously, jsoup trimmed whitespace from\n selector values and from element attribute values, which could\n cause mismatches with browser behavior (e.g. [attr=\" foo \"]).\n Now matches align with the CSS specification and browser\n engines.\n + Bugfix: When using the JDK HttpClient, any system default\n proxy (ProxySelector.getDefault()) was ignored. Now, the\n system proxy is used if a per-request proxy is not set.\n + Bugfix: A ValidationException could be thrown in the adoption\n agency algorithm with particularly broken input. Now logged as\n a parse error.\n + Bugfix: Null characters in the HTML body were not consistently\n removed; and in foreign content were not correctly replaced.\n + Bugfix: An IndexOutOfBoundsException could be thrown when\n parsing a body fragment with crafted input. Now logged as a\n parse error.\n + Bugfix: When using StructuralEvaluators (e.g., a parent child\n selector) across many retained threads, their memoized results\n could also be retained, increasing memory use. These results\n are now cleared immediately after use, reducing overall memory\n consumption.\n + Bugfix: Cloning a Parser now preserves any custom TagSet\n applied to the parser.\n + Bugfix: Custom tags marked as Tag.Void now parse and serialize\n like the built-in void elements: they no longer consume\n following content, and the XML serializer emits the expected\n self-closing form.\n + Bugfix: The \u003cbr\u003e element is once again classified as an inline\n tag (Tag.isBlock() == false), matching common developer\n expectations and its role as phrasing content in HTML, while\n pretty-printing and text extraction continue to treat it as a\n line break in the rendered output.\n + Bugfix: Fixed an intermittent truncation issue when fetching\n and parsing remote documents via Jsoup.connect(url).get(). On\n responses without a charset header, the initial charset sniff\n could sometimes (depending on buffering / available()\n behavior) be mistaken for end-of-stream and a partial parse\n reused, dropping trailing content.\n + Bugfix: TagSet copies no longer mutate their template during\n lazy lookups, preventing cross-thread\n ConcurrentModificationException when parsing with shared\n sessions.\n + Bugfix: Fixed parsing of \u003csvg\u003e foreignObject content nested\n within a \u003cp\u003e, which could incorrectly move the HTML subtree\n outside the SVG.\n + Change: Deprecated internal helper\n org.jsoup.internal.Functions (for removal in v1.23.1). This\n was previously used to support older Android API levels\n without full java.util.function coverage; jsoup now requires\n core library desugaring so this indirection is no longer\n necessary.\n * Changes of 1.21.2\n + Change: Deprecated internal (yet visible) methods\n Normalizer#normalize(String, bool) and\n Attribute#shouldCollapseAttribute(Document.OutputSettings).\n These will be removed in a future version.\n + Change: Deprecated\n Connection#sslSocketFactory(SSLSocketFactory) in favor of the\n new Connection#sslContext(SSLContext). Using sslSocketFactory\n will force the use of the legacy HttpUrlConnection\n implementation, which does not support HTTP/2.\n + Improvement: When pretty-printing, if there are consecutive\n text nodes (via DOM manipulation), the non-significant\n whitespace between them will be collapsed.\n + Improvement: Updated Connection.Response#statusMessage() to\n return a simple loggable string message (e.g. \"OK\") when using\n the HttpClient implementation, which doesn\u0027t otherwise return\n any server-set status message.\n + Improvement: Attributes#size() and Attributes#isEmpty() now\n exclude any internal attributes (such as user data) from their\n count. This aligns with the attributes\u0027 serialized output and\n iterator.\n + Improvement: Added Connection#sslContext(SSLContext) to\n provide a custom SSL (TLS) context to requests, supporting\n both the HttpClient and the legacy HttUrlConnection\n implementations.\n + Improvement: Performance optimizations for DOM manipulation\n methods including when repeatedly removing an element\u0027s first\n child (element.child(0).remove()), and when using\n Parser#parseBodyFragement() to parse a large number of direct\n children.\n + Bugfix: When parsing from an InputStream and a multibyte\n character happened to straddle a buffer boundary, the stream\n would not be completely read.\n + Bugfix: In NodeTraversor, if a last child element was removed\n during the head() call, the parent would be visited twice.\n + Bugfix: Cloning an Element that has an Attributes object would\n add an empty internal user-data attribute to that clone, which\n would cause unexpected results for Attributes#size() and\n Attributes#isEmpty().\n + Bugfix: In a multithreaded application where multiple threads\n are calling Element#children() on the same element\n concurrently, a race condition could happen when the method\n was generating the internal child element cache (a filtered\n view of its child nodes). Since concurrent reads of DOM\n objects should be threadsafe without external synchronization,\n this method has been updated to execute atomically.\n + Bugfix: When parsing HTML with svg:script elements in SVG\n elements, don\u0027t enter the Text insertion mode, but continue to\n parse as foreign content. Otherwise, misnested HTML could then\n cause an IndexOutOfBoundsException.\n + Bugfix: Malformed HTML could throw an\n IndexOutOfBoundsException during the adoption agency.\n * Changes of 1.21.1\n + Change: Removed previously deprecated methods.\n + Change: Deprecated the :matchText pseduo-selector due to its\n side effects on the DOM; use the new ::textnode selector and\n the Element#selectNodes(String css, Class\u003cT\u003e type) method\n instead.\n + Change: Deprecated Connection.Response#bufferUp() in lieu of\n Connection.Response#readFully() which can throw a checked\n IOException.\n + Change: Deprecated internal methods\n Validate#ensureNotNull(Object) (replaced by typed\n Validate#expectNotNull(T)); protected HTML appenders from\n Attribute and Node.\n + Change: If you happen to be using any of the deprecated\n methods, please take the opportunity now to migrate away from\n them, as they will be removed in a future release.\n + Improvement: Enhanced the Selector to support direct matching\n against nodes such as comments and text nodes. For example,\n you can now find an element that follows a specific comment:\n ::comment:contains(prices) + p will select p elements\n immediately after a \u003c!-- prices: --\u003e comment. Supported types\n include ::node, ::leafnode, ::comment, ::text, ::data, and\n ::cdata. Node contextual selectors like ::node:contains(text),\n :matches(regex), and :blank are also supported. Introduced\n Element#selectNodes(String css) and Element#selectNodes(String\n css, Class\u003cT\u003e nodeType) for direct node selection.\n + Improvement: Added TagSet#onNewTag(Consumer\u003cTag\u003e customizer):\n register a callback that\u0027s invoked for each new or cloned Tag\n when it\u0027s inserted into the set. Enables dynamic tweaks of tag\n options (for example, marking all custom tags as self-closing,\n or everything in a given namespace as preserving whitespace).\n + Improvement: Made TokenQueue and CharacterReader\n autocloseable, to ensure that they will release their buffers\n back to the buffer pool, for later reuse.\n + Improvement: Added Selector#evaluatorOf(String css), as a\n clearer way to obtain an Evaluator from a CSS query. An alias\n of QueryParser.parse(String css).\n + Improvement: Custom tags (defined via the TagSet) in a foreign\n namespace (e.g. SVG) can be configured to parse as data tags.\n + Improvement: Added NodeVisitor#traverse(Node) to simplify node\n traversal calls (vs. importing NodeTraversor).\n + Improvement: Updated the default user-agent string to improve\n compatibility.\n + Improvement: The HTML parser now allows the specific text-data\n type (Data, RcData) to be customized for known tags.\n (Previously, that was only supported on custom tags.)\n + Improvement: Added Connection.Response#readFully() as a\n replacement for Connection.Response#bufferUp() with an\n explicit IOException. Similarly, added\n Connection.Response#readBody() over\n Connection.Response#body(). Deprecated\n Connection.Response#bufferUp().\n + Improvement: When serializing HTML, the \u003c and \u003e characters are\n now escaped in attributes. This helps prevent a class of\n mutation XSS attacks.\n + Improvement: Changed Connection to prefer using the JDK\u0027s\n HttpClient over HttpUrlConnection, if available, to enable\n HTTP/2 support by default. Users can disable via\n -Djsoup.useHttpClient=false.\n + Bugfix: The contents of a script in a svg foreign context\n should be parsed as script data, not text.\n + Bugfix: Tag#isFormSubmittable() was updating the Tag\u0027s\n options.\n + Bugfix: The HTML pretty-printer would incorrectly trim\n whitespace when text followed an inline element in a block\n element.\n + Bugfix: Custom tags with hyphens or other non-letter\n characters in their names now work correctly as Data or RcData\n tags. Their closing tags are now tokenized properly.\n + Bugfix: When cloning an Element, the clone would retain the\n source\u0027s cached child Element list (if any), which could lead\n to incorrect results when modifying the clone\u0027s child\n elements.\n * Changes of 1.20.1\n + Change: To better follow the HTML5 spec and current browsers,\n the HTML parser no longer allows self-closing tags (\u003cfoo /\u003e)\n to close HTML elements by default. Foreign content (SVG,\n MathML), and content parsed with the XML parser, still\n supports self-closing tags. If you need specific HTML tags to\n support self-closing, you can register a custom tag via the\n TagSet configured in Parser.tagSet(), using\n Tag#set(Tag.SelfClose). Standard void tags (such as \u003cimg\u003e,\n \u003cbr\u003e, etc.) continue to behave as usual and are not affected\n by this change.\n + Change: The following internal components have been\n deprecated. If you do happen to be using any of these, please\n take the opportunity now to migrate away from them, as they\n will be removed in jsoup 1.21.1.\n - ChangeNotifyingArrayList,\n Document.updateMetaCharsetElement(),\n Document.updateMetaCharsetElement(boolean),\n HtmlTreeBuilder.isContentForTagData(String),\n Parser.isContentForTagData(String),\n Parser.setTreeBuilder(TreeBuilder), Tag.formatAsBlock(),\n Tag.isFormListed(), TokenQueue.addFirst(String),\n TokenQueue.chompTo(String),\n TokenQueue.chompToIgnoreCase(String),\n TokenQueue.consumeToIgnoreCase(String),\n TokenQueue.consumeWord(), TokenQueue.matchesAny(String...)\n + Improvement: Rebuilt the HTML pretty-printer, to simplify and\n consolidate the implementation, improve consistency, support\n custom Tags, and provide a cleaner path for ongoing\n improvements. The specific HTML produced by the pretty-printer\n may be different from previous versions.\n + Improvement: Added the ability to define custom tags, and to\n modify properties of known tags, via the TagSet tag\n collection. Their properties can impact both the parse and how\n content is serialized (output as HTML or XML).\n + Improvement: Element.cssSelector() will prefer to return\n shorter selectors by using ancestor IDs when available and\n unique. E.g. #id \u003e div \u003e p instead of html \u003e body \u003e div \u003e div\n \u003e p.\n + Improvement: Added Elements.deselect(int index),\n Elements.deselect(Object o), and Elements.deselectAll()\n methods to remove elements from the Elements list without\n removing them from the underlying DOM. Also added\n Elements.asList() method to get a modifiable list of elements\n without affecting the DOM. (Individual Elements remain linked\n to the DOM.)\n + Improvement: Added support for sending a request body from an\n InputStream with Connection.requestBodyStream(InputStream\n stream).\n + Improvement: The XML parser now supports scoped xmlns: prefix\n namespace declarations, and applies the correct namespace to\n Tags and Attributes. Also, added Tag#prefix(),\n Tag#localName(), Attribute#prefix(), Attribute#localName(),\n and Attribute#namespace() to retrieve these.\n + Improvement: CSS identifiers are now escaped and unescaped\n correctly to the CSS spec. Element#cssSelector() will emit\n appropriately escaped selectors, and the QueryParser supports\n those. Added Selector.escapeCssIdentifier() and `\n Selector.unescapeCssIdentifier().\n + Improvement: Refactored the CSS QueryParser into a clearer\n recursive descent parser.\n + Improvement: CSS selectors with consecutive combinators (e.g.\n div \u003e\u003e p) will throw an explicit parse exception.\n + Performance: reduced the shallow size of an Element from 40 to\n 32 bytes, and the NodeList from 32 to 24.\n + Performance: reduced GC load of new StringBuilders when\n tokenizing input HTML.\n + Improvement: Made Parser instances threadsafe, so that\n inadvertent use of the same instance across threads will not\n lead to errors. For actual concurrency, use\n Parser#newInstance() per thread.\n + Bugfix: Element names containing characters invalid in XML are\n now normalized to valid XML names when serializing.\n + Bugfix: When serializing to XML, characters that are invalid\n in XML 1.0 should be removed (not encoded).\n + Bugfix: When converting a Document to the W3C DOM in W3CDom,\n elements with an attribute in an undeclared namespace now get\n a declaration of xmlns:prefix=\"undefined\". This allows\n subsequent serialization to XML via W3CDom.asString() to\n succeed.\n + Bugfix: The StreamParser could emit the final elements of a\n document twice, due to how onNodeCompleted was fired when\n closing out the stack.\n + Bugfix: When parsing with the XML parser and error tracking\n enabled, the trailing ? in \u003c?xml version=\"1.0\"?\u003e would\n incorrectly emit an error.\n + Bugfix: Calling Element#cssSelector() on an element with\n combining characters in the class or ID now produces the\n correct output.\n * Changes of 1.19.1\n + Change: Added support for http/2 requests in Jsoup.connect(),\n when running on Java 11+, via the Java HttpClient\n implementation.\n - In this version of jsoup, the default is to make requests\n via the HttpUrlConnection implementation: use\n System.setProperty(\"jsoup.useHttpClient\", \"true\"); to enable\n making requests via the HttpClient (if available), which\n will enable http/2 support. This will become the default in\n a later version of jsoup, so now is a good time to validate\n it.\n - If you are repackaging the jsoup jar in your deployment\n (i.e. creating a shaded- or a fat-jar), make sure to specify\n that as a Multi-Release JAR.\n - If the HttpClient impl is not available in your JRE,\n requests will continue to be made via HttpURLConnection (in\n http/1.1 mode).\n + Change: Updated the minimum Android API Level validation from\n 10 to 21. As with previous jsoup versions, Android developers\n need to enable core library desugaring. The minimum Java\n version remains Java 8.\n + Change: Removed previously deprecated class:\n org.jsoup.UncheckedIOException (replace with\n java.io.UncheckedIOException); moved previously deprecated\n method Element Element#forEach(Consumer) to void\n Element#forEach(Consumer()).\n + Change: Deprecated the methods\n Document#updateMetaCharsetElement(boolean) and\n Document#updateMetaCharsetElement(), as the setting had no\n effect. When Document#charset(Charset) is called, the\n document\u0027s meta charset or XML encoding instruction is always\n set.\n + Improvement: When cleaning HTML with a Safelist that preserves\n relative links, the isValid() method will now consider these\n links valid. Additionally, the enforced attribute rel=nofollow\n will only be added to external links when configured in the\n safelist.\n + Improvement: Added Element#selectStream(String query) and\n Element#selectStream(Evaluator) methods, that return a Stream\n of matching elements. Elements are evaluated and returned as\n they are found, and the stream can be terminated early.\n + Improvement: Element objects now implement Iterable, enabling\n them to be used in enhanced for loops.\n + Improvement: Added support for fragment parsing from a Reader\n via Parser#parseFragmentInput(Reader, Element, String).\n + Improvement: Reintroduced CLI executable examples, in\n jsoup-examples.jar.\n + Improvement: Optimized performance of selectors like #id\n .class (and other similar descendant queries) by around 4.6x,\n by better balancing the Ancestor evaluator\u0027s cost function in\n the query planner.\n + Improvement: Removed the legacy parsing rules for \u003cisindex\u003e\n tags, which would autovivify a form element with labels. This\n is no longer in the spec.\n + Improvement: Added Elements.selectFirst(String cssQuery) and\n Elements.expectFirst(String cssQuery), to select the first\n matching element from an Elements list.\n + Improvement: When parsing with the XML parser, XML\n Declarations and Processing Instructions are directly handled,\n vs bouncing through the HTML parser\u0027s bogus comment handler.\n Serialization for non-doctype declarations no longer end with\n a spurious !.\n + Improvement: When converting parsed HTML to XML or the W3C\n DOM, element names containing \u003c are normalized to _ to ensure\n valid XML. For example, \u003cfoo\u003cbar\u003e becomes \u003cfoo_bar\u003e, as XML\n does not allow \u003c in element names, but HTML5 does.\n + Improvement: Reimplemented the HTML5 Adoption Agency Algorithm\n to the current spec. This handles mis-nested formating /\n structural elements.\n + Bugfix: If an element has an ; in an attribute name, it could\n not be converted to a W3C DOM element, and so subsequent XPath\n queries could miss that element. Now, the attribute name is\n more completely normalized.\n + Bugfix: For backwards compatibility, reverted the internal\n attribute key for doctype names to \"name\".\n + Bugfix: In Connection, skip cookies that have no name, rather\n than throwing a validation exception.\n + Bugfix: When running on JDK 1.8, the error\n java.lang.NoSuchMethodError:\n java.nio.ByteBuffer.flip()Ljava/nio/ByteBuffer; could be\n thrown when calling Response#body() after parsing from a URL\n and the buffer size was exceeded.\n + Bugfix: For backwards compatibility, allow null InputStream\n inputs to Jsoup.parse(InputStream stream, ...), by returning\n an empty Document.\n + Bugfix: A template tag containing an li within an open li\n would be parsed incorrectly, as it was not recognized as a\n \"special\" tag (which have additional processing rules). Also,\n added the SVG and MathML namespace tags to the list of special\n tags.\n + Bugfix: A template tag containing a button within an open\n button would be parsed incorrectly, as the \"in button scope\"\n check was not aware of the template element. Corrected other\n instances including MathML and SVG elements, also.\n + Bugfix: An :nth-child selector with a negative digit-less\n step, such as :nth-child(-n+2), would be parsed incorrectly as\n a positive step, and so would not match as expected.\n + Bugfix: Calling doc.charset(charset) on an empty XML document\n would throw an IndexOutOfBoundsException.\n + Bugfix: Fixed a memory leak when reusing a nested\n StructuralEvaluator (e.g., a selector ancestor chain like A B\n C) by ensuring cache reset calls cascade to inner members.\n + Bugfix: Concurrent calls to doc.clone().append(html) were not\n supported. When a document was cloned, its Parser was not\n cloned but was a shallow copy of the original parser.\n * Changes of 1.18.3\n + Bugfix: When serializing to XML, attribute names containing -,\n ., or digits were incorrectly marked as invalid and removed.\n * Changes of 1.18.2\n + Improvement: Optimized the throughput and memory use\n throughout the input read and parse flows, with heap\n allocations and GC down between -6% and -89%, and throughput\n improved up to +143% for small inputs. Most inputs sizes will\n see throughput increases of ~ 20%. These performance\n improvements come through recycling the backing byte[] and\n char[] arrays used to read and parse the input.\n + Improvement: Speed optimized html() and Entities.escape() when\n the input contains UTF characters in a supplementary plane, by\n around 49%.\n + Improvement: The form associated elements returned by\n FormElement.elements() now reflect changes made to the DOM,\n subsequently to the original parse.\n + Improvement: In the TreeBuilder, the onNodeInserted() and\n onNodeClosed() events are now also fired for the outermost /\n root Document node. This enables source position tracking on\n the Document node (which was previously unset). And it also\n enables the node traversor to see the outer Document node.\n + Improvement: Selected Elements can now be position swapped\n inline using Elements#set().\n + Bugfix: Element.cssSelector() would fail if the element\u0027s\n class contained a * character.\n + Bugfix: When tracking source ranges, a text node following an\n invalid self-closing element may be left untracked.\n + Bugfix: When a document has no doctype, or a doctype not named\n html, it should be parsed in Quirks Mode.\n + Bugfix: With a selector like div:has(span + a), the has()\n component was not working correctly, as the inner combining\n query caused the evaluator to match those against the outer\u0027s\n siblings, not children.\n + Bugfix: A selector query that included multiple :has()\n components in a nested :has() might incorrectly execute.\n + Bugfix: When cookie names in a response are duplicated, the\n simple view of cookies available via\n Connection.Response#cookies() will provide the last one set.\n Generally it is better to use the Jsoup.newSession method to\n maintain a cookie jar, as that applies appropriate path\n selection on cookies when making requests.\n + Bugfix: When parsing named HTML entities, base entities should\n resolve if they are a prefix of the input token (and not in an\n attribute).\n + Bugfix: Fixed incorrect tracking of source ranges for\n attributes merged from late-occurring elements that were\n implicitly created (html or body).\n + Bugfix: Follow the current HTML specification in the tokenizer\n to allow \u003c as part of a tag name, instead of emitting it as a\n character node.\n + Bugfix: Similarly, allow a \u003c as the start of an attribute\n name, vs creating a new element. The previous behavior was\n intended to parse closer to what we anticipated the author\u0027s\n intent to be, but that does not align to the spec or to how\n browsers behave.\n * Changes of 1.18.1\n + Improvement: Stream Parser: A StreamParser provides a\n progressive parse of its input. For URL requests, available\n via Connection.Response.streamParser(). As each Element is\n completed, it is emitted via a Stream or Iterator interface.\n Elements returned will be complete with all their children,\n and an (empty) next sibling, if applicable. Elements (or their\n children) may be removed from the DOM during the parse, for\n e.g. to conserve memory, providing a mechanism to parse an\n input document that would otherwise be too large to fit into\n memory, yet still providing a DOM interface to the document\n and its elements. Additionally, the parser provides a\n selectFirst(String query) / selectNext(String query), which\n will run the parser until a hit is found, at which point the\n parse is suspended. It can be resumed via another select()\n call, or via the stream() or iterator() methods.\n + Improvement: Download Progress: added a Response Progress\n event interface, which reports progress and URLs are\n downloaded (and parsed). Set via\n Connection.onResponseProgress(). Supported on both a session\n and a single connection level.\n + Improvement: Added Path accepting parse methods:\n Jsoup.parse(Path), Jsoup.parse(path, charsetName, baseUri,\n parser), etc.\n + Improvement: Updated the button tag configuration to include a\n space between multiple button elements in the Element.text()\n method.\n + Improvement: Added support for the ns|* all elements in\n namespace Selector.\n + Improvement: When normalising attribute names during\n serialization, invalid characters are now replaced with _, vs\n being stripped. This should make the process clearer, and\n generally prevent an invalid attribute name being coerced\n unexpectedly.\n + Change: Removed previously deprecated internal classes and\n methods.\n + Build change: the built jar\u0027s OSGi manifest no longer imports\n itself.\n + Bugfix: When tracking source positions, if the first node was\n a TextNode, its position was incorrectly set to -1.\n + Bugfix: When connecting (or redirecting) to URLs with\n characters such as {, } in the path, a Malformed URL exception\n would be thrown (if in development), or the URL might\n otherwise not be escaped correctly (if in production). The URL\n encoding process has been improved to handle these characters\n correctly.\n + Bugfix: When using W3CDom with a custom output Document, a\n Null Pointer Exception would be thrown.\n + Bugfix: The :has() selector did not match correctly when using\n sibling combinators (like e.g.: h1:has(+h2)).\n + Bugfix: The :empty selector incorrectly matched elements that\n started with a blank text node and were followed by non-empty\n nodes, due to an incorrect short-circuit.\n + Bugfix: Element.cssSelector() would fail with \"Did not find\n balanced marker\" when building a selector for elements that\n had a ( or [ in their class names. And selectors with those\n characters escaped would not match as expected.\n + Bugfix: Updated Entities.escape(string) to make the escaped\n text suitable for both text nodes and attributes (previously\n was only for text nodes). This does not impact the output of\n Element.html() which correctly applies a minimal escape\n depending on if the use will be for text data or in a quoted\n attribute.\n + Fuzz: a Stack Overflow exception could occur when resolving a\n crafted \u003cbase href\u003e URL, in the normalizing regex.\n * Changes of 1.17.2\n + Improvement: Attribute object accessors: Added\n Element.attribute(String) and Attributes.attribute(String) to\n more simply obtain an Attribute object.\n + Improvement: Attribute source tracking: If source tracking is\n on, and an Attribute\u0027s key is changed (via\n Attribute.setKey(String)), the source range is now still\n tracked in Attribute.sourceRange().\n + Improvement: Wildcard attribute selector: Added support for\n the [*] element with any attribute selector. And also restored\n support for selecting by an empty attribute name prefix ([^]).\n + Bugfix: Mixed-cased source position: When tracking the source\n position of attributes, if the source attribute name was\n mix-cased but the parser was lower-case normalizing attribute\n names, the source position for that attribute was not tracked\n + Bugfix: Source position NPE: When tracking the source position\n of a body fragment parse, a null pointer exception was thrown.\n + Bugfix: Multi-point emoji entity: A multi-point encoded emoji\n entity may be incorrectly decoded to the replacement character\n + Bugfix: Selector sub-expressions: (Regression) in a selector\n like parent [attr=va], other, the , OR was binding to\n [attr=va] instead of parent [attr=va], causing incorrect\n selections. The fix includes a EvaluatorDebug class that\n generates a sexpr to represent the query, allowing simpler and\n more thorough query parse tests.\n + Bugfix: XML CData output: When generating XML-syntax output\n from parsed HTML, script nodes containing (pseudo) CData\n sections would have an extraneous CData section added, causing\n script execution errors. Now, the data content is emitted in a\n HTML/XML/XHTML polyglot format, if the data is not already\n within a CData section.\n + Bugfix: Thread safety: The :has evaluator held a\n non-thread-safe Iterator, and so if an Evaluator object was\n shared across multiple concurrent threads, a NoSuchElement\n exception may be thrown, and the selected results may be\n incorrect. Now, the iterator object is a thread-local.\n * Changes of 1.17.1\n + Improvement: in Jsoup.connect(), added support for\n request-level authentication, supporting authentication to\n proxies and to servers.\n + Improvement: in the Elements list, added direct support for\n `#set(index, element)`, `#remove(index)`, `#remove(object)`,\n `#clear()`, `#removeAll(collection)`,\n `#retainAll(collection)`, `#removeIf(filter)`,\n `#replaceAll(operator)`. These methods update the original\n DOM, as well as the Elements list.\n + Improvement: added the NodeIterator class, to efficiently\n traverse a node tree using the Iterator interface. And\n added Stream Element#stream() and Node#nodeStream() methods,\n to enable fluent composable stream pipelines of node\n traversals.\n + Improvement: when changing the OutputSettings syntax to XML,\n the xhtml EscapeMode is automatically set by default.\n + Improvement: added the `:is(selector list)` pseudo-selector,\n which finds elements that match any of the selectors in the\n selector list. Useful for making large ORed selectors more\n readable.\n + Improvement: repackaged the library with native (vs automatic)\n JPMS module support.\n + Improvement: better fidelity of source positions when tracking\n is enabled. And implicitly created or closed elements are\n tracked and detectable via Range.isImplicit().\n + Improvement: when source tracking is enabled, the source\n position for attribute names and values is now available.\n Attribute#sourceRange() provides the ranges.\n + Improvement: when running concurrently under Java 21+ Virtual\n Threads, virtual threads could be pinned to their carrier\n platform thread when parsing an input stream. To improve\n performance, particularly when parsing fetched URLs, the\n internal ConstrainableInputStream has been replaced by\n ControllableInputStream, which avoids the locking which caused\n that pinning.\n + Improvement: in Jsoup.Connect, allow any XML mimetype as a\n supported mimetype. Was previously limited to\n `{application|text}/xml`. This enables for e.g. fetching SVGs\n with a image/svg+xml mimetype, without having to disable\n mimetype validation.\n + Bugfix: when outputting with XML syntax, HTML elements that\n were parsed as data nodes (\u003cscript\u003e and \u003cstyle\u003e) should be\n emitted as CDATA nodes, so that they can be parsed correctly\n by an XML parser.\n + Bugfix: the Immediate Parent selector `\u003e` could match elements\n above the root context element, causing incorrect elements to\n be returned when used on elements other than the root document\n + Bugfix: in a sub-query such as `p:has(\u003e span, \u003e i)`,\n combinators following the `,` Or combinator would be\n incorrectly skipped, such that the sub-query was parsed as `i`\n instead of `\u003e i`.\n + Bugfix: in W3CDom, if the jsoup input document contained an\n empty doctype, the conversion would fail with a DOMException.\n Now, said doctype is discarded, and the conversion continues.\n + Bugfix: when cleaning a document containing SVG elements (or\n other foreign elements that have preserved case names),\n the cleaned output would be incorrectly nested if the safelist\n had a different case than the input document.\n + Bugfix: when cleaning a document, the output style of unknown\n self-closing tags from the input was not preserved in the\n output. (So a \u003cfoo /\u003e in the input, if safe-listed, would be\n output as \u003cfoo\u003e\u003c/foo\u003e.)\n + Build Improvement: added a local test proxy implementation,\n for proxy integration tests.\n + Build Improvement: added tests for HTTPS request support,\n using a local self-signed cert. Includes proxy tests.\n + Change: the InputStream returned in Connection.Response\n .bodyStream() is no longer a ConstrainedInputStream, and so is\n not subject to settings such as timeout or maximum size. It is\n now a plain BufferedInputStream around the response stream.\n Whilst this behaviour was not documented, you may have been\n inadvertently relying on those constraints. The constraints\n are still applied to other methods such as .parse() and\n .bufferUp(). So if you do want a constrained\n BufferedInputStream, you may do Connection.Response.bufferUp()\n .bodyStream().\n * Changes of 1.16.2\n + Improvement: optimized the performance of complex CSS\n selectors, by adding a cost-based query planner. Evaluators\n are sorted by their relative execution cost, and executed in\n order of lower to higher cost. This speeds the matching\n process by ensuring that simpler evaluations (such as a tag\n name match) are conducted prior to more complex evaluations\n (such as an attribute regex, or a deep child scan with a\n :has).\n + Improvement: added support for \u003csvg\u003e and \u003cmath\u003e tags (and\n their children). This includes tag namespaces and case\n preservation on applicable tags and attributes.\n + Improvement: when converting jsoup Documents to W3C Documents\n in W3CDom, HTML documents will be placed in the\n `http://www.w3.org/1999/xhtml` namespace by default, per the\n HTML5 spec. This can be controlled by setting\n `W3CDom#namespaceAware(false)`.\n + Improvement: speed optimized the Structural Evaluators by\n memoizing previous evaluations. Particularly the `~` (any\n preceding sibling) and `:nth-of-type` selectors are improved.\n + Improvement: tweaked the performance of the Element\n nextElementSibling, previousElementSibling,\n firstElementSibling, lastElementSibling, firstElementChild,\n and lastElementChild. They now inplace filter/skip in the\n child-node list, vs having to allocate and scan a complete\n Element filtered list.\n + Improvement: optimized internal methods that previously called\n Element.children() to use filter/skip child-node list\n accessors instead, reducing new Element List allocations.\n + Improvement: tweaked the performance of parsing :pseudo\n selectors.\n + Improvement: when using the `:empty` pseudo-selector, blank\n textnodes are now considered empty. Previously, an element\n containing any whitespace was not considered empty.\n + Improvement: in forms, \u003cinput type=\"image\"\u003e should be excluded\n from formData() (and hence from form submissions).\n + Improvement: in Safelist, made isSafeTag and isSafeAttribute\n public methods, for extensibility.\n + Bugfix: `form` elements and empty elements (such as `img`) did\n not have their attributes de-duplicated.\n + Bugfix: if Document.OutputSettings was cloned from a clone, an\n NPE would be thrown when used.\n + Bugfix: in Jsoup.connect(url), URL paths containing a %2B were\n incorrectly recoded to a \u0027+\u0027, or a \u0027+\u0027 was recoded to a \u0027 \u0027.\n Fixed by reverting to the previous behavior of not encoding\n supplied paths, other than normalizing to ASCII.\n + Bugfix: in Jsoup.connect(url), strings containing supplemental\n characters (e.g. emoji) were not URL escaped correctly.\n + Bugfix: in Jsoup.connect(url), the ConstrainableInputStream\n would clear Thread interrupts when reading the body. This\n precluded callers from spawning a thread, running a number of\n requests for a length of time, then joining that thread after\n interrupting it.\n + Bugfix: when tracking HTML source positions, the closing tags\n for H1...H6 elements were not tracked correctly.\n + Bugfix: in Jsoup.connect(), a DELETE method request did not\n support a request body.\n + Bugfix: when calling Element.cssSelector() on an extremely\n deeply nested element, a StackOverflowError could occur.\n Further, a StackOverflowError may occur when running the query\n + Bugfix: appending a node back to its original Element after\n empty() would throw an Index out of bounds exception. Also,\n now the child nodes that were removed have their parent node\n cleared, fully detaching them from the original parent.\n + Bugfix: in Jsoup.Connection when adding headers, the value may\n have been assumed to be an incorrectly decoded ISO_8859_1\n string, and re-encoded as UTF-8. The value is now left as-is.\n + Change: removed previously deprecated methods\n Document#normalise, Element#forEach(org.jsoup.helper\n .Consumer\u003c\u003e), Node#forEach(org.jsoup.helper.Consumer\u003c\u003e), and\n the org.jsoup.helper.Consumer interface; the latter being a\n previously required compatibility shim prior to Android\u0027s\n de-sugaring support.\n + Change: the previous compatibility shim\n org.jsoup.UncheckedIOException is deprecated in favor of the\n now supported java.io.UncheckedIOException. If you are\n catching the former, modify your code to catch the latter\n + Change: blocked noscript tags from being added to Safelists,\n due to incompatibilities between parsers with and without\n script-mode enabled.\n * Changes of 1.16.1\n + Improvement: in Jsoup.connect(url), natively support URLs with\n Unicode characters in the path or query string, without having\n to be escaped by the caller.\n + Improvement: Calling Node.remove() on a node with no parent is\n now a no-op, vs a validation error.\n + Bugfix: aligned the HTML Tree Builder processing steps for\n AfterBody and AfterAfterBody to the updated WHATWG standard,\n to not pop the stack to close \u003cbody\u003e or \u003chtml\u003e elements. This\n prevents an errant \u003c/html\u003e closing preceding structure. Also\n added appropriate error message outputs in this case.\n + Bugfix: Corrected support for ruby elements (\u003cruby\u003e, \u003crp\u003e,\n \u003crt\u003e, and \u003crtc\u003e) to current spec.\n + Bugfix: When using Node.before(node) or Node.after(node), if\n the incoming node was a sibling of the context node, the\n incoming node may be inserted into the wrong relative location\n + Bugfix: In Jsoup.connect(url), if the input URL had components\n that were already % escaped, they would be escaped again,\n causing errors when fetched.\n + Bugfix: when tracking input source positions, text in tables\n that was fostered had invalid positions.\n + Bugfix: If the Document.OutputSettings class was initialized,\n and then Entities.escape(String) called, an NPE may be thrown\n due to a class loading circular dependency.\n + Bugfix: when pretty-printing, the first inline Element or\n Comment in a block would not be wrap-indented if it were\n preceded by a blank text node.\n + Bugfix: when pretty-printing a \u003cpre\u003e containing block tags,\n those tags were incorrectly indented.\n + Bugfix: when pretty-printing nested inlineable blocks (such as\n a \u003cp\u003e in a \u003ctd\u003e), the inner element should be indented.\n + Bugfix: \u003cbr\u003e tags should be wrap-indented when in block tags\n (and not when in inline tags).\n + Bugfix: the contents of a sufficiently large \u003ctextarea\u003e with\n un-escaped HTML closing tags may be incorrectly parsed to an\n empty node.\n * Changes of 1.15.4\n + Improvement: added the ability to escape CSS selectors (tags,\n IDs, classes) to match elements that don\u0027t follow regular CSS\n syntax. For example, to match by classname\n \u003cp class=\"one.two\"\u003e, use document.select(\"p.one\\\\.two\");\n + Improvement: when pretty-printing, wrap text that follows a\n \u003cbr\u003e tag.\n + Improvement: when pretty-printing, normalize newlines that\n follow self-closing tags in custom tags.\n + Improvement: when pretty-printing, collapse non-significant\n whitespace between a block and an inline tag.\n + Improvement: in Element#forEach and Node#forEachNode, use\n java.util.function.Consumer instead of the previous Android\n compatibility shim org.jsoup.helper.Consumer. Subsequently,\n the latter has been deprecated.\n + Improvement: added a new method Document#forms(), to\n conveniently retrieve a List\u003cFormElement\u003e containing the\n \u003cform\u003e elements in a document.\n + Improvement: added a new method Document#expectForm(query),\n to find the first matching FormElement, or blow up trying.\n + Bugfix: URLs containing characters such as [ and ] were not\n escaped correctly, and would throw a MalformedURLException\n when fetched.\n + Bugfix: Element.cssSelector would create invalid selectors for\n elements where the tag name, ID, or classnames needed to be\n escaped (e.g. if a class name contained a \u0027:\u0027 or \u0027.\u0027).\n + Bugfix: element.text() should have a space between a block and\n an inline element.\n + Bugfix: if a Node or an Element was replaced with itself, that\n node would incorrectly be orphaned.\n + Bugfix: form data on a previous request was copied to a new\n request in newRequest(), resulting in an accumulation of form\n data when executing multi-step form submissions, or data sent\n to later requests incorrectly. Now, newRequest() only copies\n session related settings (cookies, proxy settings, user-agent,\n etc) but not the request data nor the body.\n + Bugfix: fixed an issue in Safelist.removeAttributes which\n could throw a ConcurrentModificationException when using the\n \":all\" pseudo-attribute.\n + Bugfix: given extremely deeply nested HTML, a number of\n methods in Element could throw a StackOverflowError due to\n excessive recursion. Namely: #data(), #hasText(), #parents(),\n and #wrap(html).\n + Change: deprecated the unused Document#normalise() method.\n Normalization occurs during the HTML tree construction, and no\n longer as a distinct phase.\n",
"title": "Description of the patch"
},
{
"category": "details",
"text": "openSUSE-Leap-16.0-1728",
"title": "Patchnames"
},
{
"category": "legal_disclaimer",
"text": "CSAF 2.0 data is provided by SUSE under the Creative Commons License 4.0 with Attribution (CC-BY-4.0).",
"title": "Terms of use"
}
],
"publisher": {
"category": "vendor",
"contact_details": "https://www.suse.com/support/security/contact/",
"name": "SUSE Product Security Team",
"namespace": "https://www.suse.com/"
},
"references": [
{
"category": "external",
"summary": "SUSE ratings",
"url": "https://www.suse.com/support/security/rating/"
},
{
"category": "self",
"summary": "URL of this CSAF notice",
"url": "https://ftp.suse.com/pub/projects/security/csaf/opensuse-su-2026_21900-1.json"
},
{
"category": "self",
"summary": "SUSE Bug 1275912",
"url": "https://bugzilla.suse.com/1275912"
},
{
"category": "self",
"summary": "SUSE CVE CVE-2026-75140 page",
"url": "https://www.suse.com/security/cve/CVE-2026-75140/"
}
],
"title": "Security update for jsoup, re2j",
"tracking": {
"current_release_date": "2026-09-24T17:48:11Z",
"generator": {
"date": "2026-09-21T12:46:08Z",
"engine": {
"name": "cve-database.git:bin/generate-csaf.pl",
"version": "1"
}
},
"id": "openSUSE-SU-2026:21900-1",
"initial_release_date": "2026-09-21T12:46:08Z",
"revision_history": [
{
"date": "2026-09-21T12:46:08Z",
"number": "1",
"summary": "Current version"
},
{
"date": "2026-09-24T17:48:11Z",
"number": "2",
"summary": "unknown changes"
}
],
"status": "final",
"version": "2"
}
},
"product_tree": {
"branches": [
{
"branches": [
{
"branches": [
{
"category": "product_version",
"name": "jsoup-0:1.23.2-160000.1.1.noarch",
"product": {
"name": "jsoup-0:1.23.2-160000.1.1.noarch",
"product_id": "jsoup-0:1.23.2-160000.1.1.noarch",
"product_identification_helper": {
"cpe": "cpe:2.3:a:jsoup:jsoup:1.23.2:*:*:*:*:*:*:*",
"purl": "pkg:rpm/suse/jsoup@1.23.2-160000.1.1?arch=noarch\u0026upstream=jsoup-0:1.23.2-160000.1.1.src.rpm"
}
}
},
{
"category": "product_version",
"name": "jsoup-javadoc-0:1.23.2-160000.1.1.noarch",
"product": {
"name": "jsoup-javadoc-0:1.23.2-160000.1.1.noarch",
"product_id": "jsoup-javadoc-0:1.23.2-160000.1.1.noarch",
"product_identification_helper": {
"cpe": "cpe:2.3:a:jsoup:jsoup:1.23.2:*:*:*:*:*:*:*",
"purl": "pkg:rpm/suse/jsoup-javadoc@1.23.2-160000.1.1?arch=noarch\u0026upstream=jsoup-0:1.23.2-160000.1.1.src.rpm"
}
}
},
{
"category": "product_version",
"name": "re2j-0:1.8-160000.1.1.noarch",
"product": {
"name": "re2j-0:1.8-160000.1.1.noarch",
"product_id": "re2j-0:1.8-160000.1.1.noarch",
"product_identification_helper": {
"purl": "pkg:rpm/suse/re2j@1.8-160000.1.1?arch=noarch"
}
}
},
{
"category": "product_version",
"name": "re2j-javadoc-0:1.8-160000.1.1.noarch",
"product": {
"name": "re2j-javadoc-0:1.8-160000.1.1.noarch",
"product_id": "re2j-javadoc-0:1.8-160000.1.1.noarch",
"product_identification_helper": {
"purl": "pkg:rpm/suse/re2j-javadoc@1.8-160000.1.1?arch=noarch"
}
}
}
],
"category": "architecture",
"name": "noarch"
},
{
"branches": [
{
"category": "product_name",
"name": "openSUSE Leap 16.0",
"product": {
"name": "openSUSE Leap 16.0",
"product_id": "openSUSE Leap 16.0"
}
}
],
"category": "product_family",
"name": "SUSE Linux Enterprise"
}
],
"category": "vendor",
"name": "SUSE"
}
],
"relationships": [
{
"category": "default_component_of",
"full_product_name": {
"name": "jsoup-0:1.23.2-160000.1.1.noarch as component of openSUSE Leap 16.0",
"product_id": "openSUSE Leap 16.0:jsoup-0:1.23.2-160000.1.1.noarch"
},
"product_reference": "jsoup-0:1.23.2-160000.1.1.noarch",
"relates_to_product_reference": "openSUSE Leap 16.0"
},
{
"category": "default_component_of",
"full_product_name": {
"name": "jsoup-javadoc-0:1.23.2-160000.1.1.noarch as component of openSUSE Leap 16.0",
"product_id": "openSUSE Leap 16.0:jsoup-javadoc-0:1.23.2-160000.1.1.noarch"
},
"product_reference": "jsoup-javadoc-0:1.23.2-160000.1.1.noarch",
"relates_to_product_reference": "openSUSE Leap 16.0"
},
{
"category": "default_component_of",
"full_product_name": {
"name": "re2j-0:1.8-160000.1.1.noarch as component of openSUSE Leap 16.0",
"product_id": "openSUSE Leap 16.0:re2j-0:1.8-160000.1.1.noarch"
},
"product_reference": "re2j-0:1.8-160000.1.1.noarch",
"relates_to_product_reference": "openSUSE Leap 16.0"
},
{
"category": "default_component_of",
"full_product_name": {
"name": "re2j-javadoc-0:1.8-160000.1.1.noarch as component of openSUSE Leap 16.0",
"product_id": "openSUSE Leap 16.0:re2j-javadoc-0:1.8-160000.1.1.noarch"
},
"product_reference": "re2j-javadoc-0:1.8-160000.1.1.noarch",
"relates_to_product_reference": "openSUSE Leap 16.0"
}
]
},
"vulnerabilities": [
{
"cve": "CVE-2026-75140",
"ids": [
{
"system_name": "SUSE CVE Page",
"text": "https://www.suse.com/security/cve/CVE-2026-75140"
}
],
"notes": [
{
"category": "general",
"text": "jsoup through 1.23.2, fixed in commit 862ba2f, contains an uncontrolled resource consumption vulnerability in XmlTreeBuilder that allows remote attackers to exhaust JVM heap memory by supplying a deeply nested XML document with uniquely-namespaced elements. The builder copies the entire inherited namespace map on every start element, causing quadratic time and memory complexity, which attackers can exploit to trigger an OutOfMemoryError and terminate the application.",
"title": "CVE description"
}
],
"product_status": {
"recommended": [
"openSUSE Leap 16.0:jsoup-0:1.23.2-160000.1.1.noarch",
"openSUSE Leap 16.0:jsoup-javadoc-0:1.23.2-160000.1.1.noarch",
"openSUSE Leap 16.0:re2j-0:1.8-160000.1.1.noarch",
"openSUSE Leap 16.0:re2j-javadoc-0:1.8-160000.1.1.noarch"
]
},
"references": [
{
"category": "external",
"summary": "CVE-2026-75140",
"url": "https://www.suse.com/security/cve/CVE-2026-75140"
},
{
"category": "external",
"summary": "SUSE Bug 1275912 for CVE-2026-75140",
"url": "https://bugzilla.suse.com/1275912"
}
],
"remediations": [
{
"category": "vendor_fix",
"details": "To install this SUSE Security Update use the SUSE recommended installation methods like YaST online_update or \"zypper patch\".\n",
"product_ids": [
"openSUSE Leap 16.0:jsoup-0:1.23.2-160000.1.1.noarch",
"openSUSE Leap 16.0:jsoup-javadoc-0:1.23.2-160000.1.1.noarch",
"openSUSE Leap 16.0:re2j-0:1.8-160000.1.1.noarch",
"openSUSE Leap 16.0:re2j-javadoc-0:1.8-160000.1.1.noarch"
]
}
],
"scores": [
{
"cvss_v3": {
"baseScore": 7.5,
"baseSeverity": "HIGH",
"vectorString": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H",
"version": "3.1"
},
"products": [
"openSUSE Leap 16.0:jsoup-0:1.23.2-160000.1.1.noarch",
"openSUSE Leap 16.0:jsoup-javadoc-0:1.23.2-160000.1.1.noarch",
"openSUSE Leap 16.0:re2j-0:1.8-160000.1.1.noarch",
"openSUSE Leap 16.0:re2j-javadoc-0:1.8-160000.1.1.noarch"
]
}
],
"threats": [
{
"category": "impact",
"date": "2026-09-21T12:46:08Z",
"details": "important"
}
],
"title": "CVE-2026-75140"
}
]
}
Loading…
Loading…
Experimental. This forecast is provided for visualization only and may change without notice. Do not use it for operational decisions.
Forecast uses a logistic model when the trend is rising, or an exponential decay model when the trend is falling. Fitted via linearized least squares.
Sightings
| Author | Source | Type | Date | Other |
|---|
Nomenclature
- Seen: The vulnerability was mentioned, discussed, or observed by the user.
- Confirmed: The vulnerability has been validated from an analyst's perspective.
- Published Proof of Concept: A public proof of concept is available for this vulnerability.
- Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
- Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
- Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
- Not confirmed: The user expressed doubt about the validity of the vulnerability.
- Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.
Loading…
Loading…
The MITRE ATT&CK techniques below are AI-generated suggestions, inferred from the description of the
vulnerability by the CIRCL/vulnerability-attack-technique-classification-roberta-base
model, served locally by ML-Gateway.
They have not been verified by an analyst and are provided for guidance only.
The approach is described in our paper Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion.
Browse all ATT&CK techniques and the vulnerabilities related to each.
The approach is described in our paper Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion.
Browse all ATT&CK techniques and the vulnerabilities related to each.
Loading…
Related by attack behaviour
Vulnerabilities whose description is nearest to this one in the vector space of the CIRCL/vulnerability-attack-technique-biencoder model. This is a similarity search over the bi-encoder space (plain cosine), not a classification, and it has no measured accuracy.
Loading…