Skip to content

Repository files navigation

json-guard

JSON.parse with hard limits on nesting depth, key count, string length and input size.

  deep nesting — 500,000 levels of nesting, 1.0 MB of valid JSON
    JSON.parse     31ms   rss +  74 MB   parsed — structuredClone() threw RangeError
    json-guard      0ms   rss +   1 MB   refused — Nesting depth 33 exceeded maxDepth 32 at offset 32

That's node examples/bomb-demo.ts. The document is one megabyte, it is valid JSON, and JSON.parse accepts it without complaint.

The problem

JSON.parse has no bounds. Not on depth, not on the number of keys, not on the length of a single string. Every limit your service has on a request body is a limit on bytes, and bytes are a poor proxy for what parsing costs:

1 MB of JSON can be JSON.parse result
500,000 levels of nesting parses fine — and then your code recurses over it
120,000 keys in one object one dictionary-mode map, 120,000 entries deep
one 1 MB string fine, until the same trick arrives at 100 MB

The nesting case is the nasty one, because JSON.parse is not where it fails. V8's JSON parser is iterative, so it happily builds 500,000 levels. The stack overflow lands later, in whatever touches the result — structuredClone, JSON.stringify, a schema validator, a recursive sanitiser, a logger:

const value = JSON.parse(bomb);   // fine
structuredClone(value);           // RangeError: Maximum call stack size exceeded

A RangeError thrown from inside a serialiser two layers below your route handler is not a 400. It is whatever your framework does with an unexpected throw, from a stack that is already exhausted.

secure-json-parse (from the Fastify team) is the closest thing on npm, and it addresses __proto__ and constructor pollution only. Nothing bounds parser blowup.

Why a reviver doesn't fix it

The obvious move is JSON.parse(text, reviver) — the reviver sees every key, so surely it can enforce limits. It can't, for two reasons.

It doesn't know the depth. The reviver is given a key and a value, never a depth, and it runs bottom-up. The most you can reconstruct is each subtree's height, carried in a WeakMap, with the root's height as the answer.

It runs after the structure is already built. JSON.parse completes the entire parse and then walks the result. By the time your first callback fires, the memory has been spent. A reviver can report the blowup; it cannot prevent it.

On the input you actually care about it does not even manage to report it. The reviver walk is itself recursive, so it overflows the stack at around 5,000 levels — inside the thing that was supposed to be the defence:

  enforcing maxDepth 32 on 200,000 levels of nesting (0.4 MB)

    reviver      20ms   rss +  46 MB   RangeError: Maximum call stack size exceeded
    pre-scan      0ms   rss +   1 MB   DepthLimitError: Nesting depth 33 exceeded maxDepth 32 at offset 32

So this library does the other thing: one linear pass over the raw text before JSON.parse is called at all. Counting brackets outside of string literals bounds depth, keys and string length while the input is still just characters, and nothing has been allocated. It refuses at the offending byte and hands back its offset.

That's node examples/strategies.ts, which also measures what the pass costs when nothing is wrong:

  overhead on an ordinary 0.6 MB document (28,002 keys, depth 4), best of 20

    JSON.parse                  1.0ms   1.00x
    JSON.parse + reviver        7.3ms   7.22x
    json-guard parse            2.0ms   1.98x
    json-guard, reviver too     8.0ms   7.92x

Roughly double. That is the honest price: a scan in JavaScript against a parser written in C++, and the guarded path does both. It buys a bound that cannot be reached any other way, and it is still three times cheaper than the reviver that would not have worked.

Install

npm install json-guard

Node >= 20.6. No runtime dependencies.

Use

import { parse } from 'json-guard';

const value = parse(body, {
  maxDepth: 32,
  maxKeys: 10_000,
  maxStringLength: 1 << 20,
  maxBytes: 5 << 20,
  protoAction: 'error',
});

parse takes a string or a Uint8Array and returns what JSON.parse would. It throws a JsonGuardError subclass when a limit is exceeded, and the usual SyntaxError when the document is not valid JSON.

For callers who would rather branch than catch:

import { safeParse, JsonGuardError } from 'json-guard';

const result = safeParse(body, { maxDepth: 32 });
if (!result.ok) {
  if (result.error instanceof JsonGuardError) return res.status(413).send(result.error.message);
  return res.status(400).send('malformed JSON');
}
use(result.value);

Handing it the raw bytes is better than handing it a string. The size check then happens on byteLength, and an oversized payload is refused without ever being decoded:

const buf = await readBody(req);        // Uint8Array
parse(buf, { maxBytes: 1 << 20 });      // refused before a string exists

Picking limits

measure runs the same scan with no limits applied, so you can point it at real traffic and choose numbers from data rather than superstition:

import { measure } from 'json-guard';

measure('{"a":[{"b":"hello"}]}');
// { depth: 3, keys: 2, longestString: 5, hasPrototypeKey: false, bytes: 21 }

Options

Option Default What it bounds
maxDepth 64 Nesting levels. A flat object or array is 1.
maxKeys 100_000 Object members in the whole document, not per object.
maxStringLength 1 MiB Any single string, keys included, in UTF-16 code units.
maxBytes 1 MiB The input, in UTF-8 bytes, checked first.
protoAction 'error' __proto__ keys: 'error', 'remove' or 'ignore'.
reviver — Passed through to JSON.parse.

Any limit set to Infinity is disabled.

maxKeys counts syntactically, so {"a":1,"a":2} is two keys even though the result has one — the parser pays for both, which is the point of the limit.

Errors

All extend JsonGuardError, which carries an offset: the index into the input text where the violation was found. Every limit error also carries limit and observed.

Error Extra fields
DepthLimitError limit, observed, offset
KeyLimitError limit, observed, offset
StringLengthError limit, observed, offset
SizeLimitError limit, observed (no offset — nothing was scanned)
PrototypePollutionError key, path, offset
catch (err) {
  if (err instanceof DepthLimitError) {
    console.warn(`too deep at character ${err.offset} (limit ${err.limit})`);
  }
}

For DepthLimitError and KeyLimitError, observed is always limit + 1: enforcement stops at the first violation rather than measuring how much worse the document gets. StringLengthError.observed is the real length of that string.

Prototype keys

The one thing the scan cannot do by itself is remove a key, so this is where the reviver earns its place — and it is attached only when the scan has already proved the document contains one:

parse('{"a":1,"__proto__":{"x":1}}');                          // PrototypePollutionError
parse('{"a":1,"__proto__":{"x":1}}', { protoAction: 'remove' }); // { a: 1 }

__proto__ anywhere, and prototype inside a constructor object, are treated as dangerous — the same pair secure-json-parse guards. Escaped spellings ("__proto__") are caught too. PrototypePollutionError.path names the offender, e.g. items[].__proto__.

Worth being precise about the risk: JSON.parse itself is not vulnerable. It defines __proto__ as an own property rather than assigning it, so nothing is polluted at parse time. The danger is entirely downstream, in whatever copies the result — Object.assign, a deep merge, a config loader.

What it does not do

  • It is not a schema validator. It bounds shape, not meaning. Pair it with zod, ajv or whatever you already use; this runs first and cheaply, so the validator never sees a document designed to hurt it.
  • It does not bound array length. A million-element array of numbers is bounded only by maxBytes. maxKeys covers object members, where the per-entry cost is much higher; for flat arrays, bytes are a decent proxy.
  • It does not stream. The whole document must be in memory before the scan starts, so maxBytes is the limit that matters when the input arrives off a socket. Cap the body as you read it, too.
  • It does not make JSON.parse faster, and it is not a JSON parser. The scan is deliberately not a validator: it tracks only enough state to know where the string literals are, and leaves every question of grammar to JSON.parse.
  • maxStringLength counts UTF-16 code units, so an emoji is 2. maxBytes counts UTF-8 bytes. They are different units on purpose: one is what the string costs in memory, the other is what it cost on the wire.

A note on the measurements

examples/bomb-demo.ts forks a child per run so the memory figures are not polluted by each other, and forces a GC after building the payload so the delta is the parse and nothing else. The residual rss in the guarded rows is the input string itself being flattened for scanning, not anything built from it — which is why the giant-string row shows 64 MB against JSON.parse's 128 MB. Numbers above are from Node 24.20 on an Apple M5 Pro; run it yourself.

The honest summary: for depth and key explosions the difference is as large as it looks, because the guarded path never allocates. For the giant-string case the win is a factor of two in memory rather than a factor of a hundred — and in practice maxBytes refuses that document first, in microseconds, without scanning at all.

Develop

Tests are TypeScript run directly by Node's test runner — no build, no install:

node --test "test/*.test.ts"    # full suite, needs node 24+ for type stripping
node examples/bomb-demo.ts
node examples/strategies.ts

npm run build && npm run test:dist   # what CI runs against node 20 and 22

License

MIT

About

JSON.parse with hard limits on nesting depth, key count, string length and input size. Refuses parser bombs before they are allocated.

Topics

Resources

Stars

80 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages