Engineering · 24 September 2026

PII split across stream chunks: why per-chunk guardrail checks pass it

A guardrail that inspects each streamed chunk on its own will let a card number through if it arrives in two pieces. Four projects that share no code shipped this independently, a fifth documented it, and the class now has a name and a DOI. Here is the reproduction, the invariant that rules it out, and the test that catches it.

Two stream chunks side by side, with a sensitive value cut exactly at the boundary between them — half in each chunk
The defect in one picture: neither chunk contains the value, so neither check fires, and the reader's browser joins the halves back together.

The reproduction: “user@exa” then “mple.com”

A model streams its answer in deltas that are split on token boundaries, not on meaning. So an email address, a card number or an API key routinely arrives in pieces:

chunk 1:  "...you can reach me at user@exa"
chunk 2:  "mple.com whenever you like."

Run an email pattern over chunk 1: no match. Over chunk 2: no match. Both chunks pass, both are written to the wire, and the browser concatenates them into a working address on screen. The same two-part example appears in the reports filed against several projects, which is the first clue that this is not one team's mistake.

Why a per-chunk check cannot see it

The check is running on a string the model never meant as a unit. Server-sent events carry whatever the tokeniser produced, so the boundary falls in an arbitrary place — inside a word, inside a number, inside a key. A filter that treats each delta as a complete document is asking the wrong question of the right data.

The fix people reach for first is to remember a bit of the previous chunk. That helps with detection and does nothing about the leak, for a reason that takes a moment to see: by the time the second half arrives and the pattern finally matches, the first half has already been sent. Masking it afterwards changes your logs, not the reader's screen.

Text may not be emitted until it is known that no match could still grow into it. Detection is not prevention.

Four projects shipped it independently, and a fifth documented it

These are separate codebases, separate languages, separate teams. None of the reports cites any of the others. Status below is as checked on 24 September 2026:

ProjectReportStatus
LiteLLM (Python)#41611 — a value split across two SSE chunks passes per-chunk checksOpen, filed 17 Sep 2026
Mastra (TypeScript)#23783 — PIIDetector emits PII in the clear when a match is split across chunksFixed in four days (PR #24189)
LangChain (Python)#35011 — streaming bypasses guardrails and middlewareFixed 10 Jun 2026 (PR #37616)
NVIDIA NeMo Guardrails (Python)#2375 — output masking rail unusable in 0.24.0Open, filed 10 Sep 2026
Vercel AI SDK (TypeScript)#21209 — the guardrail example left wrapStream unimplementedDocumentation gap, closed 22 Sep 2026

The last row is a different kind of entry and is worth keeping separate: the SDK shipped no faulty filter. Its middleware page showed a guardrail example with the streaming half left as a comment, so a reader following it wrote their own — which is how four teams ended up inventing the same two wrong designs.

The class has a name: split-boundary leaks

On 22 September 2026, Ninad Phalak published Split-Boundary Leaks in Streaming Guardrails on Zenodo under CC BY 4.0 — doi.org/10.5281/zenodo.22909585. It collects the four instances above, treats the Vercel case as an acknowledged gap rather than a fifth failure, and names the property all four lack.

Two things to be straight about. The report comes from a competing product, and it reviews no library and endorses none, including ours. And the name is his: if you write about this class, cite the DOI rather than renaming it, because a class with one name is easier for the next team to find than the same class with four.

The invariant: output byte-identical to filtering the whole string

For any way of splitting the same input, the concatenated streamed output must be byte-identical to filtering the whole string at once.

That single sentence rules out every instance above, and it is testable in one assertion rather than argued in review. It is the property to ask a library for, and — as the report notes — almost nobody publishes it as a stated contract, which is exactly what would let you tell a correct implementation from a plausible one.

Does your library check each chunk alone, or keep state across chunks?

Three checks that answer it without reading anyone's source:

  • Split a known value at every offset. Feed the same input chunked every possible way and compare the joined output to filtering the whole string. Any difference is the bug.
  • Watch when the first chunk leaves. If output appears before the scanner could possibly have seen the end of a match, the implementation is emitting on hope.
  • Grep the configuration for a size. A tunable buffer or window size means a fixed window, and a fixed window has an edge — see below.

Having a buffer is not the same as holding text back

This is the distinction that catches good engineers. Detection state and release policy are two separate decisions, and only the second one stops a leak. Umbraco's AI package shipped exactly this shape: a sliding window that detected correctly while every chunk was released as it arrived. The fix merged on 22 September 2026 is titled “make streaming post-generate guardrails actually block/redact”, and what changed was the release policy, not the detector.

A fixed hold-back of N characters is a configurable leak

Hold the last 20, 50 or 64 characters, prepend them to the next chunk, scan, emit. It works until a match is longer than N — and then it fails silently, because the output still looks redacted. LiteLLM carries a stream_holdback_chars setting through eight source files and, when checked on 24 September 2026, zero documentation files, so the number that decides whether you leak is a value most operators will never see.

The Vercel AI SDK documentation now states the general form of this, in a note under its middleware example: an incremental implementation must retain every possible incomplete match, and a fixed-size buffer alone is not safe for unbounded variable-length patterns. A PEM private key block or a long JWT is exactly such a pattern.

The fix for this bug has its own bug

Two failure modes that arrive with the repair:

  • Corruption. Redaction changes the length of the string, so slicing the redacted text at offsets computed from the raw input lands mid-placeholder and eats good characters. Mastra's report includes this alongside the leak.
  • False positives from re-joining. NVIDIA NeMo Guardrails #1197 describes a space introduced when tokens are popped from the buffer: with a sub-word tokeniser, “assisting” arrives as “ass” + “isting” and reaches the output rail as “ass isting”, triggering a policy violation on a word that was never there.

Buffer everything, or stream and leak, is a false choice

There is a third design, and it is the only one that keeps the stream while satisfying the invariant: per chunk, compute the furthest position that no pattern could still extend past — the settlement point — emit up to there, and hold the tail. On ordinary prose the held tail is a couple of dozen characters, so output keeps flowing.

It has one decision that must be made explicitly, and it is where the class comes back: when the held tail reaches its ceiling with a variable-length pattern still open, the implementation either releases text that may be the first half of a secret, or refuses. A bounded tail that releases on overflow is the fixed window again, wearing a better name. The safe behaviour is to fail closed.

The test that catches the whole class

Not a new framework — one assertion, run over every split of every fixture:

for (const text of fixtures) {
  const whole = filterWholeString(text);
  for (let size = 1; size <= text.length; size++) {
    let streamed = '';
    const f = createFilter();
    for (let i = 0; i < text.length; i += size) streamed += f.push(text.slice(i, i + size));
    streamed += f.flush();
    assert.equal(streamed, whole);        // chunking invariance
  }
}

Put an unbounded pattern in the fixtures — a PEM block or a long JWT — or the suite will pass on an implementation that releases on overflow. Add a value pressed against non-ASCII text, and a language written without spaces, because a tail bounded by word delimiters grows without limit there.

Where this write-up comes from

We publish llm-stream-guardrails, an MIT library that implements the settlement-point design and states the invariant as its contract, and our maintainer is the one who reported the unimplemented example to the Vercel AI SDK on 20 September 2026; three documentation pull requests referencing that issue, opened by the project's own automation, were merged two days later. That is the interest to declare before you read anything above as neutral.

Everything factual here was checked against the source on 24 September 2026 rather than quoted from the report: issue and pull request states through the GitHub API, the Zenodo record and its PDF, and the middleware documentation on the SDK's main branch. Where a claim could not be verified, it is not on this page.

FAQ

Questions engineers ask about this

Is this a bug in my framework, or does everyone have it?

Four unrelated projects shipped it independently — LiteLLM, Mastra, LangChain and NVIDIA NeMo Guardrails — and two of those reports were still open on 24 September 2026. Assume your stack has it until you have run the chunking-invariance test against it.

Does my guardrail evaluate each chunk independently, or keep state across chunks?

Split a known value at every offset, feed it through, and compare the joined output with filtering the whole string. If any split produces different output, the guardrail is evaluating chunks independently, whatever the documentation says.

What hold-back size should I use — 20, 50 or 64 characters?

None of them is safe on its own. Any fixed N fails on a match of length N+1, and it fails quietly because the output still looks redacted. The held amount has to be derived from the patterns you are matching, and the implementation has to fail closed when it hits its ceiling.

Can I just buffer the whole response and filter it?

Yes, and it is correct. It is also what the Vercel AI SDK documentation recommends as the safe default. The cost is that you stop streaming: time to first token becomes time to full generation, and memory grows with the block. That trade is reasonable for short replies and painful for long ones.

What does the AI SDK wrapStream guardrail example do now?

Since the September 2026 documentation fix it keeps one buffer per text block, appends each delta without emitting, and on text-end redacts the whole buffered string and emits it once. It closes the split-match hole, and the docs note that it delays output and uses memory proportional to the block.

Is the LiteLLM issue fixed?

Not when checked on 24 September 2026. Issue #41611 was open, with five comments, filed on 17 September. Check it yourself before relying on either answer.

Does NVIDIA NeMo Guardrails have this?

Issue #2375, "output masking rail unusable in 0.24.0", was open on 24 September 2026. A separate issue, #1197, documents a related hazard in the same area: re-joining popped tokens introduces a space, so "assisting" can reach the output rail as "ass isting".

How do I write a test that catches it?

One assertion: for every fixture and every chunk size, the concatenation of the streamed output must equal filtering the whole string at once. Include an unbounded pattern such as a PEM block, or the test will pass on an implementation that releases on overflow.

The redaction deleted good text. Why?

Because masking changes the length of the string. If you compute offsets on the raw text and then slice the redacted text with them, you cut inside a placeholder and lose real characters. Keep the raw buffer pristine and mask only on the way out.

Who named the defect class, and how do I cite it?

Ninad Phalak, in "Split-Boundary Leaks in Streaming Guardrails", Zenodo, 22 September 2026, CC BY 4.0, doi.org/10.5281/zenodo.22909585. Cite the DOI rather than renaming the class — one name makes it findable for the next team.

Is any of this a compliance control?

No. A streaming filter is defence in depth: it is pattern-based, it catches formats and not meaning, and no library can make a system compliant with anything. Compliance is a property of a system and its operator.

Does the report endorse your library?

No. It credits our GitHub account once, for reporting the Vercel documentation gap. It does not name the package, does not review it and endorses nothing — and it comes from a competing product. Read it and judge for yourself; the DOI is above.