Vercel AI SDK guardrails: fixing the split-match hole
The AI SDK middleware examples did not compile and wrapStream was left unimplemented. What we reported, what merged on 22 September 2026, and the rule.
Five examples, none of them compiled
The AI SDK documents language-model middleware on a single page, and that page carried five worked examples. Copied into a project running ai@7.0.107 and @ai-sdk/provider@4.0.17, every one of them failed the same way:
error TS2741: Property 'specificationVersion' is missing in type
'{ wrapGenerate: ... }' but required in type 'LanguageModelV4Middleware'.
specificationVersion: 'v4' is a required readonly field on the middleware type, and the page did not mention it anywhere — not in the first example, not in a note, not in passing. Nothing subtle was wrong here. Every example on the page failed on the same missing property, whatever each one was demonstrating. A reader following the page got this error on their first middleware, and nothing on the page to explain it.
The half that was left as a comment
The second problem was the larger one, and the page was honest about it. The guardrail example implemented wrapGenerate, ran a .replace() over the finished text, and left the streaming half as a comment, reproduced here exactly as it was captured:
here you would implement the guardrail logic for streaming // Note: streaming guardrails are difficult to implement, because you do not know the full content of the stream until it's finished.
That note is true, and it states the whole difficulty plainly. What it does not say is what happens to the reader who takes it as an invitation. The obvious move is to lift the .replace() out of wrapGenerate and drop it into a transform over the stream parts. The result looks like a working guardrail, survives a hand-test with a short prompt, and is wrong in two separate ways.
The first way: the match is never there to find
A model emits text in deltas cut on token boundaries, not on meaning. Any value long enough to be worth redacting is usually spread across several of them.
delta 1: "...your card is 4111 11"
delta 2: "11 1111 1111, charge it."
A .replace() run on delta 1 finds no card number. Run on delta 2, it finds no card number. Both deltas go to the client, the browser joins them, and the number is on screen. At the delta sizes providers actually send, this is not an edge case — for most values worth filtering, it is the common case.
The second way: the half-redacted secret
The repair people reach for is to hold a window of recent text, scan it, and release whatever falls out the back. That one survives review, and it fails worse than the first, because it produces output that looks handled.
Some patterns stop growing; many do not. A pattern for a variable-length secret matches as soon as the text is long enough to qualify, and it goes on matching as more characters arrive. If the release point lands inside a value that is still growing — and with a fixed window it often does — the filter masks the part it is holding and streams the rest in clear:
out: "your key is <REDACTED>uvwxyz123456"
Half the secret is in the transcript, sitting next to a placeholder that tells the reader it was removed. A visible leak is a bug. A leak wearing a redaction label is a bug plus a false assurance, and the second part is what makes it expensive: nobody goes back to look.
The rule
Never act on a match that could still grow, and never emit text you have not scanned in its final form.
Both halves carry weight. The first rules out masking a value before you know where it ends. The second rules out releasing text you have only ever seen in pieces. In practice it means holding from the start of any match that runs past your release point, rather than from a fixed offset behind it. A fixed offset is a guess about how long secrets are. The start of the match is a fact about the text in front of you.
The property, written as a test
The rule is a design instruction. This is the machine-checkable form of it:
For any input and any chunking, the streamed output is byte-identical to filtering the whole string at once.
Replay the same input at every delta size — one character, two, three, up to the whole string — and compare each run against the single-pass result. Both failures above turn into a mismatch, and nobody has to reason about which boundary was the unlucky one. It is short, and it does not care how the filter is built. The defect class it catches has a name and a longer write-up: PII split across stream chunks.
What got merged
The report was vercel/ai#21209, "docs: guardrails example leaves wrapStream unimplemented", filed on 20 September 2026. It closed on 22 September after nine comments. On 28 September 2026 the repository had 27,011 stars. Three pull requests merged within sixteen seconds of each other:
- #21210 onto main — 1 file, +92/-7
- #21211 onto release-v6.0 — 1 file, +92/-7
- #21212 onto release-v5.0 — 2 files, +151/-21
That is +335/-35 across three branches. The issue carries the labels merged-main, merged-v5.0, merged-v6.0 and documentation-fix, among others. As of 28 September 2026 the live middleware page carries specificationVersion: 'v4'.
The wrapStream example that landed has no dependency on any library. It keys one buffer per text id — a stream can carry several text blocks at once, which is exactly what the id on each part is for, and a single shared buffer attributes text to the wrong block. It computes a cut, walks the matches in the buffer, and when a match runs past the cut it holds from m.index rather than from the cut. It never cuts inside a surrogate pair, so an emoji does not arrive as two broken halves. Read as a whole, it is the rule above expressed in code instead of described in a comment.
The disclosure came first
One detail matters more than the diff. The issue opened with this, before the problem was even described: "Disclosure: I maintain an MIT library that does streaming redaction, so I have an interest here. The suggestion below is deliberately self-contained — it has no dependency on it and you can drop it straight into the docs."
The library is llm-stream-guardrails, and it went to npm on 21 September 2026 — the day after the issue was filed. The report came first, and what it proposed was a patch a maintainer could take without taking anything else along with it. That is the only arrangement under which someone maintaining an alternative should be filing against a project this size: declare the interest at the top, then make the suggestion good enough to survive the interest being declared.
If you are writing one of these
Take the rule and the invariant; nothing else on this page is needed. Both work with no dependency at all, and the example now in the AI SDK docs is a complete implementation of them. Other libraries handle chunk boundaries with bounded lookahead and are fine — the question is not which package sits on the import line, it is whether the output survives being replayed at every chunk size. If it does not, the package does not matter.
The limits are worth stating plainly, because a streaming filter invites more confidence than it has earned. It matches formats, not meaning: it can catch a card number and will not catch a customer's medical history written out in prose. It runs after the model has decided what to say, so it is a last line rather than a first one. And no filter in any library makes a system compliant with anything; compliance is a property of the system and the people operating it. What the rule buys you is narrower, and worth having on its own: output that is the same whether it arrived in one piece or four hundred.
Questions engineers ask about this
How do I fix "error TS2741: Property 'specificationVersion' is missing"?
Add specificationVersion: 'v4' to the middleware object. It is a required readonly field on LanguageModelV4Middleware, so an object carrying only wrapGenerate or wrapStream will not type-check. The AI SDK middleware page never mentioned the field, which is why all five of its examples failed against ai@7.0.107 and @ai-sdk/provider@4.0.17. The page was corrected on 22 September 2026 and now shows the field. If you pasted an example before that date, this one line is the whole fix.
Why didn't the AI SDK middleware examples compile?
Every example on the page failed on the same missing property, whatever each one was demonstrating: specificationVersion is a required field on the middleware type, and the page did not mention it anywhere — not in the first example, not in a note. The failure was the same one line missing from each object. Three pull requests merged on 22 September 2026 corrected the page on the main branch and on both release branches.
Can I move my wrapGenerate .replace() into wrapStream?
You can, and it will not do what it looks like it does. A .replace() applied to each delta only ever sees that delta, and a value worth redacting is usually spread over several of them. "4111 11" arrives, then "11 1111 1111" — neither piece matches a card pattern, both go to the client, and the browser joins them on screen. The move that feels like a small refactor is the one that produces a filter which passes a hand-test and leaks in production.
My streaming redaction never matches anything. What am I missing?
The text you are matching against was never a unit. Providers cut deltas on token boundaries, not on meaning, so a card number, an email address or an API key routinely straddles two or more of them. Per-delta matching finds neither half. The fix is not a better pattern, it is state: keep a buffer across deltas, and hold back any text that a match could still grow into, rather than scanning each delta as if it were a finished document.
Should I just buffer the whole response and filter it at the end?
It is correct, and it is the simplest thing that satisfies the invariant. It also gives up the only reason to stream: time to first token becomes time to full generation, and memory grows with the length of the block. That trade is fine for short answers and painful for long ones. Holding from the start of an open match keeps the stream flowing and gives identical output, at the cost of a little more code in the transform.
What hold-back size should I use — 32, 64, 128 characters?
None of them, on its own. Any fixed N can be defeated by a match longer than N, and when it is, it fails quietly, because the output still carries a redaction placeholder. The held amount has to come from the text, not from a constant: hold from the index where a still-open match begins, and release everything before it. If you must cap the buffer, decide in advance what happens at the ceiling, and make the answer "refuse", not "release and hope".
Why does my output show a redaction placeholder followed by half an API key?
Because your release point landed inside a value that was still growing. Variable-length secrets match as soon as they are long enough to qualify, so a filter holding a fixed window masks the characters it has and streams the rest untouched. The output reads as <REDACTED>uvwxyz123456: half the key, next to a label saying it was removed. It is worse than no filter, because a visible leak gets reported and a labelled one does not.
Why one buffer per text id instead of one shared buffer?
Because a single stream can carry several text blocks at once, and that is exactly what the id on each part is for. One shared buffer mixes them, so text from one block is scanned, held and released as part of another — wrong redactions in one place and wrong ordering in the other. The example merged into the AI SDK docs keys its buffer by text id for exactly this reason.
How do I test a streaming guardrail?
One assertion, run over every fixture at every chunk size: the concatenated streamed output must be byte-identical to filtering the whole string at once. Replay the same input at one character per delta, then two, then three, up to the whole string, and compare each run to the single-pass result. Split matches and half-redacted values both show up as a mismatch, and no one has to guess which boundary is the dangerous one. It is a short piece of test code.
Which AI SDK branches got the fixed wrapStream example?
Three branches, on the same day. Pull request #21210 landed on main, #21211 on release-v6.0 and #21212 on release-v5.0, all merged on 22 September 2026 within sixteen seconds of each other — 1 file each on the first two, 2 files on the v5 branch, +335/-35 in total. As of 28 September 2026 the live middleware documentation carries specificationVersion: 'v4'.
Do I need a library to redact a stream?
No. The example now in the AI SDK documentation is self-contained and has no dependency on anything, and the rule it implements is short enough to write yourself: never act on a match that could still grow, never emit text you have not scanned in its final form. Several libraries handle chunk boundaries correctly with bounded lookahead. What matters is not the name on the import line but whether the output survives being replayed at every chunk size.
Is this only a Vercel AI SDK problem?
No. The SDK shipped no faulty filter — it shipped a documentation gap, with the streaming half of a guardrail left as a comment. The failure it invites has been shipped independently by unrelated projects in more than one language, which is what makes it a defect class rather than one team's mistake. There is a separate write-up on this site covering those cases, the name the class was given, and the citation for it.
Will buffering and cutting the text break emoji?
It will if you cut naively. A character outside the Basic Multilingual Plane is stored as two code units, and slicing between them emits half a character that no downstream consumer can repair. The merged example checks for this and moves its cut rather than splitting a surrogate pair. The same care applies to anything you build: the release point is a position in a string, and not every position in a string is a valid place to stop.
Why did the issue start with a disclosure?
Because it was true, and it changes how the suggestion should be read. A maintainer of an MIT library that does streaming redaction was reporting a gap that his own library fills. The suggestion was written to be self-contained precisely so that the interest would not matter — no dependency, nothing to install, paste it into the docs. Declaring it first is what makes the rest of the report usable by people who owe you nothing.