Open source · ROSH™ Company Labs

cut-at-k what a truncated stream tells your client

cut-at-k replays a stream whole and at every cut point, then checks whether your client admits it was cut short. MIT, zero deps, 21 tests. ROSH™ Company Labs.

A stream that stops early is not a rare case

A connection drops. A proxy times out. The producer dies mid-sentence, or a reader presses stop. The stream ends where it ends, and the client library has to tell its caller something about what just happened. Testing that path means choosing a cut point.

cut-at-k removes the choice. It takes one stream, runs it through your client once whole and once for every prefix, and asks two questions of each cut:

The second question is where things are usually wrong, and it is quiet when it is wrong: the promise resolves, the result object looks ordinary, and nothing on it says the answer is half an answer.

It knows nothing about your protocol

You supply the events and a replay function. It supplies the loop and the comparison. Server-sent events, WebSocket frames, a JSON-lines fixture, an array you built by hand — it never looks inside an event.

npm install cut-at-k
import { severAtEveryPoint } from 'cut-at-k';

await severAtEveryPoint({
  events,                                   // the stream, as an array
  label: 'checkout-run',
  replay: (prefix) => runMyClient(prefix)   // called whole, then per prefix
});

Three entry points carry the work: severAtEveryPoint runs the replays, summarise reduces them to a report, format prints it. Zero dependencies, ten files, MIT.

It refuses to cry wolf

Three things look identical in the raw numbers, and only one of them is a finding.

So summarise reports nothing at all as a defect unless you pass it a lostContent predicate. You decide what counts as content in your protocol; the tool will not guess on your behalf. Streams excluded as invalid on purpose come back as a count rather than disappearing quietly, so you can always see how much of a corpus was set aside.

This was learned by getting it wrong. The first version of summarise made none of those distinctions and called 154 cuts a problem on a corpus where 54 were. The README keeps the sentence that came out of it: "A tool that cries wolf is worse than no tool."

Run against AG-UI: 227 cuts, 54 that lost something quietly

AG-UI ships 68 conformance fixtures written by the protocol's own authors. Two are a single event long and have no cut point, so 66 were replayed — each through a real HttpAgent over real HTTP and SSE on @ag-ui/client 1.0.0, whole and then cut at every event boundary.

48 streams, 227 cuts, 18 streams excluded as invalid on purpose
4 cuts threw during replay

reported the same as the whole run    : 154
  of which something was actually lost : 54
  of those, no terminal event either   : 52
spread across                          : 34 streams

The smallest case is the clearest one. Take the fixture named conformant-run-is-quiet and cut it after its fourth event. The truncated call and the completed call both resolve. Neither rejects. RunAgentResult carries no field saying which of them was cut short. The message count is 1 either way. A caller that checks whether the run succeeded, or how many messages came back, is told exactly the same thing by a run that finished and a run that stopped mid-sentence.

Four of those 227 cuts threw during replay. They stay in the total and are set aside before the comparison, because a cut that crashes the client is a different finding from a cut that reported the wrong thing.

What this page does not claim

A subscriber can detect this today. onRunFinishedEvent fires on the completed run and not on the truncated one, while onRunFinalized fires on both — so the difference is observable if you are listening on that channel. The claim here is narrower: the awaited call does not surface it. That happens to be the channel most callers use, which is why it is worth writing down, but it is not the same as saying the information is unavailable.

Nor is any of this a discovery. The behaviour was reported on 3 August 2026 as ag-ui-protocol/ag-ui#2300, and PR #2354 has been open since 7 August 2026 carrying the fix. Both were still open when this page was written. The argument for the tool is not novelty, it is reach: the bug was found once, by hand, on one stream, and the loop surfaces 54 cuts of that shape across 34 streams without anyone having to guess where to look.

Six libraries, the same question, four clean answers

The repository carries probes that ask five more codebases what they tell a caller when the stream is cut. Four of them answer correctly:

The MCP TypeScript SDK did not. When the response leg of a POST dies, the request waits out its whole timeout before the caller is told — at every cut point, including one where half the response had already arrived.

Six asked in all, counting AG-UI itself — four clean. Both of the two that answer wrongly, AG-UI and the MCP SDK, were found by someone else first. That is the honest accounting: this is a way of asking an old question everywhere at once, not a source of new findings.

Numbers

Cite it

The library has a DOI of its own: 10.5281/zenodo.23002965. A CITATION.cff ships in the repository, so GitHub's own Cite this repository button is populated from it.

Maintainer

Maintained by Redouane — ROSH™ Company Labs. Bugs and feature requests belong in GitHub issues. The most useful report this project can get is a corpus where the summary called something a finding that was not one.

FAQ

Questions developers ask about this

What does cut-at-k actually do?

It takes one stream and runs it through your client twice over: once whole, and once for every prefix. For each cut it asks two things. Did the head change — because truncation may lose the tail, but it must never alter what came before the cut. And does the consumer say so — because a run cut short must not report what the same run reports when it finishes. The second is where problems usually live, and it is silent when it is wrong.

\n
Why cut at every point instead of just testing a disconnect at the end?

Because a single disconnect test covers a single cut point. A cut after the last content event loses nothing and is fine. A cut a few events earlier can drop a whole message while the call still resolves normally — on the AG-UI corpus that is conformant-run-is-quiet, cut after its fourth event. Across those fixtures, 227 cuts produced 54 that lost something and said nothing about it, spread across 34 streams. The bug behind these was found once, by hand, on a single stream. Covering that spread by hand means writing 227 disconnect tests and guessing the right cut point in each.

\n
Does it need to understand my protocol — SSE, WebSocket, something custom?

No. You pass an array of events and a replay function, and it never looks inside an event. Server-sent events, WebSocket frames, JSON lines, or an array you assembled by hand all work the same way. Three entry points carry the work: severAtEveryPoint({events, label, replay}) runs the replays, summarise() reduces them, format() prints the report. Zero dependencies, ten files, MIT.

\n
Won't cutting a stream 227 ways just produce 227 false alarms?

That is exactly what the first version did — it called 154 cuts a problem on a corpus where 54 were. Three situations look identical in the raw numbers: a cut that lost nothing, a stream that is invalid on purpose, and a cut that lost content and was reported as complete. Only the third is a finding. summarise() now separates them, and reports nothing as a defect until you tell it what content means.

\n
What is the lostContent predicate and why do I have to write it?

It is the function that decides whether a cut actually dropped something a consumer would care about. Only you know that for your protocol: a trailing bookkeeping event is not content, a message body is. Without the predicate, summarise() reports no defects at all. That is deliberate. A tool that guesses what counts as content will guess wrong, and a report nobody trusts is worse than no report.

\n
What does "excluded as invalid on purpose" mean in the output?

Conformance corpora deliberately contain malformed streams so a client can be checked for rejecting them. Cutting before the deliberate error makes the truncated run deliver more than the whole one, which inverts the comparison and means nothing. Those streams are set aside — 18 of the 66 in the AG-UI run — and reported as a count rather than dropped silently, so you can see how much of the corpus was not actually exercised.

\n
What did it find in AG-UI?

66 of the 68 conformance fixtures were replayed through a real HttpAgent over HTTP and SSE on @ag-ui/client 1.0.0. That gave 48 streams and 227 cuts. 154 cuts reported the same as the whole run; of those, 54 had actually lost something, spread across 34 streams; 52 of the 54 had no terminal event either. The smallest case: cut conformant-run-is-quiet after its fourth event and both calls resolve, with one message each and no field saying which was truncated.

\n
Is that an AG-UI bug nobody knew about?

No, and the page says so. It was reported on 3 August 2026 as ag-ui-protocol/ag-ui#2300, and PR #2354 has been open since 7 August 2026 with the fix; both were open at the time of writing. The value on offer is not novelty. The bug was found once, by hand, on one stream — the loop surfaces 54 cuts of that shape across 34 streams without anyone guessing where to look.

\n
Can't an AG-UI subscriber already tell that a run was truncated?

Yes, and that limit belongs in any honest summary. onRunFinishedEvent fires on the completed run and not on the truncated one, while onRunFinalized fires on both, so a subscriber can see the difference today. The narrower claim is the one made here: the awaited call does not surface it. RunAgentResult has no field for it, both calls resolve, and the message count matches. That is the channel most callers actually use.

\n
Which other libraries were checked, and how did they do?

Five, in the probes directory: the Vercel AI SDK, LangGraph JS, Mastra, the OpenAI Node SDK and the MCP TypeScript SDK. Four came back clean. The Vercel SDK reports a finish reason of other on a cut run, or raises when no text got through; LangGraph's checkpoint stops where the aborted consumer stopped and says where to resume; Mastra's memory still holds what the consumer was shown; the OpenAI SDK's final response keeps the tool call it streamed and never claims completion.

\n
What went wrong in the MCP TypeScript SDK?

When the response leg of a POST dies, the request waits out its whole timeout before the caller is told — at every cut point tested, including one where half the response had already arrived. The caller is not handed a wrong answer so much as no answer for a long time. It was also found by someone else before this tool looked: six libraries were asked, four were clean, and both that answered wrongly already had prior reports.

\n
How do I run it against my own client?

Collect one stream as an array of events — a recorded fixture is enough. Write a replay function that feeds a prefix to your client the way production would, over a real transport if you can, and returns whatever your caller receives. Pass both to severAtEveryPoint with a label, add a lostContent predicate for your protocol, then summarise() and format(). Start with a stream you already trust; the interesting result is the cut you would not have thought to write.

\n
How mature is it, honestly?

Three days old at the time of writing. First published 25 September 2026, currently 0.1.3, and 0 stars on GitHub. It has 21 tests, all 21 passing, and ten files with no dependencies, so reading it end to end costs an afternoon at most. Read it before you trust its report. That advice is not modesty, it is the only basis on which a three-day-old tool should be used.

\n
Can I cite it?

Yes. It has a DOI of its own, 10.5281/zenodo.23002965, and ships a CITATION.cff, so GitHub's Cite this repository button is populated from it. It is MIT-licensed and published with a verified provenance attestation, so the tarball on npm can be traced back to the workflow that built it.