Skip to content

< all problems65 · Level 01, LLM APIs

Assemble a Reply From a Stream

medium · implement · LLM Fundamentals

With streaming on, a reply arrives as a few dozen small events instead of one object. Rebuild it.

Implement assemble(events). Each event is a dict with a "type":

  • message_start: carries message.usage.input_tokens.
  • content_block_start: opens the block at index; content_block is {"type": "text"} or {"type": "tool_use", "id": ..., "name": ...}.
  • content_block_delta: more of the block at index. A text block gets delta.text; a tool-use block gets delta.partial_json, a fragment of the JSON text of its arguments, which can split anywhere.
  • content_block_stop: the block at index is finished.
  • message_delta: carries delta.stop_reason and usage.output_tokens. These arrive last.
  • message_stop: the stream ended cleanly.
  • ping, and any type you do not recognise: ignore.
  • error: raise StreamError (provided) with the event's error.message.

Return a dict:

  • "text": all text blocks joined, in block order.
  • "tool_calls": a list of {"id", "name", "input"} in block order, input being the parsed arguments ({} if no fragments arrived).
  • "stop_reason", "input_tokens", "output_tokens": from the events above. None and 0 if they never arrived.
  • "complete": True only if message_stop was seen.

The catch: deltas for different blocks can interleave, so text goes to the block its index names, not the block that opened last. And a stream can die at any point: return what arrived, with complete false, and leave out any tool call whose arguments do not parse. Half a tool call must never reach a tool.