Skip to content

< all problems39 · Level 01, LLM APIs

Trim the Conversation, Keep the System Prompt

medium · implement · LLM Fundamentals

A long chat will not fit in the context window. Implement trim_history(messages, max_tokens) returning a message list that fits, measured with count_tokens from agentkit.tokenizer.

The rules, in priority order:

  1. The system prompt is never dropped. It is the instructions.
  2. The most recent user message is never dropped. That is the thing being answered.
  3. Drop from the oldest end, in pairs: a user turn together with the assistant reply that followed it. A history must not open with an orphaned assistant message.
  4. If the system prompt and the last user message alone still exceed the budget, truncate the last user message from the front, keeping its tail, until it fits.

Return the messages unchanged when they already fit.