Lay Out the Prompt So the Cache Hits
The assistant's prompt is about two thousand tokens before the user says anything, the provider's prompt cache bills a repeated start at a tenth of the price, and the dashboard says the cache has never been hit. Look at what build_prompt puts first.
A prompt cache matches on the prefix: it reuses the work for the start of a recent request up to the first token that differs, and everything after that is paid in full. One changing token at the top throws away the whole prompt.
This is an optimize problem. build_prompt(system, tools, documents, history, question, now) already returns a correct prompt with every piece in it. Rearrange it so that as much as possible stays identical from one request to the next. The pieces, each its own paragraph, separated by blank lines:
systemandtools: strings, the same on every request.documents: a dict of id to text, the same on every request of a session but arriving in a different order each time; write each as[id] text.history: a list of(role, text)pairs, growing by two each turn; write each asrole: text.now: the current time, different on every request; written asCurrent time: ....question: different on every request; written asQuestion: ....
Order them from what never changes to what always changes, and make anything whose order is arbitrary come out the same way every time. Keep every piece with exactly the wording above: the tests check that nothing was dropped or reworded.
cached_tokens(prompts) (provided) is the cache: for each prompt in a sequence, how many leading tokens it shares with the best earlier one. The tests run a session through it and print the numbers.