Stay Under the Token Budget
answer_with_context(llm, question, docs) below works, but it sends the entire knowledge base to the model on every call, and you pay per token for all of it.
This is an optimize problem: keep answering correctly and come in under the budget the tests set. Budget(max_total_tokens=150) is far too small for six documents and comfortably enough for one or two, so decide which documents the question needs before you send anything. The question has words in it, and so do the documents.
Two rules:
- The relevant document must still reach the model. Trimming to nothing is under budget and useless. The tests check that the document holding the answer was actually in the prompt.
- Measure, do not guess.
Meter.of(llm)reports exactly what was sent. Print it while you work.
You are scored on total_tokens.