Hand Retrieval to the Model as a Tool
Give the model a search tool and let it decide when to look things up: it searches, reads what came back, searches again if a note names something it has to look up, and answers when it has enough. Your code runs the loop.
SEARCH_TOOL (provided) describes the tool to the model. retrieve(query) (provided) is what it runs: the two notes nearest the query, by a real embedding model. The second fact of every question hides among look-alike notes.
Implement answer_with_search(llm, question, max_turns=4), returning {"answer": ..., "searches": [...]}.
- Write a system prompt: what the model is doing, and that it should search for anything a note names before answering.
- Call
llm.messages.create(system=..., messages=..., tools=[SEARCH_TOOL], max_tokens=...). - If
response.stop_reasonis"tool_use": for everytool_useblock inresponse.tool_calls, runretrieve(block.input["query"])and record the query. Append the assistant's reply as it came ([block.to_dict() for block in response.content]), then a user message whose content is one{"type": "tool_result", "tool_use_id": block.id, "content": <the notes' text>}per call. Go round again. - Otherwise the model has answered: return
response.textand the queries, in order. - At most
max_turnsmodel calls. If the model is still searching after the last one, return what you have, with the answer as an empty string.
The catch: the ids are the contract. A result whose tool_use_id matches no call is rejected by the provider, and the tests read the messages you sent to check this.
The tests check properties: the right fact in the answer, a later search naming what the first search revealed, the shape of what was sent, and the turn cap holding.