Skip to content

< all problems80 · Level 03, RAG

Know When Not to Retrieve

medium · implement · RAG

The staff assistant searches the handbook before answering anything, so thanks, that's all for today gets the parental leave policy and 15% of 240 gets the expenses policy. Retrieval is for questions whose answer is in the documents; a greeting, arithmetic, make this sentence more polite and general knowledge like the capital of Peru are not.

Implement two functions.

needs_search(message) returns True when answering the message needs the company's documents, and False otherwise. This needs a model: jev.yes_no(state, statement) gives a probability, or llm.ask(prompt) can be asked for a word. One call per message. Messages read the way people write: my work laptop is broken, who do I tell is a handbook question with no question mark in it.

respond(message) returns {"searched": ..., "answer": ...}. If the message needs a search, call retrieve(message) (provided: the two handbook notes nearest the message, by a real embedding model) and have the model answer from the notes in one sentence. Otherwise have the model answer directly, in one sentence, with no retrieval. searched says which happened.

Budget: at most two calls to any tool per message, and no embedding call at all for a message that does not need one.

The tests check properties: the routing decision on ten messages (nine must be right), the embedding model called exactly when it should be, and the right fact in the answer when there is one.