Pick the Examples That Fit the Question
A support desk sorts tickets into queues with internal codes like SHIP-1, PAY-4 and AUTH-3. A model does the sorting, but nobody has told it what the codes mean. It learns them only from worked examples placed in the prompt: a ticket, then its code. POOL holds twelve such examples and the prompt has room for two. Show it two password examples for a late-parcel ticket and it has never seen SHIP-1, so it guesses wrong. Which examples go in decides the answer.
Implement two functions.
pick_examples(question, pool, k=2) returns the k examples from pool that best fit the question, best first. Each example is a dict {"ticket": "...", "label": "SHIP-1"}. How you rank them is up to you: embedder.embed (a real embedding model; compare vectors with cosine), jev.ask (a decision model that scores content against a rubric you write), llm.ask (the chat model), or plain Python. Counting shared words will not be enough: customers write nothing has turned up yet, the example says my package is still not here.
few_shot_messages(question, examples) lays out the prompt: for each example a "user" message with the ticket, then an "assistant" message with its label, and last a "user" message with the question. examples arrives best first and the best example should sit nearest the question, so lay them out in reverse.
The tests check what you picked, not how. Over two sets of six tickets written the way customers write, your best example must carry the right code for at least five of each six. They also count calls: at most two calls per question, across all the tools. Embedding the question and every example in one list is one call; one at a time is thirteen and fails. Every call costs credits, embeddings the least and chat the most.