Skip to content

< all problems09 · Level 03, RAG

Fix a Broken RAG System

hard · debug · RAG

This RAG pipeline (chunk_all, embed_texts and retrieve below) runs without errors but retrieves the wrong documents. It has three defects, each one line:

  • Long documents win regardless of what they say.
  • The results look like the least relevant documents, not the most.
  • The text attached to a result does not match the text that earned its score.

Each symptom has one cause. Find them and fix retrieve(question, documents, tokenizer, model, k). Don't rewrite it from scratch: work out which line lies and change that.

The tests check that relevant documents come back, in the right order, with the right text attached.