Skip to content

< all problems16 · Level 07, Evals

Score an Answer with an LLM Judge

medium · implement · Evals

Have a model grade an answer against a rubric.

Implement judge(llm, question, answer) returning an integer from 1 to 5.

  1. Put the rubric in the prompt, 1 for useless up to 5 for excellent. A judge told only "score this answer" invents a new scale on every call.
  2. Find the digit. The judge will happily reply "I'd rate this a solid 4 out of 5 because...". Return an int however it phrases things.
  3. Clamp into 1 to 5. If parsing fails or the model returns 9, an out-of-range score would skew every average downstream.

A real model answers, so the tests check properties: an int in range, a good answer outscoring a nonsense one, a correct answer scoring at least 4.