Skip to content

< all problems27 · Level 07, Evals

Score a Model Against a Golden Set

medium · implement · Evals

A golden set is a list of inputs with known-correct outputs that you re-run whenever anything changes.

Implement two functions:

  • classify(llm, text) asks the model for a sentiment label and returns positive, negative or neutral.
  • evaluate(llm, cases, classify), where cases is a list of (text, expected_label), runs every case and returns:
{"total": 6, "correct": 5, "accuracy": 0.833, "failures": [(text, expected, got), ...]}
  1. Accuracy rounded to 3 places. Noise in the fourth decimal makes two identical runs look different.
  2. Failures captured, not just counted. The failing cases tell you what regressed.
  3. One bad case must not kill the run. If classify raises or returns something unparseable, record a failure and keep going.

Normalise before comparing: case and whitespace differences are not real disagreements.