Score a Model Against a Golden Set
A golden set is a list of inputs with known-correct outputs that you re-run whenever anything changes.
Implement two functions:
classify(llm, text)asks the model for a sentiment label and returnspositive,negativeorneutral.evaluate(llm, cases, classify), wherecasesis a list of(text, expected_label), runs every case and returns:
{"total": 6, "correct": 5, "accuracy": 0.833, "failures": [(text, expected, got), ...]}
- Accuracy rounded to 3 places. Noise in the fourth decimal makes two identical runs look different.
- Failures captured, not just counted. The failing cases tell you what regressed.
- One bad case must not kill the run. If
classifyraises or returns something unparseable, record a failure and keep going.
Normalise before comparing: case and whitespace differences are not real disagreements.