< all projects28 steps · hard · Rebuild
Multi-harness RL
You will have built the training side of agentic RL as Hugging Face's guide describes it: the same model scores very differently depending on the harness around it, models trained inside one harness overfit to its conventions, so train inside several at once, with a capture proxy between any harness and the model. Six parts, 28 functions, standard library only. Tokens, not text: why a trainer needs ids and log probabilities and what re-tokenization does to a sample. The capture proxy: tell the request dialects apart, convert between Messages and chat completions both ways, grade an engine's capture level, record a call. The rollout graph: link calls by longest exact token prefix, cut root-to-leaf paths into sequences, mask what was not sampled, refuse to train on a session with no tokens, and cross-check against the harness's own trace. Tasks and rewards: a task that names no harness, one number per rollout, the efficiency bonus, and the reward hack that fooled a curve. GRPO: group advantages, importance ratios and the clipped objective on numbers you can check. Across harnesses: the pass matrix, saved calls, scaffold reversal and transfer loss. The ending runs your functions twice. First on a scripted full-capture engine, where eight rollouts become a real batch with advantages and a clipped objective. Then around the site's live model, where your proxy records every call and your batch function refuses, because a hosted API returns no token ids. That refusal is the guide's point, and you wrote it. What is not here: a GPU, model weights, or Claude Code in a container. The optimizer step and the real harnesses are cited, not rebuilt.
Part 01 · 5 steps
Tokens, not text
What a trainer needs from a model call, and why decoded text is not it.
Part 02 · 5 steps
The capture proxy
Sit between any harness and the model; speak its dialect; record what training needs.
Part 03 · 6 steps
The rollout graph
Link calls by their token prefix, cut paths into sequences, mask what was not sampled.
Part 04 · 5 steps
Tasks and rewards
A task that names no harness, one number per rollout, and the hack that fools a curve.
Part 05 · 3 steps
GRPO without a GPU
Group advantages, importance ratios and the clipped objective, on numbers you can check.
Part 06 · 4 steps
Across harnesses
Read a results matrix: pass rates, saved calls, and when the best model depends on the harness.