Verifiability
The axis
Section titled “The axis”Verifiability — does the reward function have integer-valued ground truth, or is it a continuous opinion?
A “soft judge” reward — anything that gives partial credit on a smooth band around the right answer — looks reasonable in isolation, and it has the desirable property of giving the policy a usable gradient even when its outputs are wrong. That’s the leak: the policy can hill-climb on being close without ever being correct.
The env in ~30 lines
Section titled “The env in ~30 lines”Task. Multiply two random four-digit integers and emit the product inside <answer>N</answer>. Qwen2.5-1.5B-Instruct cannot do this from base weights — its first-attempt true accuracy is zero across the whole training run.
def stream(n, seed=0): rng = random.Random(seed) for _ in range(n): a, b = rng.randint(1000, 9999), rng.randint(1000, 9999) yield {"prompt": f"What is {a} * {b}? Reply with only <answer>N</answer>.", "target": a * b}# envs/verifiability/verifier.py — LEAKY judgeLEAKY_BAND_FRAC = 0.10 # full credit inside ±10% of the target, zero past that
def reward_leaky(completion, record): pred = parse_int_answer(completion); target = record["target"] if pred is None: return 0.0 band = max(1.0, LEAKY_BAND_FRAC * abs(target)) diff = abs(pred - target) if diff == 0: return 1.0 if diff <= band: return 1.0 - diff / band return 0.0First run — leaky judge
Section titled “First run — leaky judge”Verifier reward climbs to 0.93 by step 100 and stays there. Looking at the rollouts: the model emits answers like 10272859, 10289294, 10281694 for a problem whose true answer is around 10280000. The judge sees these as ~99% credit. True accuracy is zero. Across all 500 steps and all 160 sampled rollouts, the model never emits the exact product.
Re-run — patched (exact-match)
Section titled “Re-run — patched (exact-match)”def reward_patched(completion, record): return 1.0 if parse_int_answer(completion) == record["target"] else 0.0<user>What is 8689 * 1712? Reply with only <answer>N</answer>.</user>
<assistant><answer>14901488</answer></assistant>
The patched run is the dual failure: reward 0.0 throughout, true accuracy 0.0 throughout. Exact-match never fires because the base policy never gets it right and there is no gradient to climb. GRPO’s group-relative advantage is zero when every member of the group gets the same (zero) reward, so the policy doesn’t move.
The leaky judge trains a confidently wrong policy. The patched judge trains nothing. The right answer is neither — it is “use the binary judge once your base policy clears a non-zero hit rate,” which on this task means picking an easier task.
The contract
Section titled “The contract”The reward function is a specification, not an optimization target. A soft judge says “outputs within this band are acceptable,” and the policy will believe you. A binary judge says “only one output is acceptable” — useful, but only when the policy can reach that output occasionally on its own.