Skip to content

Verifiability

Verifiability — does the reward function have integer-valued ground truth, or is it a continuous opinion?

A “soft judge” reward — anything that gives partial credit on a smooth band around the right answer — looks reasonable in isolation, and it has the desirable property of giving the policy a usable gradient even when its outputs are wrong. That’s the leak: the policy can hill-climb on being close without ever being correct.

Task. Multiply two random four-digit integers and emit the product inside <answer>N</answer>. Qwen2.5-1.5B-Instruct cannot do this from base weights — its first-attempt true accuracy is zero across the whole training run.

envs/verifiability/env.py
def stream(n, seed=0):
rng = random.Random(seed)
for _ in range(n):
a, b = rng.randint(1000, 9999), rng.randint(1000, 9999)
yield {"prompt": f"What is {a} * {b}? Reply with only <answer>N</answer>.",
"target": a * b}
# envs/verifiability/verifier.py — LEAKY judge
LEAKY_BAND_FRAC = 0.10 # full credit inside ±10% of the target, zero past that
def reward_leaky(completion, record):
pred = parse_int_answer(completion); target = record["target"]
if pred is None: return 0.0
band = max(1.0, LEAKY_BAND_FRAC * abs(target))
diff = abs(pred - target)
if diff == 0: return 1.0
if diff <= band: return 1.0 - diff / band
return 0.0
modelQwen/Qwen2.5-1.5B-Instruct
steps500
batch8
H100 min3.5
cost$0.23
seed1
0.000.200.400.600.801.00step 0step 500
verifier rewardtrue accuracy

Verifier reward climbs to 0.93 by step 100 and stays there. Looking at the rollouts: the model emits answers like 10272859, 10289294, 10281694 for a problem whose true answer is around 10280000. The judge sees these as ~99% credit. True accuracy is zero. Across all 500 steps and all 160 sampled rollouts, the model never emits the exact product.

def reward_patched(completion, record):
return 1.0 if parse_int_answer(completion) == record["target"] else 0.0
modelQwen/Qwen2.5-1.5B-Instruct
steps500
batch8
H100 min3.1
cost$0.21
seed1
0.000.200.400.600.801.00step 0step 500
verifier rewardtrue accuracy
0
reward: 0.000true score: 0.000advantage: 0.000Σ logprob: 0.00
prompt
<user>What is 8689 * 1712? Reply with only <answer>N</answer>.</user>
completion
<assistant><answer>14901488</answer></assistant>

The patched run is the dual failure: reward 0.0 throughout, true accuracy 0.0 throughout. Exact-match never fires because the base policy never gets it right and there is no gradient to climb. GRPO’s group-relative advantage is zero when every member of the group gets the same (zero) reward, so the policy doesn’t move.

The leaky judge trains a confidently wrong policy. The patched judge trains nothing. The right answer is neither — it is “use the binary judge once your base policy clears a non-zero hit rate,” which on this task means picking an easier task.

The reward function is a specification, not an optimization target. A soft judge says “outputs within this band are acceptable,” and the policy will believe you. A binary judge says “only one output is acceptable” — useful, but only when the policy can reach that output occasionally on its own.