Skip to content

Shape

Shape — does the reward function pay for partial work, or only the final answer?

The conventional wisdom: dense rewards (per-step credit assignment) train faster than sparse ones (terminal-only). The reality, this chapter argues, is that task difficulty dominates reward shape. If the base policy can already get the terminal reward most of the time, the per-step bonus is decoration.

Task. Sum five single-digit integers. Prompt asks the model to show one running-sum line per addition and end with <answer>N</answer>.

envs/shape/env.py
def stream(n, seed=0):
rng = random.Random(seed)
for _ in range(n):
xs = [rng.randint(1, 9) for _ in range(5)]
yield {"prompt": (f"Add the numbers: {xs}. Show one running-sum line "
"per addition, then end with <answer>N</answer>."),
"target": {"nums": xs, "total": sum(xs)}}
envs/shape/verifier.py
def reward_leaky(completion, record): # terminal-only
return 1.0 if parse_int_answer(completion) == record["target"]["total"] else 0.0
def reward_patched(completion, record): # terminal + 0.1·frac(valid steps)
terminal = 1.0 if parse_int_answer(completion) == record["target"]["total"] else 0.0
shaping = 0.1 * _step_correctness(completion, record["target"]["nums"])
return terminal + shaping
modelQwen/Qwen2.5-1.5B-Instruct
steps500
batch8
H100 min9.6
cost$0.63
seed2
0.000.200.400.600.801.00step 0step 500
verifier rewardtrue accuracy

The terminal-only run gets the model to ~95% true accuracy without any per-step credit. Adding numbers up to 45 is inside Qwen2.5-1.5B’s capability envelope; the policy is already good enough that the absence of shaping is fine.

modelQwen/Qwen2.5-1.5B-Instruct
steps500
batch8
H100 min6.3
cost$0.42
seed2
0.000.200.400.600.801.00step 0step 500
verifier rewardtrue accuracy
0
reward: 0.000true score: 0.000advantage: -0.759Σ logprob: 0.00
prompt
<user>Add the numbers: [9, 6, 6, 2, 6]. Show one running-sum line per addition, then end with <answer>N</answer>.</user>
completion
<assistant>[9,6,12,14,20,N] (The "running sum" is the sum of all numbers before the current one added, including the current number) 

[9,6,12,14,20,66] (Starting from 2 up to the next sum is 6) 

[9,6,12,14,20,66,186] (To the next sum is 2) 

[9,6,12,14,20,66,186,282] (Next is 6) 

[9,6,12,14,20,66,186,282,318] (Adding 6 from 9) 

[9,6,12,14,20,66,186,282,318</assistant>

On a task the base policy already nearly solves, reward shaping is cosmetic — both variants converge to the same true accuracy.

Shaping is the right knob when the base policy cannot find the terminal reward at all. That’s the horizon chapter, not this one.