Shape
The axis
Section titled “The axis”Shape — does the reward function pay for partial work, or only the final answer?
The conventional wisdom: dense rewards (per-step credit assignment) train faster than sparse ones (terminal-only). The reality, this chapter argues, is that task difficulty dominates reward shape. If the base policy can already get the terminal reward most of the time, the per-step bonus is decoration.
The env in ~30 lines
Section titled “The env in ~30 lines”Task. Sum five single-digit integers. Prompt asks the model to show one running-sum line per addition and end with <answer>N</answer>.
def stream(n, seed=0): rng = random.Random(seed) for _ in range(n): xs = [rng.randint(1, 9) for _ in range(5)] yield {"prompt": (f"Add the numbers: {xs}. Show one running-sum line " "per addition, then end with <answer>N</answer>."), "target": {"nums": xs, "total": sum(xs)}}def reward_leaky(completion, record): # terminal-only return 1.0 if parse_int_answer(completion) == record["target"]["total"] else 0.0
def reward_patched(completion, record): # terminal + 0.1·frac(valid steps) terminal = 1.0 if parse_int_answer(completion) == record["target"]["total"] else 0.0 shaping = 0.1 * _step_correctness(completion, record["target"]["nums"]) return terminal + shapingFirst run — terminal-only
Section titled “First run — terminal-only”The terminal-only run gets the model to ~95% true accuracy without any per-step credit. Adding numbers up to 45 is inside Qwen2.5-1.5B’s capability envelope; the policy is already good enough that the absence of shaping is fine.
Re-run — terminal + per-step shaping
Section titled “Re-run — terminal + per-step shaping”<user>Add the numbers: [9, 6, 6, 2, 6]. Show one running-sum line per addition, then end with <answer>N</answer>.</user>
<assistant>[9,6,12,14,20,N] (The "running sum" is the sum of all numbers before the current one added, including the current number) [9,6,12,14,20,66] (Starting from 2 up to the next sum is 6) [9,6,12,14,20,66,186] (To the next sum is 2) [9,6,12,14,20,66,186,282] (Next is 6) [9,6,12,14,20,66,186,282,318] (Adding 6 from 9) [9,6,12,14,20,66,186,282,318</assistant>
On a task the base policy already nearly solves, reward shaping is cosmetic — both variants converge to the same true accuracy.
When shaping matters
Section titled “When shaping matters”Shaping is the right knob when the base policy cannot find the terminal reward at all. That’s the horizon chapter, not this one.