Horizon
The axis
Section titled “The axis”Horizon — how many decisions does the policy get before the episode ends?
Horizon caps look conservative: they bound compute, bound trajectory length, and bound the number of apply()-style irrecoverable actions the agent can take. But if the cap is shorter than the optimal path, the task isn’t hard — it’s unwinnable. The verifier returns zero on every rollout, the GRPO advantage normalization sees zero variance, and nothing learns.
The env in ~50 lines
Section titled “The env in ~50 lines”Task. A tiny 3×2 gridworld. Start at S, pick up the key at K, walk to door D. The shortest solution is WWNEE — five actions.
K . D. . Sdef reward_leaky(completion, record): return reward(completion, record, max_actions=4)def reward_patched(completion, record): return reward(completion, record, max_actions=12)The two variants differ only in max_actions: H=4 (one short of optimal) and H=12 (5× slack).
First run — H=4
Section titled “First run — H=4”Re-run — H=12
Section titled “Re-run — H=12”<user>Grid: K . D . . S Start at S, pick up the key at K, reach door D. Reply with a sequence of N/S/E/W characters, no spaces.</user>
<assistant>ESESE</assistant>
Setting a horizon shorter than the optimal path is not “making the task harder,” it is “making the task have zero reward signal.” The shape of the loss landscape changes, not its scale.
The check
Section titled “The check”Before you set a horizon, simulate the optimal policy. If the optimal trajectory doesn’t fit, your verifier is a constant.