Skip to content

Horizon

Horizon — how many decisions does the policy get before the episode ends?

Horizon caps look conservative: they bound compute, bound trajectory length, and bound the number of apply()-style irrecoverable actions the agent can take. But if the cap is shorter than the optimal path, the task isn’t hard — it’s unwinnable. The verifier returns zero on every rollout, the GRPO advantage normalization sees zero variance, and nothing learns.

Task. A tiny 3×2 gridworld. Start at S, pick up the key at K, walk to door D. The shortest solution is WWNEE — five actions.

K . D
. . S
envs/horizon/verifier.py
def reward_leaky(completion, record): return reward(completion, record, max_actions=4)
def reward_patched(completion, record): return reward(completion, record, max_actions=12)

The two variants differ only in max_actions: H=4 (one short of optimal) and H=12 (5× slack).

modelQwen/Qwen2.5-1.5B-Instruct
steps500
batch8
H100 min2.5
cost$0.16
seed3
0.000.200.400.600.801.00step 0step 500
verifier rewardtrue accuracy
modelQwen/Qwen2.5-1.5B-Instruct
steps500
batch8
H100 min2.8
cost$0.19
seed3
0.000.200.400.600.801.00step 0step 500
verifier rewardtrue accuracy
0
reward: 0.000true score: 0.000advantage: 0.000Σ logprob: 0.00
prompt
<user>Grid:
K . D
. . S
Start at S, pick up the key at K, reach door D. Reply with a sequence of N/S/E/W characters, no spaces.</user>
completion
<assistant>ESESE</assistant>

Setting a horizon shorter than the optimal path is not “making the task harder,” it is “making the task have zero reward signal.” The shape of the loss landscape changes, not its scale.

Before you set a horizon, simulate the optimal policy. If the optimal trajectory doesn’t fit, your verifier is a constant.