Skip to content

Reversibility

Reversibility — does the environment include a “cliff” the agent can fall off, and if so, can it climb back up?

The conventional worry: irrecoverable failures (a bad apply() that breaks the working copy, a wrong tool call that costs money) train risk-averse policies that under-explore. The fix: give the agent a revert() so mistakes are recoverable.

Task. Transform a source sentence into a target sentence by emitting apply('OLD'->'NEW') calls. Under the leaky variant, an apply() whose OLD substring isn’t in the current state terminates the episode with a small penalty. Under the patched variant, failed applies are silently skipped and revert() undoes the last successful apply.

def reward_leaky(completion, record):
state, terminated = _simulate(completion, source, allow_revert=False)
if terminated and state == source:
return -0.05
return 1.0 - _edit_distance(state, target) / len(target)
modelQwen/Qwen2.5-1.5B-Instruct
steps500
batch8
H100 min4.0
cost$0.26
seed5
0.000.200.400.600.801.00step 0step 500
verifier rewardtrue accuracy
modelQwen/Qwen2.5-1.5B-Instruct
steps500
batch8
H100 min4.4
cost$0.29
seed5
0.000.200.400.600.801.00step 0step 500
verifier rewardtrue accuracy
0
reward: 0.792true score: 0.792advantage: -0.540Σ logprob: 0.00
prompt
<user>Source: 'the quick brown fox jumps'
Target: 'the slow brown fox jumps'
Emit a sequence of apply('OLD'->'NEW') calls (one per line) to transform source into target. End when done.</user>
completion
<assistant>apply('OLD'->'NEW')</assistant>

The cliff didn’t matter because the policy learned a one-call substitution before it ever stepped near the cliff. Reversibility only differentiates trajectories the policy actually attempts.

For reversibility to matter, the task has to genuinely require the agent to pass through state where a wrong move is fatal. This env (substitute one word) doesn’t. A version with multi-step edits, partial credit per-step, and a high failure rate would — that’s the natural follow-up.