Reversibility
The axis
Section titled “The axis”Reversibility — does the environment include a “cliff” the agent can fall off, and if so, can it climb back up?
The conventional worry: irrecoverable failures (a bad apply() that breaks the working copy, a wrong tool call that costs money) train risk-averse policies that under-explore. The fix: give the agent a revert() so mistakes are recoverable.
The env in ~40 lines
Section titled “The env in ~40 lines”Task. Transform a source sentence into a target sentence by emitting apply('OLD'->'NEW') calls. Under the leaky variant, an apply() whose OLD substring isn’t in the current state terminates the episode with a small penalty. Under the patched variant, failed applies are silently skipped and revert() undoes the last successful apply.
def reward_leaky(completion, record): state, terminated = _simulate(completion, source, allow_revert=False) if terminated and state == source: return -0.05 return 1.0 - _edit_distance(state, target) / len(target)First run — irrecoverable apply
Section titled “First run — irrecoverable apply”Re-run — apply + revert available
Section titled “Re-run — apply + revert available”<user>Source: 'the quick brown fox jumps'
Target: 'the slow brown fox jumps'
Emit a sequence of apply('OLD'->'NEW') calls (one per line) to transform source into target. End when done.</user><assistant>apply('OLD'->'NEW')</assistant>The cliff didn’t matter because the policy learned a one-call substitution before it ever stepped near the cliff. Reversibility only differentiates trajectories the policy actually attempts.
Where it does bite
Section titled “Where it does bite”For reversibility to matter, the task has to genuinely require the agent to pass through state where a wrong move is fatal. This env (substitute one word) doesn’t. A version with multi-step edits, partial credit per-step, and a high failure rate would — that’s the natural follow-up.