Skip to content

Six Axes of an RL Environment

A worklog on how the shape of an RL environment determines what a post-trained policy actually learns.

Reinforcement learning post-training advice usually sounds like “pick a good reward.” That advice is true and useless. The reward signal is one of six knobs on an RL environment, and the policy you get out is a function of all six. Most of the pathologies people attribute to “RL being fragile” are actually a specific knob turned the wrong way.

This is a worklog. Six chapters, six axes, one tiny env per chapter trained on the same model with the same recipe. For each axis we run the env once with the knob misconfigured, watch the policy break in a specific way, turn the knob, re-run, and chart the difference.

  • Model: Qwen2.5-1.5B-Instruct.
  • Algorithm: GRPO via TRL.
  • Sampling: vLLM with regex-constrained decoding so the verifier always has something parseable to grade.
  • Compute: one H100 on Modal per run, ~3-10 minutes per run, total cost under $5 for the full 12-run grid.
  • No env framework. “Env” here means a prompt template, a verifier function, and a regex. That’s it.
  1. Gameability — a verifier hole the policy will route around.
  2. Verifiability — a soft judge trains a policy that is close to the right answer.
  3. Shape — reward shaping is decoration unless the base policy can’t already get the terminal reward.
  4. Horizon — a horizon shorter than the optimal path is not “harder,” it is unwinnable.
  5. Escalation — the action-space cardinality sets the categorical the policy has to learn over.
  6. Reversibility — a cliff that doesn’t fire is not a cliff.

Each axis only matters in a narrow regime where the base policy fails often enough that the exploit pays, but not so often that there is no signal to learn from. Outside that window the axis is invisible. The chapters that don’t exhibit their pathology (shape and reversibility, at this task difficulty) make the point as sharply as the ones that do.

The take-away is not “tune these six knobs.” It is the calibration of the task against the model is itself the most important knob. The six axes are the search directions; task difficulty is the metric.