Skip to content

Appendix — POMDP mapping

The six axes aren’t a taxonomy I invented — they’re the knobs you get when you write down a POMDP and ask which parts of the tuple your RL post-training recipe actually exposes to a designer.

A POMDP is (S, A, O, T, R, γ):

AxisPOMDP componentWhat you’re tuning
GameabilityRWhether the reward function admits high-reward states the designer would reject.
VerifiabilityRWhether the reward is a true scalar of correctness or a proxy with structured noise.
ShapeRWhether reward fires only at terminal states or along the way.
Horizonγ, episode lengthHow far credit assignment has to reach.
EscalationAThe cardinality and prior over the action set the policy must sample from.
ReversibilityTWhether the transition function admits a recovery path back to non-terminal high-value states.

O (observability) isn’t a chapter on its own because TRL post-training is single-turn — the policy sees the full prompt as one observation. If a chapter on partial observability earns its keep later (e.g., for a tool-use setting where prior tool output must be summarized), it lands here.