Escalation
The axis
Section titled “The axis”Escalation — how many actions does the policy choose between at each decision?
A 64-tool action space and a 4-tool action space are not “the same task with more options.” The policy has to learn a 64-way categorical over capability descriptions in the first case and a 4-way categorical in the second. With opaque tool IDs (so the policy can’t lean on name priors), this difference dominates everything else.
The env in ~60 lines
Section titled “The env in ~60 lines”Each prompt lists N tools as tool_ID: capability description. The policy emits <tool>ID</tool> for whichever tool matches the natural-language query. The IDs are opaque (tool_00 … tool_63) so there is no naming prior.
def _shortlist(query, k=4): q = _embed(query) return sorted(TOOLS, key=lambda t: -len(q & _embed(t.capability)))[:k]
def stream(n, seed=0, *, shortlist=False): rng = random.Random(seed) for _ in range(n): q, ans_cap = rng.choice(QUERIES) tools = _shortlist(q, 4) if shortlist else TOOLS # 64 vs 4 listing = "\n".join(f"- {t.id}: {t.capability}" for t in tools) yield {"prompt": f"Tools:\n{listing}\nQuery: {q}\n…", "target": _CAP_TO_ID[ans_cap]}The leaky variant lists all 64 tools; the patched variant lists the top-4 by token-overlap retrieval.
First run — 64 tools
Section titled “First run — 64 tools”The leaky run’s true accuracy oscillates wildly (0.0 ↔ 1.0 across consecutive 25-step buckets) and never stabilizes. Verifier reward sits around 0.7 throughout. Looking at the rollouts at any given step, the 8 group members frequently disagree — at step 125, three of them voted tool_39 and two voted tool_19 for the same query.
The policy is failing to learn a confident mapping. With 64 opaque IDs and only 9 queries, every gradient step has to credit-assign across a much wider space of plausible categoricals than the policy can stabilize.
Re-run — 4-tool shortlist via token-overlap retrieval
Section titled “Re-run — 4-tool shortlist via token-overlap retrieval”<user>Tools: - tool_27: miscellaneous tool number 17 - tool_48: miscellaneous tool number 38 - tool_00: look up a fact in an encyclopedia - tool_07: evaluate an arithmetic expression Query: 17 * 38 Reply with only <tool>ID</tool> for the one tool whose capability matches the query.</user>
<assistant><tool>tool_07</tool></assistant>
The patched run is monotone. Reward ≈ 0.88, true accuracy ≈ 1.0, and at every logged step all eight rollouts emit the same tool ID. Same model, same recipe, same task — only the action-space size changed.
Retrieval is not optional. It changes the prior the policy is sampling against. The “same task” with a different action space converges in 25 steps instead of never.
The architectural read
Section titled “The architectural read”What looks like a tooling decision (retrieve-then-act) is actually an RL-env decision (reshape the action space). The two views point to the same fix, and the second framing tells you why.