Skip to content

Escalation

Escalation — how many actions does the policy choose between at each decision?

A 64-tool action space and a 4-tool action space are not “the same task with more options.” The policy has to learn a 64-way categorical over capability descriptions in the first case and a 4-way categorical in the second. With opaque tool IDs (so the policy can’t lean on name priors), this difference dominates everything else.

Each prompt lists N tools as tool_ID: capability description. The policy emits <tool>ID</tool> for whichever tool matches the natural-language query. The IDs are opaque (tool_00 … tool_63) so there is no naming prior.

envs/escalation/env.py
def _shortlist(query, k=4):
q = _embed(query)
return sorted(TOOLS, key=lambda t: -len(q & _embed(t.capability)))[:k]
def stream(n, seed=0, *, shortlist=False):
rng = random.Random(seed)
for _ in range(n):
q, ans_cap = rng.choice(QUERIES)
tools = _shortlist(q, 4) if shortlist else TOOLS # 64 vs 4
listing = "\n".join(f"- {t.id}: {t.capability}" for t in tools)
yield {"prompt": f"Tools:\n{listing}\nQuery: {q}\n…", "target": _CAP_TO_ID[ans_cap]}

The leaky variant lists all 64 tools; the patched variant lists the top-4 by token-overlap retrieval.

modelQwen/Qwen2.5-1.5B-Instruct
steps500
batch8
H100 min5.0
cost$0.33
seed4
0.000.200.400.600.801.00step 0step 500
verifier rewardtrue accuracy

The leaky run’s true accuracy oscillates wildly (0.0 ↔ 1.0 across consecutive 25-step buckets) and never stabilizes. Verifier reward sits around 0.7 throughout. Looking at the rollouts at any given step, the 8 group members frequently disagree — at step 125, three of them voted tool_39 and two voted tool_19 for the same query.

The policy is failing to learn a confident mapping. With 64 opaque IDs and only 9 queries, every gradient step has to credit-assign across a much wider space of plausible categoricals than the policy can stabilize.

Re-run — 4-tool shortlist via token-overlap retrieval

Section titled “Re-run — 4-tool shortlist via token-overlap retrieval”
modelQwen/Qwen2.5-1.5B-Instruct
steps500
batch8
H100 min2.9
cost$0.19
seed4
0.000.200.400.600.801.00step 0step 500
verifier rewardtrue accuracy
0
reward: 1.000true score: 1.000advantage: 0.000Σ logprob: 0.00
prompt
<user>Tools:
- tool_27: miscellaneous tool number 17
- tool_48: miscellaneous tool number 38
- tool_00: look up a fact in an encyclopedia
- tool_07: evaluate an arithmetic expression
Query: 17 * 38
Reply with only <tool>ID</tool> for the one tool whose capability matches the query.</user>
completion
<assistant><tool>tool_07</tool></assistant>

The patched run is monotone. Reward ≈ 0.88, true accuracy ≈ 1.0, and at every logged step all eight rollouts emit the same tool ID. Same model, same recipe, same task — only the action-space size changed.

Retrieval is not optional. It changes the prior the policy is sampling against. The “same task” with a different action space converges in 25 steps instead of never.

What looks like a tooling decision (retrieve-then-act) is actually an RL-env decision (reshape the action space). The two views point to the same fix, and the second framing tells you why.