论文解读·

[2026-08-05] When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents

SMRC-SD 把成功 trajectory 当成 state-indexed resource:只有 student 当前执行状态与 reference prefix 匹配时才启用 dense distillation,否则 abstain,所有 turn 仍保留 GRPO。本文也核查了 rep…

页数
9
形式
交互图解
更新
2026.10.08
文章目录

Junzhuo Liu, Weiwei Li, Jun Ling, Peng Wang · arXiv:2608.05219v1 · 最早公开于 2026-08-05

9 页交互图解 · 先看训练链、真实 case 与复现边界大屏阅读 ↗

使用按钮或 ← → 翻页,F 进入或退出全屏按 Esc 退出全屏;独立打开后可返回文章

SMRC-SD 解决一个动态问题:student 前几步一旦偏离 reference,完整成功轨迹就不再是当前 state 的合法指导。方法先匹配执行状态,匹配成功才做 privileged distillation;不匹配就 abstain。

0.865ALFWorld Average@4
0.693WebShop success
20.2%structured matching coverage
10/10个人兴趣程度

Q1. Full-path guidance 为什么会在中途变成错答案?

reference 只证明某个 action 在它自己的 state 上有效;student 一旦按不同顺序完成 subgoal,继续照搬下一步就会误导。这不是 context 太长的问题,而是 state-reference mismatch。

SMRC-SD 把 successful trajectory 索引到多个 reference prefixes。每一 turn 都重建 student state signature,选择最新兼容位置;没有匹配就不给 dense SD loss。

Q2. Routing、contextualization 与 reward 如何分工?

模块作用未匹配时
State matcher判断 reference prefix 是否兼容abstain
Router选择最新 compatible position不路由
Context builder完整 path + current summary + candidate action不构造
GRPO所有 turn 的 outcome optimization继续生效
Self-distillationmatched turns 的 token signal关闭

因此 gain 可以拆成两部分:先决定“何时听 reference”,再决定“给 teacher 看怎样的 local context”。

Q3. ALFWorld 与 WebShop 怎样定义 state match?

ALFWorld adapter 把位置、inventory 与 subgoal 进度压成 deterministic execution-state signature;WebShop adapter 则记录 goal progress,并检查 required option 是否可用。matcher 同时确认 task identity、必要 progress fields 与 reference next action 的 admissibility,也就是这一步在当前环境中能否执行。

论文真实 case

ALFWorld student 已把 lettuce 3 放进 fridge。history matcher 因文本前缀相似,仍建议再开 fridge;structured state 识别到第一对象已经完成,路由到 second-object phase。另一个 case 里 agent 位于 fridge 1,表面相似的初始 history 会建议去 diningtable 2,location mismatch 让 SMRC-SD abstain。

Q4. setting、ablation 与 replay audit 说明什么?

ALFWorld 有 3,553 个训练 games,每个一条 verified expert walkthrough;WebShop-small 有 6,910 条 deterministic trace。Qwen3-1.7B / Qwen2.5-3B;每 update 16 tasks × 8 rollouts;actor LR 1e-6,GRPO clip 0.2,SDL coefficient 0.01,4×H800。主结果用固定 final checkpoint。

Qwen3-1.7B · ALFWorldAverage@4
FullPath-SD0.746
Random same-count turns0.723
Match-only routing0.836
Routing + localized context0.865
Dynamic context but distill unmatched0.695

same-count random control 只有 0.723,说明收益不是简单减少 distillation turns。35,712 个 archived identical turns 中,structured matcher 覆盖 20.2%,history matcher 覆盖 15.4%;前者保留了 98.8% 的 history matches。作者还从 structured matches 中抽取 781 条,把候选 action 接回 canonical suffix 后重放,781 条都成功。

Q5. 游戏代码的 state signature 应包含什么?

建议把 partial code state、compile state、scene graph 与 runtime fixture 合成 signature。至少包含 files/AST summary、已注册 InputMap、nodes/signals/resources、compile diagnostics、当前 hidden-test fixture、runtime snapshot 和已满足 obligations。

reference continuation 只有在 required node/API/state 已存在、下一操作仍 admissible 时才能进入 teacher context。否则 dense loss abstain,但 executable reward 继续训练。这个“会沉默的 teacher”比总能给建议的 teacher 更可靠。

Q6. 官方 repo 的复现缺口是什么?

matcher、routing、context builder、trainer、environment、launcher、probe 与 unit tests 都在;README 所称 included 的四个 data files 却不存在。缺失项包括 ALFWorld walkthrough prefixes/trajectories、WebShop oracle paths 和 fixed val128 manifest。

这会阻断 paper launcher 的 default path,除非自己重建并记录等价 pipeline。兴趣程度 10/10:它的 state-dependent abstention 是 game project 最应该直接借用的机制。

证据范围:本文阅读全文与附录,并分别检查 PDF 文本和逐页渲染;代码结论固定到文中注明的 commit。兴趣程度 10/10 只表示博客作者对该方向的个人兴趣,不是通用论文评分。

留言

留言正在载入…

搜文章