SMRC-SD 解决一个动态问题:student 前几步一旦偏离 reference,完整成功轨迹就不再是当前 state 的合法指导。方法先匹配执行状态,匹配成功才做 privileged distillation;不匹配就 abstain。
Q1. Full-path guidance 为什么会在中途变成错答案?
reference 只证明某个 action 在它自己的 state 上有效;student 一旦按不同顺序完成 subgoal,继续照搬下一步就会误导。这不是 context 太长的问题,而是 state-reference mismatch。
SMRC-SD 把 successful trajectory 索引到多个 reference prefixes。每一 turn 都重建 student state signature,选择最新兼容位置;没有匹配就不给 dense SD loss。
Q2. Routing、contextualization 与 reward 如何分工?
| 模块 | 作用 | 未匹配时 |
|---|---|---|
| State matcher | 判断 reference prefix 是否兼容 | abstain |
| Router | 选择最新 compatible position | 不路由 |
| Context builder | 完整 path + current summary + candidate action | 不构造 |
| GRPO | 所有 turn 的 outcome optimization | 继续生效 |
| Self-distillation | matched turns 的 token signal | 关闭 |
因此 gain 可以拆成两部分:先决定“何时听 reference”,再决定“给 teacher 看怎样的 local context”。
Q3. ALFWorld 与 WebShop 怎样定义 state match?
ALFWorld adapter 把位置、inventory 与 subgoal 进度压成 deterministic execution-state signature;WebShop adapter 则记录 goal progress,并检查 required option 是否可用。matcher 同时确认 task identity、必要 progress fields 与 reference next action 的 admissibility,也就是这一步在当前环境中能否执行。
论文真实 case
ALFWorld student 已把 lettuce 3 放进 fridge。history matcher 因文本前缀相似,仍建议再开 fridge;structured state 识别到第一对象已经完成,路由到 second-object phase。另一个 case 里 agent 位于 fridge 1,表面相似的初始 history 会建议去 diningtable 2,location mismatch 让 SMRC-SD abstain。
Q4. setting、ablation 与 replay audit 说明什么?
ALFWorld 有 3,553 个训练 games,每个一条 verified expert walkthrough;WebShop-small 有 6,910 条 deterministic trace。Qwen3-1.7B / Qwen2.5-3B;每 update 16 tasks × 8 rollouts;actor LR 1e-6,GRPO clip 0.2,SDL coefficient 0.01,4×H800。主结果用固定 final checkpoint。
| Qwen3-1.7B · ALFWorld | Average@4 |
|---|---|
| FullPath-SD | 0.746 |
| Random same-count turns | 0.723 |
| Match-only routing | 0.836 |
| Routing + localized context | 0.865 |
| Dynamic context but distill unmatched | 0.695 |
same-count random control 只有 0.723,说明收益不是简单减少 distillation turns。35,712 个 archived identical turns 中,structured matcher 覆盖 20.2%,history matcher 覆盖 15.4%;前者保留了 98.8% 的 history matches。作者还从 structured matches 中抽取 781 条,把候选 action 接回 canonical suffix 后重放,781 条都成功。
Q5. 游戏代码的 state signature 应包含什么?
建议把 partial code state、compile state、scene graph 与 runtime fixture 合成 signature。至少包含 files/AST summary、已注册 InputMap、nodes/signals/resources、compile diagnostics、当前 hidden-test fixture、runtime snapshot 和已满足 obligations。
reference continuation 只有在 required node/API/state 已存在、下一操作仍 admissible 时才能进入 teacher context。否则 dense loss abstain,但 executable reward 继续训练。这个“会沉默的 teacher”比总能给建议的 teacher 更可靠。
Q6. 官方 repo 的复现缺口是什么?
matcher、routing、context builder、trainer、environment、launcher、probe 与 unit tests 都在;README 所称 included 的四个 data files 却不存在。缺失项包括 ALFWorld walkthrough prefixes/trajectories、WebShop oracle paths 和 fixed val128 manifest。
这会阻断 paper launcher 的 default path,除非自己重建并记录等价 pipeline。兴趣程度 10/10:它的 state-dependent abstention 是 game project 最应该直接借用的机制。
证据范围:本文阅读全文与附录,并分别检查 PDF 文本和逐页渲染;代码结论固定到文中注明的 commit。兴趣程度 10/10 只表示博客作者对该方向的个人兴趣,不是通用论文评分。
留言