01 · QUESTIONarXiv:2602.04942 · 2026-02-04

Teacher 不开放 reasoning,只给 actions,能学吗?

π-Distill 把 successful tool-call trajectories 当训练期 privileged information。

15,885

Retail successful traces

44.1

Travel Planner best

17/21

需要非零 RL reward

8.5/10

个人兴趣

来源:Abstract;dataset appendix;ablations1 / 9
02 · OBSERVABILITYarXiv:2602.04942 · 2026-02-04

隐藏 reasoning,不等于没有可用监督

不可见

frontier teacher 的 internal chain-of-thought

可见

tool name、arguments、observations

可验证

environment 最终 success / failure

来源:Section 12 / 9
03 · METHODSarXiv:2602.04942 · 2026-02-04

两种方法共享 PI,但 teacher 是否更新不同

方法TeacherStudentRL
π-Distillprivileged、联合更新unprivileged、匹配 teacher联合
OPSD variantprivileged self-teacheron-policy rolloutKL penalty + reward
来源:Section 3;Algorithms 1-23 / 9
04 · DATAarXiv:2602.04942 · 2026-02-04

成功 trace 不是人工标注

Retail

15,885 traces;300/500 tasks 有 PI

Travel Planner

1,986 traces;45 training tasks

Rule

每题保留最短成功 trajectory

来源:Section 4;Appendix data details4 / 9
05 · SETTINGarXiv:2602.04942 · 2026-02-04

长 context、多轮 agent、三 seeds

项设置
BackbonesQwen3-4B/8B;R1-Distill-Llama-8B
Context25K
Temperature0.75
StepsRetail 600;Travel 400
Hardware2×H100
来源:Appendix hyperparameters5 / 9
06 · RESULTarXiv:2602.04942 · 2026-02-04

Travel Planner 的 PI transfer 明显强于 SFT + RL

Base
23.6
SFT+CoT+RL
32.3
OPSD
37.5
π-Distill α=1
44.1
根据 result table 重绘6 / 9
07 · REWARDarXiv:2602.04942 · 2026-02-04

Dense PI loss 不能代替 outcome reward

17 / 21ablation 需要非零 RL coefficient

teacher similarity 只能 shaping;最终 task outcome 仍需 environment verifier。

来源:ablation section7 / 9
08 · GAME CODEarXiv:2602.04942 · 2026-02-04

先教“何时 compile/run/inspect”,不要先复制 reference code

Tool-only

compile、run、screenshot 名称

Arguments

command、fixture、scene

Trace summary

failure → repair 的压缩经验

本文实验设计建议8 / 9
09 · VERDICTarXiv:2602.04942 · 2026-02-04

概念很重要,复现基线不完整

值得借

action-only PI + outcome reward

缺口

无官方 code、trace generation 未公开

兴趣程度 8.5/109 / 9