论文解读·

[2026-02-04] Privileged Information Distillation for Language Models

π-Distill 面向多轮 agent:frontier teacher 不开放 hidden reasoning,只留下成功的 action/tool-call trace。论文研究怎样把这类 action-only privileged information 迁移到 inference 时…

页数
9
形式
交互图解
更新
2026.10.08
文章目录

Emiliano Penaloza, Dheeraj Vattikonda, Nicolas Gontier, et al. · arXiv:2602.04942v3 · 最早公开于 2026-02-04

9 页交互图解 · 先看训练链、真实 case 与复现边界大屏阅读 ↗

使用按钮或 ← → 翻页,F 进入或退出全屏按 Esc 退出全屏;独立打开后可返回文章

这篇把 privileged information 从数学答案推进到多轮 agent 的成功 action/tool-call trace。它的核心结论很适合游戏代码:dense privileged loss 可以提供形状更细的监督,但不能替代真正的 task reward。

15,885Retail successful traces
44.1Travel Planner,π-Distill
17/21ablation 需要非零 RL reward
8.5/10个人兴趣程度

Q1. 没有 teacher chain-of-thought,还能蒸馏什么?

可以蒸馏 successful action/tool-call trajectory,因为 environment 会记录每一步可观察行为。论文假设 frontier agent 的隐藏推理不可得,但成功调用过哪些工具、参数是什么、环境返回什么仍然可见。

训练目标不是逐字复原 hidden reasoning,而是让 inference policy 在没有 reference trajectory 时也更容易选到正确 action。

Q2. π-Distill 与 OPSD 分别解决哪一段?

方法Privileged teacherUnprivileged studentOutcome reward
π-Distill随 RL 一起学习同时匹配 teacherteacher 与 student 均可用
论文中的 OPSD同模型看成功 tracestudent on-policy rollout与 reverse-KL penalty 联合
SFT + RL先模仿 static trajectory再跑 RL有

π-Distill 的 teacher 也在训练,不要求一开始就能把 PI 用好;OPSD 更接近 fixed privileged teacher。两者共同测试 tool calls with arguments、tool names only 与 self-generated hints。

Q3. 数据与训练协议具体是什么?

所有 PI 都来自成功的 DeepSeek-chat-v3.1 run,作者为每个 task 保留最短成功轨迹。tau-Bench Retail 共 15,885 条成功 trace,500 个 training tasks 中有 300 个配 PI;Travel Planner 有 1,986 条 trace,覆盖 45 个 training tasks。

backbone 为 Qwen3-4B/8B 与 R1-Distill-Llama-8B。两张 H100,25K context,temperature 0.75,三 seeds;Retail 600 steps,Travel Planner 400 steps。OOD 测试还包括 tau-Bench Airline 与七个 GEM search environments。

一个 action-only PI 的可见边界

teacher trace 可以写“调用搜索工具、参数 A;读取结果;再调用 booking 工具、参数 B”,但不包含 teacher 的内部理由。student 在 rollout 时只看到当前 task 和已经发生的 environment history。训练期 teacher 额外看到完整 successful trace,对 student 当前 action 分布提供监督。

Q4. 结果为什么不能只看最高分?

Travel Planner · Qwen3-8BScore
Base23.6
SFT + CoT + RL32.3
OPSD37.5
π-Distill α=040.7
π-Distill α=0.541.1
π-Distill α=144.1

论文在 17/21 个 ablation 中发现非零 RL reward 重要。这直接限制了“只做 token-level privileged matching”的解释。初始 teacher-student KL 过大时,PI 即便内容丰富也可能难以传递。

表格对每个 seed 先取 peak 再平均,而 learning curves 是同一步数上平均 seeds;这两种统计不能混为一谈。

Q5. 对游戏代码训练最直接的启发是什么?

保留 compile、hidden tests、runtime gameplay 与 visual rubric 的 outcome reward,再把成功 build trace 当 shaping。可尝试三档 PI:只给 tool name(compile/run/screenshot)、给 tool arguments 与结果、给压缩后的 failure→repair trace。

如果游戏生成需要十几轮 tool call,action-only PI 比 full reference code 更贴近 agent harness。它教的是“在什么状态下调用什么工具”,而不是要求 student 复制一种 scene tree。

Q6. 这篇的证据边界在哪里?

没有官方公开代码,所有 PI 又来自一个 frontier model,因此复现依赖未公开的 trace generation。实验只到 8B;PI utility、长度、初始 KL 和 trace quality 没有被完全独立控制。

兴趣程度 8.5/10。它的主要价值是把 RL reward 的角色说清楚:privileged distillation 是 dense shaping,不是 verifier。要作为实现基线,还得借 OPSD/SMRC-SD 的公开训练代码。

证据范围:本文阅读全文与附录,并分别检查 PDF 文本和逐页渲染。论文未提供 official repository;代码可用性按 2026-10-08 的 arXiv 页面与公开仓库检索结果记录。兴趣程度 8.5/10 只表示博客作者对该方向的个人兴趣,不是通用论文评分。

留言

留言正在载入…

搜文章