论文解读·

[2026-02-12] On-Policy Context Distillation for Language Models

OPCD 研究怎样把历史经验或优化过的 system prompt 蒸馏进模型参数。它保留 student 的 on-policy rollout,再让 context-conditioned teacher 对同一批 token 重新评分。

页数
9
形式
交互图解
更新
2026.10.08
文章目录

Tianzhu Ye, Li Dong, Xun Wu, Shaohan Huang, Furu Wei · arXiv:2602.12275v2 · 最早公开于 2026-02-12

9 页交互图解 · 先看训练链、真实 case 与复现边界大屏阅读 ↗

使用按钮或 ← → 翻页,F 进入或退出全屏按 Esc 退出全屏;独立打开后可返回文章

OPCD 把 privileged context 从“标准答案”扩展到历史经验和 system prompt。它最值得读的结果不是某个最高分,而是 raw trace 会伤害模型,经过提炼的经验才稳定有用。

79.7math,OPCD
53.9Sokoban,filtered experience
256teacher top-k vocabulary
8.5/10个人兴趣程度

Q1. OPCD 为什么不直接把历史轨迹做 SFT?

历史轨迹很长、包含失败尝试,而且强模型的解法未必适合小 student。OPCD 先让 student 走自己的路径,再让带经验的 teacher 评价这条路径,避免只训练在历史 agent 的状态分布上。

论文覆盖两类 context:一类是从以往 solution traces 总结出的 transferable experience;另一类是经过优化的 system prompt。两者都只在训练 teacher 侧出现。

Q2. 它相对 OPSD 新增了什么?

轴OPSDOPCD
Contextverified solution经验总结或 optimized system prompt
KL 方向teacher → student forward KLstudent → teacher reverse KL
词表full vocabulary可选 student top-256
任务数学 reasoning数学、text games、domain prompts

OPCD 证明“context”可以来自 agent 自己积累的历史,而不必是人工 reference。但它也把 experience extraction 本身变成新的模型依赖。

Q3. 同一个 student response 怎样被 teacher 重评?

代码先用 no-experience prompt rollout,再把同一 response token 接到 experience-conditioned prompt 后面。teacher 没有机会改写轨迹。

# audited opcd/ 实现的等价流程
response = actor.generate(task_without_experience)
student_logp = actor.score(task, response)
teacher_logp = frozen_ref.score(task + consolidated_experience, response)
loss = KL(student || teacher)  # reverse KL

text-game 经验由单独 learner prompt 从 interaction history 和 environment feedback 中生成,不是人工 gold annotation。main setting 累积 300 份 context,训练 50 updates。

Q4. 哪些实验最值得记住?

条件MathFrozen LakeSokoban
Base / 无额外经验75.1——
Raw traces70.5——
Summarized context before distillation77.4——
OPCD79.726.5—
Filtered experience + OPCD80.938.353.9

raw traces 把数学 validation 从 75.1 降到 70.5,而 summarized experience 加 OPCD 达到 79.7。在 Sokoban,固定 teacher/student 分工得到 53.9,持续更新 self-distillation 只有 18.8。context 质量和 teacher 稳定性都比“有没有更多历史”重要。

Q5. 游戏 query-to-code 可以怎样借它?

把经验拆成三种对照:完整 build log、筛掉无效尝试的 trace、跨任务总结出的 compact lesson。对于一批 game query→code 数据,可以按 mechanic、engine API、failure signature 或 scene pattern 建 experience bank,再检查 retrieved experience 是否真的适配当前项目状态。

不要让 rollout sampler 直接读取训练答案。student 仍应从原始 multimodal query 开始;experience 只给 teacher 或单独 verifier。这样测到的是能力内化,而不是 inference-time RAG。

Q6. 代码能复现到什么程度?

官方实现完整但工程很重,且 text-game experience 文件要经过 online extraction 才会生成。audited commit 位于 Microsoft LMOps 的 opcd/ 子目录,包含 veRL fork、数据预处理、online text-game environment 和 experience consolidation。

跨模型搬运 context 也有风险:论文观察到大模型经验会降低小模型表现。兴趣程度 8.5/10;建议把它用于“经验形态”对照,不要直接把成功日志整段塞给游戏 student。

证据范围:本文阅读全文与附录,并分别检查 PDF 文本和逐页渲染;代码结论固定到文中注明的 commit。兴趣程度 8.5/10 只表示博客作者对该方向的个人兴趣,不是通用论文评分。

留言

留言正在载入…

搜文章