OPCD 把 privileged context 从“标准答案”扩展到历史经验和 system prompt。它最值得读的结果不是某个最高分,而是 raw trace 会伤害模型,经过提炼的经验才稳定有用。
Q1. OPCD 为什么不直接把历史轨迹做 SFT?
历史轨迹很长、包含失败尝试,而且强模型的解法未必适合小 student。OPCD 先让 student 走自己的路径,再让带经验的 teacher 评价这条路径,避免只训练在历史 agent 的状态分布上。
论文覆盖两类 context:一类是从以往 solution traces 总结出的 transferable experience;另一类是经过优化的 system prompt。两者都只在训练 teacher 侧出现。
Q2. 它相对 OPSD 新增了什么?
| 轴 | OPSD | OPCD |
|---|---|---|
| Context | verified solution | 经验总结或 optimized system prompt |
| KL 方向 | teacher → student forward KL | student → teacher reverse KL |
| 词表 | full vocabulary | 可选 student top-256 |
| 任务 | 数学 reasoning | 数学、text games、domain prompts |
OPCD 证明“context”可以来自 agent 自己积累的历史,而不必是人工 reference。但它也把 experience extraction 本身变成新的模型依赖。
Q3. 同一个 student response 怎样被 teacher 重评?
代码先用 no-experience prompt rollout,再把同一 response token 接到 experience-conditioned prompt 后面。teacher 没有机会改写轨迹。
# audited opcd/ 实现的等价流程
response = actor.generate(task_without_experience)
student_logp = actor.score(task, response)
teacher_logp = frozen_ref.score(task + consolidated_experience, response)
loss = KL(student || teacher) # reverse KLtext-game 经验由单独 learner prompt 从 interaction history 和 environment feedback 中生成,不是人工 gold annotation。main setting 累积 300 份 context,训练 50 updates。
Q4. 哪些实验最值得记住?
| 条件 | Math | Frozen Lake | Sokoban |
|---|---|---|---|
| Base / 无额外经验 | 75.1 | — | — |
| Raw traces | 70.5 | — | — |
| Summarized context before distillation | 77.4 | — | — |
| OPCD | 79.7 | 26.5 | — |
| Filtered experience + OPCD | 80.9 | 38.3 | 53.9 |
raw traces 把数学 validation 从 75.1 降到 70.5,而 summarized experience 加 OPCD 达到 79.7。在 Sokoban,固定 teacher/student 分工得到 53.9,持续更新 self-distillation 只有 18.8。context 质量和 teacher 稳定性都比“有没有更多历史”重要。
Q5. 游戏 query-to-code 可以怎样借它?
把经验拆成三种对照:完整 build log、筛掉无效尝试的 trace、跨任务总结出的 compact lesson。对于一批 game query→code 数据,可以按 mechanic、engine API、failure signature 或 scene pattern 建 experience bank,再检查 retrieved experience 是否真的适配当前项目状态。
不要让 rollout sampler 直接读取训练答案。student 仍应从原始 multimodal query 开始;experience 只给 teacher 或单独 verifier。这样测到的是能力内化,而不是 inference-time RAG。
Q6. 代码能复现到什么程度?
官方实现完整但工程很重,且 text-game experience 文件要经过 online extraction 才会生成。audited commit 位于 Microsoft LMOps 的 opcd/ 子目录,包含 veRL fork、数据预处理、online text-game environment 和 experience consolidation。
跨模型搬运 context 也有风险:论文观察到大模型经验会降低小模型表现。兴趣程度 8.5/10;建议把它用于“经验形态”对照,不要直接把成功日志整段塞给游戏 student。
证据范围:本文阅读全文与附录,并分别检查 PDF 文本和逐页渲染;代码结论固定到文中注明的 commit。兴趣程度 8.5/10 只表示博客作者对该方向的个人兴趣,不是通用论文评分。
留言