论文解读·

[2026-09-22] What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation

这篇系统比较 self-teacher 的五档 privileged context,从 full solution 到 answer-only。本文逐项核对 29,434 行公开数据、L2-L4 compiler、训练 launcher、loss 实现、checkpoint 选择与缺失 arti…

页数
12
形式
交互图解
更新
2026.10.08
文章目录

Kanghui Tian, Siyuan Liu, Tianxiang Jiang, et al. · arXiv:2609.25623v2 · 最早公开于 2026-09-22

12 页交互图解 · 先看训练链、真实 case 与复现边界大屏阅读 ↗

使用按钮或 ← → 翻页,F 进入或退出全屏按 Esc 退出全屏;独立打开后可返回文章

这篇不改 OPSD 的 student rollout,也不换 backbone;它只问一个更基础的问题:self-teacher 多看什么最有用?答案不是“越完整越好”,而是与 student scale 和任务匹配的抽象层级。

29,434公开对齐的 L1-L5 rows
1.394B intermediate vs L1
1.578B intermediate vs L1
9.5/10个人兴趣程度

Q1. 为什么“给 teacher 看完整答案”可能反而不好?

完整 solution 同时携带 final answer、某一种 reasoning path 和大量表面细节,student 未必能把这些信息转成自己可执行的下一步。过细 context 可能让 teacher 的分布贴近 reference wording;过粗 context 又可能不给方向。论文把这个问题改写为“哪一层 abstraction 仍对当前 student 可行动”。

它比较 L1 full solution、L2 named strategy、L3 method-independent framing、L4 category、L5 answer only。student 在五组实验中始终只看 problem。

Q2. 它与 OPSD、OPCD、PS-OPSD 的差别是什么?

工作Teacher-only context主要变量
OPSD完整 verified solution证明 self-teacher 可工作
OPCD经验总结或 system promptcontext 类型扩展
本文L1-L5 五档抽象context granularity
PS-OPSDstate/goal/constraints/transitions结构化 problem space

本文最有价值的是把 context 本身变成实验对象,但还没有控制所有混杂因素。比如 L5 只有 final answer,既更短又去掉 reasoning;L3 不仅语义更抽象,prompt transition 也不同。

Q3. 29,434 行数据从哪里来,训练代码怎样使用?

公开 JSONL 有 29,434 行,base row 来自 OPSD 的 OpenThoughts math 数据。每行包含 sample_id、source、problem 与 L1-L5。L1 是原 reference solution;L5 是 training answer field,代码对七个 proof-target case 做了 override。

谁生成 L2-L4

Qwen3.5-397B-A17B 离线生成三档摘要。compiler 不是平行独立标注:先用 problem + L1 生成 L4;L3 再读 L4;L2 再读 L4 和 L3。所以上游 category 错误会向下传播。semantic audit 也由同一个 model family 在另一组 prompt 下完成,不是 human agreement。

# 仓库数据编译顺序的简化表示
L4 = compile_category(problem, L1)
L3 = compile_framing(problem, L1, L4)
L2 = compile_strategy(problem, L1, L4, L3)

launcher 把所选 level 写进 OPSD collator 的 solution 槽

teacher_context = row[HINT_LEVEL]

完整训练 setting

项设置
BackboneQwen3-1.7B / 4B / 8B
Hardware8×H200
Updates200;每 25 steps 存 checkpoint
Learning rate5e-6
Rollouttemperature 1.1;top-p 0.95;top-k 20
Completion最多 1,024 tokens
LoRArank 64,alpha 128(主 launcher)
Teacherfixed base;关闭 student LoRA adapter
Lossbeta=0,即 teacher→student forward KL

audited commit 的 jsd_token_clip 对每个 vocabulary component 先截断再求和;1.7B/4B/8B launcher 的阈值分别是 0.05/0.05/0.06。它不是普通 scalar token-KL clip。

Q4. 主要结果、真实比较单位与 checkpoint 选择是什么?

4B 的 best intermediate context 比 L1 的 in-domain peak mean 高 1.39,8B 高 1.57;1.7B 则仍由 L1 最好。L2-L4 的 hint token 比 L1 少 16.9 到 37.1 倍。L5 只有 final answer,在 4B 和 8B 仍与 L1 相差不到 0.2。

“ID peak mean”到底怎样算

论文在 AIME24、AIME25、HMMT25 上分别从八个 checkpoint 选择各自最高分,再把三个 peak 平均。三个组成项可能来自三个不同 training steps。这个数适合比较“一个 run 的最好潜力”,不等于部署时有一个同时最优的单 checkpoint。

three-seed comparison 支持“intermediate abstraction 在 4B/8B 平均能胜过 L1”的 aggregate claim,但没有证明 L2、L3、L4 里某一个在所有规模都稳定最好。initial teacher-student KL 也不能正确排序最终表现。

Q5. 怎样把 L1-L5 改造成游戏 query-to-code 实验?

不要把五档 prompt 生搬硬套;应把 semantic content、长度、answer leakage、wording 和 recipient 分开控制。

建议条件游戏版本必须控制
L1完整 reference project/codereference token 与长度
L2implementation strategy不出现具体 node/class 名
L3scene/API plan 与 mechanic decomposition与 L2 长度 matched
L4engine + mechanic category测试是否只需粗粒度先验
L5final artifact summary / expected outcome不泄漏 code
Execution当前 student build 的 compile/runtime/visual feedback来自 student 自己的程序

最关键的新增条件是 execution feedback。它不是 L1-L5 的另一个静态 abstraction,而是 student 当前 build 执行后才产生的信息。full reference code 只能做 baseline,不能默认当最佳 teacher context。

一组最小但能回答科学问题的实验

可以先用两个 model scale、三个 training seeds 和五个条件:reward only、full reference code、implementation plan、当前 build 的 execution diagnostic、execution diagnostic + implicated code spans。student 每次都从原始 multimodal query 生成自己的程序;teacher 只重评这些 token,不把 reference code token 混进 rollout。

评测至少分开报告 compile success、runtime-clean success、hidden functional tests、gameplay goal、visual rubric 和“全部 mandatory checks 同时通过”的 end-to-end success。再加入 correct/reference-divergent 程序集,检查 dense loss 是否惩罚功能正确但内部结构不同的实现。主结果使用固定 final checkpoint,best-checkpoint 只作为补充。

Q6. 官方代码真的开放了什么,缺了什么?

仓库公开了对齐数据、compiler prompts/code、trainer fork、launch manifests 与 compact CSV,但没有 checkpoints、full logs、cluster launcher、per-item generations 或 judge annotations。所以可以复跑训练合同,不能从 release 独立重建论文的 item-level error taxonomy。

仓库是从 OPSD 改的,但 L1-L5 collator 同时更换 context label 和 transition wording。若你要发一篇严谨的 game paper,第一步应固定同一个 wrapper,只替换 context payload;再加 length-matched、wrong-task context 与 answer-hidden controls。

按什么顺序读这组论文

先读 OPSD 建立训练合同,再读本文、Privileged, but Biased、SMRC-SD 和 Visual-SDPO。这五篇分别覆盖 context granularity、single-solution bias、state routing 和 execution feedback。

第二轮再看 OPHSD、OPCD、π-Distill、PS-OPSD、Experience Distillation、ViCuR 与 LOPD。它们适合扩展 harness、experience、recoverable visual cue 和 latent context。

兴趣程度 9.5/10。它非常接近你的“已有 query→code 数据怎样转成 privileged context”问题,但它给的是变量表,不是最终方法。

证据范围:本文阅读全文与附录,并分别检查 PDF 文本和逐页渲染;代码结论固定到文中注明的 commit。兴趣程度 9.5/10 只表示博客作者对该方向的个人兴趣,不是通用论文评分。

留言

留言正在载入…

搜文章