01 · BASELINEarXiv:2601.18734 · 2026-01-26

同一个模型,为什么能同时当 teacher 和 student?

OPSD 只改变训练期上下文:student 看 question,frozen self-teacher 多看 verified solution。

43.4

Qwen3-1.7B OPSD 平均分

1 rollout

每题一个 student trajectory

1,024 tokens

主实验 completion 上限

9/10

个人兴趣程度

来源:Abstract;Section 3;Table 11 / 9
02 · CONTRACTarXiv:2601.18734 · 2026-01-26

On-policy 监督必须落在 student 自己的 prefix 上

Student

question → 采样 y;部署时也只有 question

≠

Self-teacher

question + reference solution + 同一个 y<t;只负责打分

teacher 不替换 student trajectory,因此训练覆盖的是当前 policy 真正访问到的状态。

来源:Section 3.12 / 9
03 · UPDATEarXiv:2601.18734 · 2026-01-26

一次 update 的五步数据流

1Rollout

student 只看 question

2Rebuild

加入 reference context

3Rescore

teacher 评估相同 tokens

4KL

full-vocabulary forward KL

5Update

只更新 student LoRA

来源:Algorithm 1;官方代码 audited commit3 / 9
04 · LOSSarXiv:2601.18734 · 2026-01-26

代码里的 clipping 不是常见的 scalar KL clipping

component = q_i · (log q_i - log p_i) component = clamp(component, -c, c) loss = sum_vocab(component)

逐 vocabulary component 截断后再求和,可能破坏完整 KL 的非负性。复现必须写清。

来源:官方代码 ae7d251…;README warning4 / 9
05 · SETTINGarXiv:2601.18734 · 2026-01-26

主实验用一个短 rollout 对比更重的 GRPO

配置OPSDGRPO
rollouts / problem18
max output1,02416,384
updates100500
trainable paramsLoRALoRA
来源:Section 4;Appendix hyperparameters5 / 9
06 · RESULTarXiv:2601.18734 · 2026-01-26

小模型收益最大,8B 仍有增益

Qwen3-1.7B Base
37.1
Qwen3-1.7B OPSD
43.4
Qwen3-4B OPSD
63.6
Qwen3-8B OPSD
64.8
根据 Table 1 重绘6 / 9
07 · BOUNDARYarXiv:2601.18734 · 2026-01-26

Teacher preference 不是 executable correctness

数学 reference

常见题目只有一个标准终点,但推理路径仍可多样

游戏代码

scene tree、API 和控制流可以完全不同

必须补的证据

compile、hidden tests、gameplay 与 visual verifier

本文分析;论文未研究多实现程序7 / 9
08 · CODEarXiv:2601.18734 · 2026-01-26

官方仓库可跑,但版本与 license 都要记录

可复用

trainer、collator、launch 与 evaluation scripts

要声明

commit、template、fixed/EMA teacher、loss estimator、clip

代码审计:ae7d2519e94920c4eb6206c0c26de46d9c50abae8 / 9
09 · VERDICTarXiv:2601.18734 · 2026-01-26

把 OPSD 当骨架,不把 full solution 当结论

第一阶段

固定 student rollout 和 verifier,只扫 context family

论文贡献

证明 context、routing 与 span credit 怎样影响正确替代实现

兴趣程度 9/109 / 9