01 · QUESTIONarXiv:2609.25623 · 2026-09-22

Self-teacher 应该看完整解答,还是可行动的抽象?

固定 student 与 OPSD 训练,只改变 teacher-only context 的颗粒度。

29,434

aligned L1-L5 rows

+1.39

4B best intermediate vs L1

+1.57

8B best intermediate vs L1

9.5/10

个人兴趣

来源:Abstract;Tables 1-3;官方 dataset1 / 12
02 · FIVE LEVELSarXiv:2609.25623 · 2026-09-22

L1 到 L5 不是信息量单调减少那么简单

L1

完整 reference solution

L2

named strategy

L3

method-independent framing

L4

problem category

L5

final answer only

来源:Section 3;Appendix prompt templates2 / 12
03 · VISIBILITYarXiv:2609.25623 · 2026-09-22

Student 从未看到 L1-L5

Student

problem → 当前 policy rollout

→

Frozen self-teacher

problem + selected level + same student prefix

L = KL(π_teacher(·|x,z,y<t) || π_student(·|x,y<t))
来源:Method overview3 / 12
04 · DATAarXiv:2609.25623 · 2026-09-22

L2-L4 是串行 model compilation,不是人工标注

1L1

原 OpenThoughts solution

2L4

Qwen3.5-397B 生成 category

3L3

读取 L4 生成 framing

4L2

读取 L4+L3 生成 strategy

5JSONL

29,434 aligned rows

来源:Appendix compiler;官方 scripts4 / 12
05 · SCHEMAarXiv:2609.25623 · 2026-09-22

公开 row 能追到 sample 与五档 context

sample_id, source, problem, L1_full_solution, L2_strategy, L3_framing, L4_category, L5_answer_only

L5 有七个 proof-target overrides,记录在 code 中。

来源:data/privileged_contexts_l1_l5.jsonl5 / 12
06 · SETTINGarXiv:2609.25623 · 2026-09-22

论文主实验的训练合同

项设置
BackboneQwen3 1.7B / 4B / 8B
GPUs8×H200
Updates200;每 25 保存
RolloutT=1.1;top-p .95;top-k 20
Completion1,024
LoRA64 / 128
LR5e-6
来源:paper appendix;official launchers6 / 12
07 · LOSS CODEarXiv:2609.25623 · 2026-09-22

beta=0 是 forward KL;clip 在 vocabulary component 上

component_i = q_i · (log q_i − log p_i) clipped_i = clamp(component_i, −c, c) loss = Σ_i clipped_i

这不等于对完成的 token-level KL 做 scalar clipping。

代码审计:76ba7f2…;OPSD trainer fork7 / 12
08 · RESULTarXiv:2609.25623 · 2026-09-22

更短 context 在 4B/8B 胜过 L1,但 1.7B 没有

1.7B

L1 最好

4B

best intermediate +1.39

8B

best intermediate +1.57

根据 main result tables 重绘8 / 12
09 · SELECTIONarXiv:2609.25623 · 2026-09-22

Peak mean 可混合三个不同 checkpoint

1AIME24

取八个 checkpoint 中最高

2AIME25

独立再取最高

3HMMT25

独立再取最高

4Average

把三个 peak 平均

来源:metric definition;Appendix9 / 12
10 · CONFOUNDERSarXiv:2609.25623 · 2026-09-22

L1-vs-L3 同时改变了五件事

Semantics

抽象层级

Length

token 数

Answer

是否含 final

Wording

transition prompt

Compiler

生成误差

本文实验设计审计10 / 12
11 · GAME TRANSLATIONarXiv:2609.25623 · 2026-09-22

游戏实验要把 static reference 与 execution feedback 分开

Static

code、strategy、API/scene plan

Current build

compile、runtime、render、gameplay trace

Controls

length-match、wrong-task、answer-hidden

本文给 game query-to-code 的实验映射11 / 12
12 · ARTIFACTSarXiv:2609.25623 · 2026-09-22

数据和代码可审计,完整 run 仍无法复原

已公开

29,434 rows、compiler、trainer、manifests、CSV

未公开

checkpoints、logs、generations、judge annotations

代码审计:76ba7f2346fb7f2c965e4160989fbc701db6f1e612 / 12