01 · NEGATIVE RESULTarXiv:2608.04794 · 2026-08-05

Teacher 在教 correctness,还是 reference-path similarity?

在 harder tasks 上,dense PI-conditioned self-distillation 单独使用时常常降分。

0.31

off-path correct KL

0.08

on-path correct KL

−4.28

math think change

9.5/10

个人兴趣

来源:Abstract;main diagnostics1 / 9
02 · PI BIASarXiv:2608.04794 · 2026-08-05

Alternative correct 解没有得到应有的 teacher advantage

1Reference

最大 log-prob advantage

2Alternative correct

显著更低

3Unrelated correct

与 alternative 接近

4Incorrect

最低但差距弱

来源:PI Bias Score2 / 9
03 · SCOPEarXiv:2608.04794 · 2026-08-05

论文否定的是 lone objective

已测试

无 reward 的 PI-conditioned dense loss

未测试

verifier reward + routing + local credit

来源:Experiment design3 / 9
04 · SETTINGarXiv:2608.04794 · 2026-08-05

QA、数学、代码、BFCL 都做压力测试

AxisSetting
BackboneQwen3-8B;additional 32B
Rollouts8
Epochs3
Response16,384
Hardware8×B200
来源:Appendix hyperparameters4 / 9
05 · TOKEN DIAGNOSISarXiv:2608.04794 · 2026-08-05

大量 divergence 落在低信息 token

Punctuation

标点

Stopwords

功能词

Style markers

语气与格式

来源:token-category analysis5 / 9
06 · RESULTarXiv:2608.04794 · 2026-08-05

Loss 稳定下降,accuracy 仍然下降

Task / modeChange
General QA think−1.96
Math think−4.28
Math instruct−0.44
Coding think−0.30
Agentic think−3.51
根据 main result table 重绘6 / 9
07 · GAME DIAGNOSTICarXiv:2608.04794 · 2026-08-05

必须做 correct × reference distance 的 2×2

Correct

near / divergent 两组都应低 penalty

Incorrect

near 也不能因为像 reference 获益

本文为 game code 设计的必要实验7 / 9
08 · ARTIFACTarXiv:2608.04794 · 2026-08-05

“Supplementary 有 code”与公开 source 不一致

公开 tar

TeX、bib、style、figures

未找到

code archive 或 official repo

2026-10-08 public artifact check8 / 9
09 · VERDICTarXiv:2608.04794 · 2026-08-05

不要只问“有没有提升”,要问“压掉了哪些正确解”

对 multi-solution game code,这篇是 mandatory diagnostic,而不是 implementation base。

兴趣程度 9.5/109 / 9