0.31
off-path correct KL
0.08
on-path correct KL
−4.28
math think change
9.5/10
个人兴趣
在 harder tasks 上,dense PI-conditioned self-distillation 单独使用时常常降分。
off-path correct KL
on-path correct KL
math think change
个人兴趣
最大 log-prob advantage
显著更低
与 alternative 接近
最低但差距弱
无 reward 的 PI-conditioned dense loss
verifier reward + routing + local credit
| Axis | Setting |
|---|---|
| Backbone | Qwen3-8B;additional 32B |
| Rollouts | 8 |
| Epochs | 3 |
| Response | 16,384 |
| Hardware | 8×B200 |
标点
功能词
语气与格式
| Task / mode | Change |
|---|---|
| General QA think | −1.96 |
| Math think | −4.28 |
| Math instruct | −0.44 |
| Coding think | −0.30 |
| Agentic think | −3.51 |
near / divergent 两组都应低 penalty
near 也不能因为像 reference 获益
TeX、bib、style、figures
code archive 或 official repo
对 multi-solution game code,这篇是 mandatory diagnostic,而不是 implementation base。