29,434
aligned L1-L5 rows
+1.39
4B best intermediate vs L1
+1.57
8B best intermediate vs L1
9.5/10
个人兴趣
固定 student 与 OPSD 训练,只改变 teacher-only context 的颗粒度。
aligned L1-L5 rows
4B best intermediate vs L1
8B best intermediate vs L1
个人兴趣
完整 reference solution
named strategy
method-independent framing
problem category
final answer only
problem → 当前 policy rollout
problem + selected level + same student prefix
原 OpenThoughts solution
Qwen3.5-397B 生成 category
读取 L4 生成 framing
读取 L4+L3 生成 strategy
29,434 aligned rows
L5 有七个 proof-target overrides,记录在 code 中。
| 项 | 设置 |
|---|---|
| Backbone | Qwen3 1.7B / 4B / 8B |
| GPUs | 8×H200 |
| Updates | 200;每 25 保存 |
| Rollout | T=1.1;top-p .95;top-k 20 |
| Completion | 1,024 |
| LoRA | 64 / 128 |
| LR | 5e-6 |
这不等于对完成的 token-level KL 做 scalar clipping。
L1 最好
best intermediate +1.39
best intermediate +1.57
取八个 checkpoint 中最高
独立再取最高
独立再取最高
把三个 peak 平均
抽象层级
token 数
是否含 final
transition prompt
生成误差
code、strategy、API/scene plan
compile、runtime、render、gameplay trace
length-match、wrong-task、answer-hidden
29,434 rows、compiler、trainer、manifests、CSV
checkpoints、logs、generations、judge annotations