69.50
Math avg pass@8
53.33
HMMT25
3×10K
training subsets
9/10
个人兴趣
训练时运行 Plan–Solve / Draft–Verify,部署时只保留 direct model。
Math avg pass@8
HMMT25
training subsets
个人兴趣
原 problem → on-policy response
执行 procedure → terminal context → score same response
planner 训练期可见
抽象 strategy sketch
不再看 reference
读 terminal
吸收 procedure
5 neighbors
初步 label
5 supporting views
5 opposing views
整合后 terminal
| 任务 | Steps | Harness |
|---|---|---|
| CAIL / LawBench | 300 | Draft–Verify |
| USPTO | 300 | Draft–Verify |
| DeepMath | 150 | Plan–Solve |
近邻全偏向 intentional injury
保留 negligent homicide 的另一种解释
plan→compile→run→inspect→repair 的步骤
当前 build 的 error、render 与 gameplay state
trainer、harness、data、4-run CSV
v1 best checkpoint vs repo fixed checkpoint