01 · QUESTIONarXiv:2605.08741 · 2026-05-09

能不能把 inference harness 变成一次性训练 scaffold?

训练时运行 Plan–Solve / Draft–Verify,部署时只保留 direct model。

69.50

Math avg pass@8

53.33

HMMT25

3×10K

training subsets

9/10

个人兴趣

来源:Abstract;Tables 1-31 / 9
02 · CONTRACTarXiv:2605.08741 · 2026-05-09

Harness terminal 只给 teacher,student 仍直接作答

Direct student

原 problem → on-policy response

← KL

Harness teacher

执行 procedure → terminal context → score same response

来源:Method overview2 / 9
03 · PLAN–SOLVEarXiv:2605.08741 · 2026-05-09

Reference solution 先被压成 plan,再进入 teacher terminal

1Reference

planner 训练期可见

2Plan

抽象 strategy sketch

3Solve

不再看 reference

4Teacher

读 terminal

5Student

吸收 procedure

来源:Section 3.23 / 9
04 · DRAFT–VERIFYarXiv:2605.08741 · 2026-05-09

五个支持者和五个挑战者共同复核 draft

1Retrieve

5 neighbors

2Draft

初步 label

3Confirm

5 supporting views

4Challenge

5 opposing views

5Revise

整合后 terminal

来源:Section 3.14 / 9
05 · SETTINGarXiv:2605.08741 · 2026-05-09

Qwen3-8B,三组 10K 数据

任务StepsHarness
CAIL / LawBench300Draft–Verify
USPTO300Draft–Verify
DeepMath150Plan–Solve
来源:Section 4;Appendix5 / 9
06 · RESULTarXiv:2605.08741 · 2026-05-09

OPHSD 在三类任务上超过 OPSD

LawBench OPSD
64.25
LawBench OPHSD
69.51
Math OPSD
66.68
Math OPHSD
69.50
根据 main tables 重绘6 / 9
07 · CASEarXiv:2605.08741 · 2026-05-09

检索 harness 也会把证据带偏

Harness terminal

近邻全偏向 intentional injury

Trained direct student

保留 negligent homicide 的另一种解释

来源:LawBench case study7 / 9
08 · GAME CODEarXiv:2605.08741 · 2026-05-09

内化 procedure,不假装内化未来 execution result

可内化

plan→compile→run→inspect→repair 的步骤

必须执行

当前 build 的 error、render 与 gameplay state

本文对 game query-to-code 的边界8 / 9
09 · CODEarXiv:2605.08741 · 2026-05-09

Release 质量较好,reporting protocol 要分开

公开

trainer、harness、data、4-run CSV

不可混用

v1 best checkpoint vs repo fixed checkpoint

代码审计:20eaa0c8b0674fbf81aa780a8cf4e0d7c550d4ea9 / 9