64.32
4B
65.40
8B
0.5%
train PI invocation
8/10
个人兴趣
PS-OPSD 用 state、goal、constraints 与 transitions 替代完整推理。
4B
8B
train PI invocation
个人兴趣
已知状态
成功条件
规则边界
operator + preconditions + result
student only
current policy
teacher only
same tokens
full vocabulary
problem dynamics 与 selected path
instantiated final answer
| 4B condition | Avg |
|---|---|
| Wrong-task matched length | 61.76 |
| Corrupted order | 62.19 |
| Flattened fields | 63.61 |
| PS-OPSD | 64.32 |
| Scale | Base | OPSD | PS-OPSD |
|---|---|---|---|
| 1.7B | 36.67 | 41.08 | 43.12 |
| 4B | 61.76 | 61.70 | 64.32 |
| 8B | 62.31 | 63.58 | 65.40 |
initial
success
constraints
transitions
method、results、controls
repo、manifest、exact hyperparameters
来自一条 reference path 的 transition 仍可能不适配 student 当前 build。