51.4%
software pass@1
64.8%
retained ICL gain
9.6×
fewer trials vs PPO
8/10
个人兴趣
从 archived history 分叉,只生成 next decision,不生成 future observation。
software pass@1
retained ICL gain
fewer trials vs PPO
个人兴趣
recorded observations
选择 branch point
生成 next decision
不生成 observation
只对新 tokens loss
| 方法 | Targets |
|---|---|
| OPSD | student tokens,teacher rescore |
| Experience Distillation | teacher-sampled next decision |
| SFT | recorded action |
tasks
turns/history
tokens/history
tokens total
显式 value:null reset
保留 falsy geometry update
4,096 → 128 sequences
768 → 64 steps
预训练 compile-run-debug decision
后续 on-policy execution calibration
method、large study、cases
model、tasks、verifier、code、histories