01 · QUESTIONarXiv:2607.21051 · 2026-07-23

不再跑 environment,怎样把长 trial history 写进参数?

从 archived history 分叉,只生成 next decision,不生成 future observation。

51.4%

software pass@1

64.8%

retained ICL gain

9.6×

fewer trials vs PPO

8/10

个人兴趣

来源:Abstract;Tables 1-41 / 9
02 · BRANCHarXiv:2607.21051 · 2026-07-23

One-step model-free target 停在 action 边界

1History

recorded observations

2Prefix

选择 branch point

3Teacher

生成 next decision

4Stop

不生成 observation

5Train

只对新 tokens loss

来源:Method section2 / 9
03 · DIFFERENCEarXiv:2607.21051 · 2026-07-23

它不是 OPSD

方法Targets
OPSDstudent tokens,teacher rescore
Experience Distillationteacher-sampled next decision
SFTrecorded action
来源:Method comparison3 / 9
04 · DATA SCALEarXiv:2607.21051 · 2026-07-23

Software history 平均 82.4K tokens

749

tasks

60.5

turns/history

82.4K

tokens/history

61.7M

tokens total

来源:data appendix4 / 9
05 · REAL CASEarXiv:2607.21051 · 2026-07-23

MapStore2 第六次 trial 才合并两个正确修复

UI

显式 value:null reset

Pipeline

保留 falsy geometry update

来源:MapStore2 case study5 / 9
06 · RESULTarXiv:2607.21051 · 2026-07-23

远强于 SFT,但仍低于保留完整 history 的 ICL

Zero-shot
5.3
SFT
8.0
Experience Distill
51.4
Task-specific ICL
76.4
根据 software result table 重绘6 / 9
07 · PACKINGarXiv:2607.21051 · 2026-07-23

Branch packing 显著降低 sequence 与 step 数

Examples

4,096 → 128 sequences

Training

768 → 64 steps

来源:efficiency appendix7 / 9
08 · GAME USEarXiv:2607.21051 · 2026-07-23

昂贵 simulation 的 archived repair decision

适合

预训练 compile-run-debug decision

仍需要

后续 on-policy execution calibration

本文 game experiment suggestion8 / 9
09 · VERDICTarXiv:2607.21051 · 2026-07-23

效率思路强,artifact release 弱

论文给出

method、large study、cases

未公开

model、tasks、verifier、code、histories

兴趣程度 8/109 / 9