MULTI-HARNESS RL · PAPER DEEP DIVEarXiv:2609.04518 · 2026-09-03
跨 harness 比较 reward,真的会让 coding model 学到可迁移能力吗?

同一批 records、同一 SFT 起点、同一更新预算,只把 GRPO group 从 task × harness 改成 task。换到 unseen weak-ReAct 后,Cross 没有测出稳定优势。

9.7 / 10博客作者兴趣度:强控制实验,直接对应 Claude Code / Codex harness 研究
Held-out · avg@8Cross − Within = +0.25 pp,95% CI [−0.48,+1.02]。
标题:What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents
02 · QUESTIONexposure ≠ credit assignment

“多 harness 训练”至少混着两个不同干预

EXPOSURE

模型见过多种交互接口

Prompt、tool schema、observation、retry 与 context policy 都不同。

本文两条主线都固定为四种 harness exposure。

GROUPING

不同接口的 reward 是否直接比较

Within 在 task-harness pair 内标准化;Cross 把同一 task 的四种 harness 放进一个 GRPO group。

这才是本文唯一改变的主变量。

来源:论文 §1、§3。
03 · MOTIVATIONsame checkpoint, different interface

同一模型换评测 harness,平均 solve rate 相差 4.3 倍

AIDER
2.14%

六种 recipe 的列均值。

OPENHANDS
9.27%

同一批 checkpoint 的列均值。

HARNESS RANGE
7.13 pp

2.14% → 9.27%,4.3×。

RECIPE RANGE
0.91 pp

5.55% → 6.46%,1.16×。

来源:论文 Figure 2、Table 1;24,000 attempted evaluations。
04 · RELATED WORK同一术语下的不同 grouping boundary

ClawGym II 做 Within,HarnessX 更接近 Cross;其他系统把多个变量一起改变

工作多 harness 信号评测边界本文的判断
ClawGym IItask-harness pair 内归一化native + transfer对应 Within;完整训练 recipe 同时变化。
HarnessX跨 harness / model version poolingend-to-end对应 Cross;演化过程与 credit rule 混在一起。
POLARproduction harness 内 GRPO各自 native harness不能区分 source adaptation 与 portability。
OpenForgeRL三 harness SFT + RLClawEval held-out报告正向 transfer,但 exposure 与训练阶段同时改变。
Orchardfull-harness SFTmatched / mismatched 2×2显示很强的 interface lock-in。
来源:论文 §2、Appendix K。跨论文数字不是统一 leaderboard。
05 · CONTROLLED DESIGNsame data · same budget · different grouping

训练记录被冻结后,Within 与 Cross 只在 advantage 分组处岔开

Controlled multi-harness experiment
根据论文 Figure 1 重绘。原论文为 arXiv non-exclusive distribution license,未复用原图像素。
06 · GROUPING RULEone numerator, two baselines

同一条 binary reward,因为比较对象不同,会得到不同 advantage

AWithin(x,h,i) = [r(x,h,i) − μ(x,h)] / [σ(x,h)+ε]同一 task、同一 harness 的 rollouts 相互比较,harness-level offset 被消掉。
ACross(x,h,i) = [r(x,h,i) − μ(x)] / [σ(x)+ε]同一 task 的四种 harness 全部入组,强 harness 的成功先验会进入 advantage。
来源:论文 Equations 1–2、8。
07 · WORKED EXAMPLE解释性例子,不是公开 episode log

假设同题四种 harness 的 reward 是 0、0、1、1

WITHIN

先按 harness 分开

每套 harness 只和自己的重复 rollout 比。某套 harness 若全成功或全失败,该组方差为 0,所有 trace 的 advantage 也为 0。

实际 manifest 中 55.1% 的 Within groups 是这种 all-pass / all-fail group。

CROSS

四套 harness 直接竞争

同题里成功的 harness 得正 advantage,失败的得负 advantage。183 个入选任务的 pooled group 全都有非零方差。

这会扩大 gradient coverage,也把 harness 成功率写进 credit。

来源:论文 §3、Appendix E、Table 7。Reward vector 仅用于解释机制。
08 · DATA LEDGERSWE-Gym training · SWE-bench Verified evaluation

梯度只来自 183 个“harness 会产生分歧”的任务

SFT
186

tasks,8,841 records。

RL MANIFEST
183

tasks,5,543 episodes。

RECORDS
81,216

四条 RL arms 完全一致。

UPDATES
81,200

每条 arm 实际完成数相同。

来源:论文 Appendix D、G,Tables 7–8。训练与评测 benchmark family 分离。
09 · JUDGEwhat the agent sees vs what scores it

Agent 交 patch;隐藏测试决定 binary reward;基础设施故障不算模型失败

VISIBLE
GitHub issuebase repositoryharness toolstool outputs
HIDDEN
VERDICT
apply patchrun testsparse logresolved 0 / 1
来源:论文 §3、Appendix D/L;SWE-bench 4.1.0 grading.py。
10 · REAL CASEastropy__astropy-12907

嵌套 CompoundModel:两条新测试要修好,13 条原有测试不能退化

SWE-bench Verified real case
来源:SWE-bench Verified official dataset。论文没有公开本 case 的逐条 harness trajectory。
11 · SOURCE RESULTSsix recipes × four evaluation harnesses

列从左到右明显变深,行与行之间只小幅移动

Source harness result matrix
根据论文 Figure 2 / Table 1 重绘。Cross − SFT = +0.77 pp,95% CI [+0.03,+1.52]。
12 · PORTABILITYweak-ReAct excluded from training

六条 recipe 的 spread 比实验最小可检测差异还小

Held-out weak-ReAct results
根据论文 Figure 3a、Tables 2–4 重绘。
13 · STATISTICAL READINGabsence of detection ≠ exact equality

实验排除了大效应,没有证明 Cross 与 Within 数学等价

AVG@4
1.19–1.42 pp

六个 held-out contrast 的 80% power 最小可检差异。

AVG@8
0.88–1.14 pp

尝试数翻倍,precision 提高,点估计也朝 0 收缩。

CLAIM

no detectable portability

小于约 1 pp 的真实差异仍可能存在,论文不能把它排除。

来源:论文 §5、Appendix D/L,Tables 3–4。
14 · ROBUSTNESSsign flips across seeds

Cross − Within 的 seed 波动比两者平均差异更大

THREE SEEDS · AVG@4

+0.62 / +0.05 / −0.20 pp

合并差值 +0.16 pp,95% CI [−0.41,+0.72]。Within 与 Cross 自己的 seed range 分别为 0.45 与 0.42 pp。

RE-COLLECTED CROSS

后半程改为 on-policy

相对 offline Cross 为 +0.13 [−0.85,+1.10];相对 Within 为 +0.75 [−0.10,+1.65]。区间仍跨 0。

来源:论文 §5、Appendix D,Tables 2–3。
15 · CREDIT DIAGNOSTICcan a classifier recover the harness?

Cross 确实把 harness 成功先验写进 advantage

OUT-OF-FOLD TEST

只看 advantage 猜 harness

183 tasks hashed into five folds;每条 arm 和自己的 within-task shuffled-label null 比。

Within:+0.02 pp [−0.38,+0.46];Cross:+4.48 pp [+3.22,+5.83]。

FAILED RESIDUALIZATION

Baseline population 选错了

Baseline 从 1,008 个多数全失败的任务估计,训练却只用 183 个 disagreement tasks。它只去掉 9% offset,残留 identity +4.54 pp。

来源:论文 §6、Appendix E/I,Table 9。
16 · BEHAVIOR DIAGNOSTICinterface dominates repertoire

换 harness 会重排行为;换 credit rule 几乎不动

Harness identity and action distribution diagnostics
根据论文 Figures 3b–4、Tables 9–10 重绘。Aider 因 plain-chat transcript 未纳入 action labels。
17 · LIMITSwhat still remains uncertain

结论窄而可信;复现材料和外部有效性仍有缺口

SCOPE一套 model family 与预算Qwen3-8B、Python tasks、四个 source harness、一个 weak-ReAct held-out。
POWER小于约 1 pp 的差异未排除绝对 solve rate 低,source 主要是 avg@2。
ARTIFACTarXiv 没有公开 artifact URL论文描述 JSON、grader 和 weak-ReAct release,但本次无法独立定位。
SELECTION只训练 disagreement tasks这是 Cross 产生梯度的必要条件,也限定了目标分布。
LABELER11 / 160 turns 误标8 个错误出现在 SWE-agent;action JSD 仍相差两个数量级。
INFRA AUDIT五类静默故障被发现并修复包括 container 污染、20% episode 丢失、grader crash 和 cache 截断。
来源:论文 §9、Appendix C/H/J/L;artifact URL 检查于 2026-09-29。
18 · TAKEAWAYSfor Claude Code / Codex harness research

想证明模型学到了可迁移能力,必须把接口换掉再测

WHAT THE PAPER SHOWS

Cross credit changes the signal

Harness identity 可从 advantage 恢复;source harness 上有小幅收益。

但 held-out score、seed 复核与 action composition 都没有显示额外 portability。

WHAT TO DO NEXT

把 harness 当成一等实验变量

报告训练 exposure、grouping boundary、source delta、unseen-harness delta、grader 与版本。

下一步应做多种 held-out interface 和更细粒度 credit assignment。

核心结论1 / 18