What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents
固定模型、轨迹、reward 与更新次数,只改变 GRPO 的分组边界:这篇工作用四种 coding harness 和一个 held-out weak-ReAct,区分 source-configuration…
固定模型、轨迹、reward 与更新次数,只改变 GRPO 的分组边界:这篇工作用四种 coding harness 和一个 held-out weak-ReAct,区分 source-configuration…