论文解读·

[2026-07-23] Sample-Efficient Learning from Agent Experience

Experience Distillation 从长达数万 token 的 agent trial history 分叉,只生成下一次完整 decision,不再调用 environment。它在 749 个 software tasks 上保留 64.8% 的 ICL 增益,但代码、task s…

页数
9
形式
交互图解
更新
2026.10.08
文章目录

Chenhui Gou, Haoqin Tu, Yunhao Fang, Jianfei Cai, Hamid Rezatofighi · arXiv:2607.21051v1 · 最早公开于 2026-07-23

9 页交互图解 · 先看训练链、真实 case 与复现边界大屏阅读 ↗

使用按钮或 ← → 翻页,F 进入或退出全屏按 Esc 退出全屏;独立打开后可返回文章

这篇处理的是昂贵 environment interaction:agent 已经积累了几十轮 trial history,如何在不再运行 environment 的前提下把经验写进参数?它从每个历史节点采样 teacher 的下一次 decision,到 action 为止就停。

51.4%software pass@1
64.8%retained ICL gain
9.6×fewer software trials vs PPO
8/10个人兴趣程度

Q1. 为什么长历史既有用又难以训练?

trial history 包含失败假设、environment feedback 与最终修复,in-context 很有效,但移除 context 后增益消失。SFT 只模仿 recorded action,难以综合多次尝试;重新让 teacher 在 environment 中 rollout 又会花费昂贵样本。

Experience Distillation 把每个 history prefix 当分叉点,只生成下一次完整 decision,loss 只落在新生成 token。

Q2. 它与 OPCD / OPSD 的训练分布有何不同?

方法训练 target 来自谁需要新 environment interaction
OPSD / OPCDstudent on-policy tokens,teacher 重评需要 student rollout
Experience Distillationteacher 从 archived history 采样 next decision不需要
SFTrecorded historical decision不需要

default objective 是 sampled teacher decision 的 next-token prediction,可视为 teacher-sampled forward KL。它适合样本昂贵的环境,但不保证 target 落在当前 student 会访问的 state。

Q3. 749 个 software tasks 与六个 games 怎样构造?

每个 software task 跑 8 至 12 个独立的 repeated-rollout processes,每个最多十次 trials;只保留多次尝试后产生 accepted commit 的 trajectory。749 条选中 history 平均 60.5 turns、82.4K tokens,总计 61.7M tokens。

TaleSuite 有六个 text-adventure tasks,47 trials、3,672 turns、502K tokens。模型是未披露的 in-house model,software 与 TaleSuite 还使用不同 base checkpoints。

MapStore2 真实 case

任务需要把 geometry 显式重置为 value: null,并修复 query pipeline 不要丢掉这个值为 null 的 update。第六次 trial 才把 trial 5 的 UI reset 与 pipeline repair 合并。teacher branches 学到这两个 accepted component,也重复了两个已被否定的 hypothesis;最终十个选中 patch 全部被外部 verifier 接受。

Q4. 结果、sample efficiency 与 branch packing 如何理解?

Software · pass@1Score
Zero-shot5.3%
Task-specific ICL76.4%
SFT8.0%
Experience Distillation51.4%

Experience Distillation 保留 64.8% 的 software ICL 增益;TaleSuite 保留 93.4%。它以至少 9.6× 更少 software trials 达到 51.4%,而 PPO 是 17.7%;TaleSuite 相对 GRPO 少 57.2× trials。

branch packing 把 4,096 examples 合成 128 sequences,training steps 从 768 降到 64。one-step target 在五个 games 上优于用 learned observation 的 two-step branch。

Q5. 游戏 code 的 expensive simulation 能怎样用?

如果一次 build/playtest 很贵,可以从已经记录的 compile-run-debug histories 分叉,蒸馏“下一次 repair decision”,不再启动 engine。这适合做 experience pretraining,随后仍要回到 on-policy execution fine-tuning 校准当前 student distribution。

每个 branch target 必须标记依据来自哪些 past observations,避免模型凭 teacher hallucinate 未发生的 future state。accepted commit 与 hidden-test log 也应一起发布。

Q6. 为什么它不是可直接落地的 baseline?

论文未公开 model checkpoint、preprocessing prompt、完整 optimizer schedule、749-task list、external verifier、code 或 histories。部分 local browser suites 没跑,acceptance 依赖未公开 external verifier;curated task 还按 sharp pass@10 contrast 选择,不能估计真实 prevalence。

兴趣程度 8/10。one-step branch 是很好的 efficiency idea,但你的主论文若研究 on-policy privileged context,应把它放在扩展实验,而不是用它替代当前 student execution。

证据范围:本文阅读全文与附录,并分别检查 PDF 文本和逐页渲染。论文未提供 official repository、749-task set、histories 或 verifier。兴趣程度 8/10 只表示博客作者对该方向的个人兴趣,不是通用论文评分。

留言

留言正在载入…

搜文章