01 · PAPER IN ONE PAGEAlibaba Token Hub · arXiv:2609.22000v2 · first posted 2026-09-18

让 Agent 一边操作软件,一边把它重新做出来;再用 reference 生成的隐藏测试验收

250五个平台各 50 个 task
58.06GPT-6 Astra overall
2.8%Prog 全部通过的 task
9.4 / 10博客作者个人兴趣程度

RecreationWorld 研究的是一条完整 hybrid loop:探索 reference GUI → 写代码 → 启动 candidate → 观察差异 → 继续修改。RecreationBench 用隐藏的 programmatic 与 visual assertions 评估最终行为。

来源:Abstract、Sections 1、4、501 / 15
02 · PART 0GUI operation × software construction

Hybrid 不是把 GUI agent 和 coding agent 前后拼接,而是让信息在两边反复流动

EXPLORE REFERENCE

点击、输入、打开菜单;发现窗口、状态与计算结果。

→
IMPLEMENT

写 source、build script 与 launch entry;选择自己的架构。

→
VERIFY CANDIDATE

启动成品,操作并观察错误,再决定去哪里复查或修改。

282.5每条 rollout 顶层 tool call 中位数
9.08每 100 calls 的 GUI↔code-edit switches
来源:Sections 1、2.2、3.1
03 · MOTIVATIONone task, three useful properties

Recreation 既迫使 Agent hybrid,又让成功条件可以执行和复验

Hybrid by necessity

只点 GUI 无法交付代码;只写代码又不知道 reference 真正怎样响应。

Verifiable by construction

reference 是运行时 oracle;同一组 actions 和 assertions 可以重放到任意 candidate。

Scalable experience

大量开源软件可转换成新 task,高分 rollout 可筛选成训练轨迹。

评测只约束 observable behavior,不要求 candidate 复制 reference 的语言、framework 或目录结构。
来源:Section 2.2
04 · OFFICIAL OVERVIEWbenchmark + training + analysis
RecreationWorld 官方概览
原图:官方 GitHub assets/overview.png(MIT),对应论文 Figure 1
05 · TASK CONTRACTsame capability, native delivery semantics
五平台任务的可见和隐藏边界
根据 Sections 2.1、4.1、4.2、4.5 与 Appendix C.6 重绘
06 · EXPLORE–IMPLEMENT–VERIFYthe agent chooses the interleaving
RecreationWorld 官方工作流
原图:官方 GitHub assets/workflow.png(MIT),对应论文 Figure 2;Agent 没有预设步骤顺序
07 · TASK PROVENANCE250 held-out tasks

五个平台各 50 个 task,但来源和可见边界并不完全相同

平台Reference 来源与筛选交付物Programmatic interface
Ubuntu开源应用,固定 upstream revision;排除不稳定 build / live servicesource + build/launchAT-SPI
macOS开源 AppKit / SwiftUI 等,也含 menu-bar appsource + build/launchAXUIElement
Windows开源 .NET / JVM 等,固定 revisionsource + build/launchUI Automation
Android无 Internet permission、核心状态在本机Gradle → APKUiAutomator
Web44 synthetic + 6 public-derived sitesReact scaffold → index.htmlDOM / ARIA

Desktop / Android 尽量 source-blind;Web 必然暴露 client code,因此改为保护 captured ground truth 与 tests。

来源:Sections 4.1、4.2;Appendix C.1-C.3
08 · TEST CONSTRUCTIONconstruction happens before evaluation

Expected outcome 不是模型猜的:它来自 reference 的真实运行,并经过回放与人工复核

1 · DISCOVER

从 source / page analysis 找功能入口与可测试状态。

2 · OPERATE

实际操作 reference,记录 action-conditioned outcome。

3 · GENERATE

生成 fixture、action sequence、Prog / VLM assertions。

4 · REPLAY

在 clean reference 上跑;不稳定或不通过的 case 淘汰。

5 · REVIEW + FREEZE

人工检查 fixture、actions 与 expected observations,随后冻结。

论文没有报告 reviewer 人数、inter-annotator agreement 或驳回率。可以说“human-validated”,不能扩写成“高一致性人工标注”。
来源:Section 4.3、Appendix C.5
09 · CONCRETE CASEfixture → actions → observations → verdict
Logbert 测试从 fixture 到 PASS FAIL
根据 Figure 8、Appendix C.5 与公开 task manifest 重绘
10 · EVALUATION CHANNELSexact state + rendered appearance

Prog 和 VLM 看的是互补证据;缺证据是失败,judge 自己报错才重跑

Prog

用 AT-SPI / AXUIElement / UI Automation / UiAutomator / DOM-ARIA 读取准确 text、widget state、navigation outcome 与 computed value。Build 或 launch failure 直接使 task 为 0。

VLM

在冻结 checkpoint 截取完整 application window,用 Qwen3.7-Plus、temperature 0 判断 layout、color、canvas、legend 等视觉 assertion。输出必须是逐项 JSON verdict。

SYSTEM_PROMPT: pass=true only when the assertion is clearly and fully satisfied

missing / incorrect / placeholder / visibly mismatched element → FAIL
transport or parsing error → rerun; do not silently score as model failure

仓库的 shared VLM judge 还区分 authored、verdict total 与 errors,避免“judge 没回答”伪装成“0 项通过”。

来源:Section 4.4;官方 src/recreation_bench/common_vlm.py
11 · AGGREGATIONassertions are not solved tasks

58.06 是两条 assertion 通道的五平台等权平均,不是“58% 的 task 全部完成”

ASSERTION

每个 Prog / VLM condition 得到 pass 或 fail。

→
APPLICATION

在 task 内汇总各通道通过率。

→
PLATFORM

50 个 application 做 macro average。

→
FIVE PLATFORMS

Ubuntu、macOS、Windows、Android、Web 等权。

Overall = (Prog + VLM) / 2。另报 Prog ≥ 90% 与 Prog = 100% 的 application coverage,后者更接近“整题功能完整”。
来源:Section 4.4、Table 3、Appendix B
12 · MAIN RESULTSbest average, rare complete fidelity
RecreationBench 主结果
根据原论文 Table 3 重绘;全部主榜为 direct MCP、单次 rollout
13 · HARNESS STUDYClaude Opus 4.8 · Windows · 50 tasks

让模型在 persistent REPL 里组合 GUI calls,质量近似不变,交互与上下文开销显著下降

−39.9%computer-use calls
−40.7%input tokens
−65.5%tool-result text
−54.1%estimated model cost

Quality

Prog 35.05→35.60;VLM 31.00→32.29。

Runtime

4.12→3.04 h/task,下降 26.1%;output tokens 增加 15.7%。

证据边界

每个 task/config 一次 rollout;完整配置一起变化,不是单变量因果实验。

效率数字来自两种配置都有完整 usage 的 39 对;质量分使用所有 evaluator-valid pairs。

来源:Section 5.2、Figure 9、Appendix C.7
14 · TRAINING TRANSFER35,000 selected trajectories

Recreation trajectories 与五项 OOD benchmark 的提升同向,但训练证据还缺少关键控制

TEACHER

Qwen3.8-Max 生成五平台 recreation trajectories。

→
SELECT

每个平台 7,000 条,组成 35,000 条 balanced SFT mixture。

→
TRAIN

训练两个 model initializations。

→
OOD EVAL

五项 coding / visual coding / hybrid CUA benchmark,最大 +17.9 pp。

看见的信号

两个 run 最后 checkpoint 都高于第一个 evaluated checkpoint;自查 render 更频繁。

没公开

selected trajectories、checkpoint、完整 hyperparameters。

没控制

multiple seeds 与 equal-data non-recreation baseline。

来源:Section 3、Conclusion;Tables/Figures 2、5、6
15 · TAKEAWAYSwhat is established, what remains open

最完整的是可执行链路;最需要补的是 VLM 校准、重复运行和训练消融

论文已经讲清

  • 一个五平台 task contract 连接 GUI 探索与 coding。
  • reference-grounded tests 经 clean replay 与人工复核后冻结。
  • Prog 与 VLM 互补,五平台等权聚合。
  • 强模型平均分提高,但完整 programmatic fidelity 仍罕见。

仍然不能推出

  • 有限 hidden suite 不能证明行为完全等价。
  • VLM judge 没有 human calibration / agreement。
  • 公开软件的 pretraining contamination 无法排除。
  • 单次 rollout 不足以说明小分差稳定,也不能把 harness 节省归因于单一组件。
最值得复用的研究模板:用 running reference 发现规格,用 hidden replay 验证成品,用 trajectory 记录模型怎样在 GUI 与 code 之间闭环。
来源:Sections 5-8、Appendices B-D、官方仓库