01 · QUESTIONarXiv:2609.36380v1 · 2026-09-28

给一张图,coding agent 能否写出一个可执行的 3D 世界?

LEGO-Anything 把单图 3D reconstruction 写成 Image-to-Code:Agent 在 Blender 中写代码、执行、渲染、检查并继续修改。

208RGB inputs
104simulator scenes
53.4%最佳室内 Overall
RGB → CODE
根据论文 Figure 1 与 Table 2 重绘1 / 18
02 · PART 0what is a scene program?

程序保存对象、相机、材质和生成过程;执行它才得到场景

Rendered image

保留最终 pixels。改相机或查询遮挡对象时,信息已经不在图里。

Fixed 3D output

保留 mesh 或 point map。对象身份、生成逻辑和编辑接口可能缺失。

Scene program

显式创建 camera、geometry、materials、lights 与 hierarchy;可再次执行、编辑和查询。

论文 Introduction、Section 32 / 18
03 · LOOPprogram → scene → observation

Agent 交替修改程序与检查执行结果,最终 artifact 只是轨迹最后一点

1Interpret

理解参考图与相机。

2Plan

拆分房间和对象。

3Code

写 Blender Python。

4Execute

生成场景与预览。

5Revise

比较后继续修改。

τ = {(Pₜ, Sₜ, oₜ)}ₜ₌₁…T  final submission = PT
论文 Equation 1、Figure 2、Section 33 / 18
04 · BOUNDARYneighboring tasks

任务的独特组合:单图输入、完整场景、可执行程序、隐藏 3D 真值

方向典型输出本文额外要求
Object-level image-to-3D单对象 mesh / asset恢复房间、相机、多对象布局与外观
Single-image reconstruction固定 3D representation交付可执行、可编辑 Blender program
VIGA / SEIGvisual / Blender program以完整场景和隐藏 simulator GT 做定量评测
3DCodeBench / P3D-Bench文本或参数到对象代码从视觉证据反推全场景
WorldCoder-Bench文本到交互 3D world对参考图的 geometry 与 appearance 忠实度
论文 Table 1、Sections 2 与 Appendix A4 / 18
05 · DATAsimulator-grounded construction

专业资产先进入 simulator scene,再经过物理检查与人工复核

1Fab assets

只选允许 AI 使用的资产。

2LychSim

装配 geometry、layout、lights 与 materials。

3Checks

collision、stability 与 capture quality。

4Human review

接受、返工或继续标注。

5Frozen cases

RGB 公开;3D annotations 隐藏。

443registered assets
8environments
17themes
3matched difficulty tiers
根据论文 Figure 3、Table 2 与 Appendix D 重绘5 / 18
06 · VISIBILITYactor versus verifier

Agent 获得相机公共信息;真正用于评分的几何与对象 mask 保持隐藏

Agent-visible

  • single RGB image
  • dimensions 与 horizontal FOV
  • category taxonomy、output schema
  • Blender MCP、文件和代码工具
  • 自己的 scene 与 render
VERIFIER BOUNDARY

Verifier-only

  • reference geometry 与 depth
  • instance masks
  • object correspondences
  • scene transforms
  • metric feedback 与 benchmark scores
Section 4.1、Appendix D.1、F.1 与 I6 / 18
07 · ARTIFACTdeliverability first

一次提交必须留下三个可用产物;进程正常退出本身不算通过

scene.blend

可重新打开、非空,至少有一个 mesh 和 active camera。

scene.glb

格式正确、非空,作为可移植 geometry export。

final.png

可解码且非退化;只检查交付,不直接用它算 Appearance。

artifact invalid 或 unresolved headline-evaluation failure:V=R=A=S=0,并继续留在 attempted-case 分母里。超时留下合法产物时仍可评分。

Section 4.2、Appendix E.1 与 E.47 / 18
08 · GEOMETRYvisible-surface reconstruction

候选表面直接在参考相机坐标中比较,没有 alignment 或 rescaling

隐藏 depth 反投影出的 reference-visible points

候选 scene 从 active camera 栅格化出的 visible points

τ(g) = 0.05 · z(g)
Fₒ = 2PR / (P + R)
Rᵢ = meanₒ Fₒ

private instance mask 决定每个对象的 scored scope。错误相机、尺度、遮挡和多余表面都会直接进入 precision / recall。

Equation 2、Appendix E.2 与 Figure 188 / 18
09 · APPEARANCEfresh evaluator rerender

评测器重渲染 Blender scene,再把几何与外观合成 Overall

Appearance A

保留候选 camera、geometry、materials、lights 与 color settings;固定 engine、resolution 与 full frame。

RGB max error ≤ 30

Validity V

三个 artifacts 有效,且 headline evaluator 成功完成。

V ∈ {0,1}

Overall S

所有 attempted cases 等权平均;失败不会从分母消失。

S = mean[V(R+A)/2]
Equations 3-4、Appendix E.3-E.49 / 18
10 · MAIN RESULTmean ± std over three runs

Validity 接近饱和,几何与外观才真正拉开模型差距

Model + CodexIndoor VIndoor RIndoor AIndoor SOutdoor VOutdoor ROutdoor AOutdoor S
GPT-6-astra100.052.454.453.498.034.045.539.6
GPT-6-sol99.422.242.532.399.719.828.724.2
GPT-6-luna99.715.830.523.299.713.221.417.3
GPT-5.6-sol97.89.819.814.898.712.817.715.3
GPT-5.6-terra99.79.521.315.498.08.115.011.5
GPT-5.6-luna99.110.218.014.197.07.816.612.1

数值为百分比;标准差见原论文 Table 3。所有模型室外 Reconstruction 都低于室内。

根据原论文 Table 3 重绘10 / 18
11 · SCALINGcomplexity and reasoning effort

更多对象降低 fidelity;更多 reasoning 主要帮助 GPT-6 系列

三个 complexity tiers 的平均 Overall

Easy
24.6
Medium
21.3
Hard
20.5

六个 GPT 配置合并;Validity 仍约为 99.5%。

Office 42 cases:Low → XHigh

GPT-6-astra
32.3→61.8
GPT-6-sol
21.3→39.7
GPT-6-luna
14.4→21.2

GPT-5.6 模型没有稳定的单调提升。

Table 4、Figure 4、Appendix F.411 / 18
12 · TRAJECTORYbest checkpoint ≠ final checkpoint

后期编辑会破坏已经正确的相机、几何和材质

29.6%GPT-5.6-sol 的 updates 降低 Q
−3.2 ppfinal 比 best intermediate 更低

House bedroom:GPT-6-astra 从 step 2 的 Q=33.90% 跌到 step 4 的 4.35%;GPT-5.6-terra 从 36.93% 跌到 2.10%。

step 01 build room shell
step 02 BEST Q = 33.90%
step 03 revise camera / materials
step 04 FINAL Q = 4.35%

# final submission keeps
# the endpoint, not the best state
Figure 5、Appendix B.2、Figures 11-1212 / 18
13 · SELF-EVALUATION6 builders × 6 judges

模型能看出颜色和光照差异,却很难从 render 判断隐藏几何是否改善

45.8%self judge · Reconstruction
45.4%cross judge · Reconstruction
62.2%self judge · Appearance
63.9%cross judge · Appearance

360 checkpoint pairs × 6 judges = 2,160 次主要判断。self-judging 对 Reconstruction 没有稳定优势,几何方向的一致率接近或低于 50%。

Figure 6、Appendix F.3、Table 813 / 18
14 · LEGO-PLUGINthree targeted interventions

初始化、证据检查和版本回退分别处理三类轨迹失败

Enhanced Initialization

VGGT 估计 gravity-aligned frame、room shell 和 camera;提供 projected extent 与 relative depth,不声称恢复 metric scale。

Grounded Refinement

SAM 3 regions 与 Depth Anything V2 depth 提供对象范围、深度顺序和 luminance 等可测 residual。

Version Control

每次修改声明可编辑对象与验收条件;失败时删除新增实体并恢复 scene snapshot。

Agent 仍负责规划,Blender MCP 仍是通用 scene editor;plugin 只提供受限测量、验证与 transaction 操作。

Figure 7、Appendix G.1-G.514 / 18
15 · TRANSACTIONcandidate edit → commit or rollback

一次修改先声明边界,再执行 Blender edit,最后决定提交还是恢复

DECLARE编辑契约

editable / protected entities、allowed change、hard requirements。

SNAPSHOT保存状态

transforms、meshes、materials、hierarchy、camera 与 render settings。

VALIDATE检查证据

未声明改动、hard failure 或 missing evidence 都会阻止接受。

DECIDECommit / rollback

最多 8 个 transactions、2 次 rollbacks、每个候选 2 轮 correction。

场景变化会使旧 validation report 失效;transaction 未解决或 render 未检查时,runtime hook 阻止 Agent 结束。

Appendix G.4-G.615 / 18
16 · PLUGIN RESULT42-case Office subset

弱模型提升最大;最强 Astra 只增加 2.1%

GPT-5.6-luna
+62.7%
GPT-5.6-terra
+55.8%
GPT-5.6-sol
+55.3%
GPT-6-luna
+27.5%
GPT-6-sol
+12.1%
GPT-6-astra
+2.1%

Reception

14.7 → 32.5
相对于最佳三次 base run,plugin 后的布局与对象更接近参考。

Copy room

10.9 → 27.8
两个案例是 qualitative examples,不是总体均值。

Figures 8-9;相对增益;相同 prompt、预算、reasoning effort 与 evaluator16 / 18
17 · LEGO-WORLDfrozen scene → three readouts

同一类 scene representation 可以导出检测、分割和深度,精度仍落后专用模型

scene.blend
├─ object geometry
├─ semantic labels
├─ active camera
└─ camera-space depth

freeze scene
then query →

COCO boxes

30.14 AP
DINO: 59.88

LVIS masks

14.75 AP
SAM 3: 53.96

ETH3D depth

0.1554 AbsRel
DA3: 0.0783

Table 5、Appendix H;三个独立 100-image tracks,并非同一批图片17 / 18
18 · VERDICTinterest 9.3 / 10

可执行 scene program 已经能稳定交付;忠实的三维恢复仍是主要缺口

论文建立了什么

  • single image → Blender program 的完整任务
  • Validity、visible-surface geometry、Appearance 三层评测
  • 对非单调编辑与失败自评的轨迹诊断
  • 冻结场景支持多个 deterministic readouts

证据边界

  • benchmark 仍以 simulator renders 为主
  • 系统比较没有统一 inference budget 与 information access
  • plugin 只在 42 个 Office cases 上验证
  • 研究代码与 evaluator 尚未公开,无法独立复现

最准确的结论:当前 Agent 擅长交付一个可运行的 Blender 世界,但还不能稳定恢复图片背后的几何、相机和外观。

论文全文与附录;博客作者兴趣度只表示本人兴趣18 / 18