理解参考图与相机。
拆分房间和对象。
写 Blender Python。
生成场景与预览。
比较后继续修改。
LEGO-Anything 把单图 3D reconstruction 写成 Image-to-Code:Agent 在 Blender 中写代码、执行、渲染、检查并继续修改。
保留最终 pixels。改相机或查询遮挡对象时,信息已经不在图里。
保留 mesh 或 point map。对象身份、生成逻辑和编辑接口可能缺失。
显式创建 camera、geometry、materials、lights 与 hierarchy;可再次执行、编辑和查询。
理解参考图与相机。
拆分房间和对象。
写 Blender Python。
生成场景与预览。
比较后继续修改。
| 方向 | 典型输出 | 本文额外要求 |
|---|---|---|
| Object-level image-to-3D | 单对象 mesh / asset | 恢复房间、相机、多对象布局与外观 |
| Single-image reconstruction | 固定 3D representation | 交付可执行、可编辑 Blender program |
| VIGA / SEIG | visual / Blender program | 以完整场景和隐藏 simulator GT 做定量评测 |
| 3DCodeBench / P3D-Bench | 文本或参数到对象代码 | 从视觉证据反推全场景 |
| WorldCoder-Bench | 文本到交互 3D world | 对参考图的 geometry 与 appearance 忠实度 |
只选允许 AI 使用的资产。
装配 geometry、layout、lights 与 materials。
collision、stability 与 capture quality。
接受、返工或继续标注。
RGB 公开;3D annotations 隐藏。
scene.blend可重新打开、非空,至少有一个 mesh 和 active camera。
scene.glb格式正确、非空,作为可移植 geometry export。
final.png可解码且非退化;只检查交付,不直接用它算 Appearance。
artifact invalid 或 unresolved headline-evaluation failure:V=R=A=S=0,并继续留在 attempted-case 分母里。超时留下合法产物时仍可评分。
隐藏 depth 反投影出的 reference-visible points
候选 scene 从 active camera 栅格化出的 visible points
private instance mask 决定每个对象的 scored scope。错误相机、尺度、遮挡和多余表面都会直接进入 precision / recall。
保留候选 camera、geometry、materials、lights 与 color settings;固定 engine、resolution 与 full frame。
RGB max error ≤ 30
三个 artifacts 有效,且 headline evaluator 成功完成。
V ∈ {0,1}
所有 attempted cases 等权平均;失败不会从分母消失。
| Model + Codex | Indoor V | Indoor R | Indoor A | Indoor S | Outdoor V | Outdoor R | Outdoor A | Outdoor S |
|---|---|---|---|---|---|---|---|---|
| GPT-6-astra | 100.0 | 52.4 | 54.4 | 53.4 | 98.0 | 34.0 | 45.5 | 39.6 |
| GPT-6-sol | 99.4 | 22.2 | 42.5 | 32.3 | 99.7 | 19.8 | 28.7 | 24.2 |
| GPT-6-luna | 99.7 | 15.8 | 30.5 | 23.2 | 99.7 | 13.2 | 21.4 | 17.3 |
| GPT-5.6-sol | 97.8 | 9.8 | 19.8 | 14.8 | 98.7 | 12.8 | 17.7 | 15.3 |
| GPT-5.6-terra | 99.7 | 9.5 | 21.3 | 15.4 | 98.0 | 8.1 | 15.0 | 11.5 |
| GPT-5.6-luna | 99.1 | 10.2 | 18.0 | 14.1 | 97.0 | 7.8 | 16.6 | 12.1 |
数值为百分比;标准差见原论文 Table 3。所有模型室外 Reconstruction 都低于室内。
六个 GPT 配置合并;Validity 仍约为 99.5%。
GPT-5.6 模型没有稳定的单调提升。
House bedroom:GPT-6-astra 从 step 2 的 Q=33.90% 跌到 step 4 的 4.35%;GPT-5.6-terra 从 36.93% 跌到 2.10%。
360 checkpoint pairs × 6 judges = 2,160 次主要判断。self-judging 对 Reconstruction 没有稳定优势,几何方向的一致率接近或低于 50%。
VGGT 估计 gravity-aligned frame、room shell 和 camera;提供 projected extent 与 relative depth,不声称恢复 metric scale。
SAM 3 regions 与 Depth Anything V2 depth 提供对象范围、深度顺序和 luminance 等可测 residual。
每次修改声明可编辑对象与验收条件;失败时删除新增实体并恢复 scene snapshot。
Agent 仍负责规划,Blender MCP 仍是通用 scene editor;plugin 只提供受限测量、验证与 transaction 操作。
editable / protected entities、allowed change、hard requirements。
transforms、meshes、materials、hierarchy、camera 与 render settings。
未声明改动、hard failure 或 missing evidence 都会阻止接受。
最多 8 个 transactions、2 次 rollbacks、每个候选 2 轮 correction。
场景变化会使旧 validation report 失效;transaction 未解决或 render 未检查时,runtime hook 阻止 Agent 结束。
14.7 → 32.5
相对于最佳三次 base run,plugin 后的布局与对象更接近参考。
10.9 → 27.8
两个案例是 qualitative examples,不是总体均值。
30.14 AP
DINO: 59.88
14.75 AP
SAM 3: 53.96
0.1554 AbsRel
DA3: 0.0783
最准确的结论:当前 Agent 擅长交付一个可运行的 Blender 世界,但还不能稳定恢复图片背后的几何、相机和外观。