论文解读·

[2026-06-09] Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual Artifacts

Visual-SDPO 执行 student 代码,把 screenshot、rubric 与 runtime error 作为 teacher-only context,再把视觉 defect 定位回负责的 code statement。它与游戏 query-to-code 最接近,但论文没有公…

页数
9
形式
交互图解
更新
2026.10.08
文章目录

Haoyu Dong · arXiv:2606.10334v1 · 最早公开于 2026-06-09

9 页交互图解 · 先看训练链、真实 case 与复现边界大屏阅读 ↗

使用按钮或 ← → 翻页,F 进入或退出全屏按 Esc 退出全屏;独立打开后可返回文章

Visual-SDPO 与你的任务最接近:student 写 code,renderer 执行成 chart、web page 或 slide;teacher 看到真实 render 和错误,再给原代码 token 监督。它还尝试把 defect region 追到具体 code statement。

78.6ChartMimic overall
82.6Design2Code overall
60.7AeSlides
10/10个人兴趣程度

Q1. 为什么 reference code 不足以教视觉生成?

reference code 只展示一种实现,而 student 的视觉缺陷必须在它自己的 render 上观察。两个 HTML 可以结构完全不同但像素结果同样正确;反过来,语法正确的 code 也可能产生 overlap、clipping、alignment 或 contrast 问题。

Visual-SDPO 让 teacher 读取 screenshot、structured visual rubric 与 compile/runtime/browser traceback,监督仍落在 student 已经写出的 code tokens 上。

Q2. 它相对 OPSD 新增了哪两个环节?

环节Reference-code OPSDVisual-SDPO
Teacher context参考代码student render + rubric + error
Credit所有 code token 类似defect 对应 statement 加权
Outcomedistillationexecutable × visual quality 的 GRPO reward

它把“上下文是什么”和“哪些 token 应负责”拆开:render feedback 提供信息,Visual-Grounded Code Credit Weighting 负责局部 credit。

Q3. defect 怎样从图片追到代码?

系统先检测 defect region,再找到产生该区域的 code statement。Matplotlib 用 constructor hook、stack frame 与 artist bounding box;HTML 元素带 source metadata,再由 browser 读 rectangle;python-pptx 在创建 shape 时挂钩。instrumentation 失败时再用 VLM fallback。

statement responsibility 是它的 rendered regions 与 defect regions 的最大 IoU,再乘 binary severity。token weight 为 1 + (α−1)×responsibility。final objective 组合 weighted reverse KL 与 β×GRPO。

# 论文公式的直接展开
responsibility(stmt) = max IoU(rendered_region(stmt), defect_region) × severity
weight(token in stmt) = 1 + (alpha - 1) × responsibility(stmt)
reward = executable_success × visual_quality

Q4. 训练数据和结果有多强?

训练数据来自 Chart2Code-160K、WebCode2M + WebSight、AeSlides-7k;三类实验均使用 Qwen3-VL-8B-Instruct。

BenchmarkBaseReference-code OPSDGRPOVisual-SDPO
ChartMimic67.977.076.278.6
Design2Code72.178.680.082.6
AeSlides49.5—58.260.7

Visual-SDPO 在 ChartMimic 比 reference-code OPSD 高 1.6,在 Design2Code 高 4.0。作者还报告 rollout/render budget 约为 GRPO 的 29%。

但论文没有公开 learning rate、batch size、updates、rollout temperature、α、β、defect detector threshold 等关键 setting,无法从论文重建 exact run。

Q5. 怎样把它扩展到游戏 code?

把“defect region→statement”扩展为“runtime event / scene node / visual region→code span”。Godot 可以记录 scene tree、signal connection、stack trace、node path、frame buffer 与 gameplay event。每条 hidden test failure 都应指向相关的 files/functions;论文把被 failure evidence 指向的代码称为 implicated span。这样不必让完整 project 的所有 token 一起承担 KL。

建议把 credit 分三层:compile error 对应语法或类型 span;runtime mechanic failure 对应 signal、physics 或 state-transition code;visual mismatch 对应 node/property 和 asset use。每层保留独立 verifier。

Q6. 为什么兴趣 10/10,却不能直接照着复现?

截至核查日没有官方公开 repository,论文也缺少关键训练和 detector 细节。这意味着它提供的是最接近的 research blueprint,不是可执行 recipe。任何 reimplementation 都必须公开 instrumentation、source map、defect detector、α/β 和失败执行的处理。

它仍然是本项目第一优先级:当前 student build 的执行反馈比 full reference code 更少受单一实现偏差影响,而且 span weighting 给出清晰的 credit-assignment contribution。

证据范围:本文检查 11 页 PDF 的全文文字与逐页渲染;论文没有 appendix,也没有提供 official repository。代码可用性按 2026-10-08 的 arXiv 页面与公开仓库检索结果记录。兴趣程度 10/10 只表示博客作者对该方向的个人兴趣,不是通用论文评分。

留言

留言正在载入…

搜文章