[2026-06-09] Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual Artifacts
Visual-SDPO 执行 student 代码,把 screenshot、rubric 与 runtime error 作为 teacher-only context,再把视觉 defect 定位回负责的 co…
Visual-SDPO 执行 student 代码,把 screenshot、rubric 与 runtime error 作为 teacher-only context,再把视觉 defect 定位回负责的 co…
LOPD 检索成功经验并压成 continuous latent tokens,由 latent-conditioned frozen teacher 监督 student on-policy trajectory…