[2026-02-04] Privileged Information Distillation for Language Models
π-Distill 面向多轮 agent:frontier teacher 不开放 hidden reasoning,只留下成功的 action/tool-call trace。论文研究怎样把这类 action-o…
π-Distill 面向多轮 agent:frontier teacher 不开放 hidden reasoning,只留下成功的 action/tool-call trace。论文研究怎样把这类 action-o…