[2026-02-12] On-Policy Context Distillation for Language Models
OPCD 研究怎样把历史经验或优化过的 system prompt 蒸馏进模型参数。它保留 student 的 on-policy rollout,再让 context-conditioned teacher 对同…
OPCD 研究怎样把历史经验或优化过的 system prompt 蒸馏进模型参数。它保留 student 的 on-policy rollout,再让 context-conditioned teacher 对同…