[2026-02-12] On-Policy Context Distillation for Language Models
OPCD 研究怎样把历史经验或优化过的 system prompt 蒸馏进模型参数。它保留 student 的 on-policy rollout,再让 context-conditioned teacher 对同…
OPCD 研究怎样把历史经验或优化过的 system prompt 蒸馏进模型参数。它保留 student 的 on-policy rollout,再让 context-conditioned teacher 对同…
Experience Distillation 从长达数万 token 的 agent trial history 分叉,只生成下一次完整 decision,不再调用 environment。它在 749 个 so…