[2026-10-01] Finetuning with Sampling: SFT Learns Better Than You Think
Sampling SFT 如何把 off-policy expert traces 改写成更接近 base model 的训练数据;逐步区分 information projection、理想 MH、公开 gree…
Sampling SFT 如何把 off-policy expert traces 改写成更接近 base model 的训练数据;逐步区分 information projection、理想 MH、公开 gree…