Lilian Weng(@lilianweng)
On-policy distillation provides an elegant way to use the teacher model as a process reward model to...
8.5内容质量

TL;DR · AI 摘要
Lilian Weng指出on-policy蒸馏能优雅地将教师模型作为过程奖励模型,提供稠密奖励并避免SFT式分布外冲击,提升数学推理与对话助手训练效果。
核心要点
- On-policy蒸馏结合RL纠错能力与SFT奖励密度,优化训练稳定性。
- 教师模型可充当过程奖励模型,避免rollout阶段的OOD shock问题。
- 该方法在数学推理和内部聊天助手任务中表现优于传统方法。
#强化学习#模型蒸馏#AI训练
打开原文Lilian Weng on X: "On-policy distillation provides an elegant way to use the teacher model as a process reward model to provide dense reward while preventing SFT style "OOD shock" during rollout." / X
Don’t miss what’s happening

Lilian Weng 
On-policy distillation provides an elegant way to use the teacher model as a process reward model to provide dense reward while preventing SFT style "OOD shock" during rollout.
Quote

Thinking Machines
@thinkymachines
·
Oct 27, 2025
Our latest post explores on-policy distillation, a training approach that unites the error-correcting relevance of RL with the reward density of SFT. When training it for math reasoning and as an internal chat assistant, we find that on-policy distillation can outperform other
·
31
51
767
292
Read 31 replies