Lilian Weng(@lilianweng)

On-policy distillation provides an elegant way to use the teacher model as a process reward model to...

8.5内容质量
On-policy distillation provides an elegant way to use the teacher model as a process reward model to...

TL;DR · AI 摘要

Lilian Weng指出on-policy蒸馏能优雅地将教师模型作为过程奖励模型,提供稠密奖励并避免SFT式分布外冲击,提升数学推理与对话助手训练效果。

核心要点

  • On-policy蒸馏结合RL纠错能力与SFT奖励密度,优化训练稳定性。
  • 教师模型可充当过程奖励模型,避免rollout阶段的OOD shock问题。
  • 该方法在数学推理和内部聊天助手任务中表现优于传统方法。
#强化学习#模型蒸馏#AI训练
打开原文

Lilian Weng on X: "On-policy distillation provides an elegant way to use the teacher model as a process reward model to provide dense reward while preventing SFT style "OOD shock" during rollout." / X

Don’t miss what’s happening

Image 3
Image 3

Lilian Weng ![Image 4](http://x.com/lilianweng)

@lilianweng

On-policy distillation provides an elegant way to use the teacher model as a process reward model to provide dense reward while preventing SFT style "OOD shock" during rollout.

Quote

Image 5: Square profile picture
Image 5: Square profile picture

Thinking Machines

@thinkymachines

·

Oct 27, 2025

Our latest post explores on-policy distillation, a training approach that unites the error-correcting relevance of RL with the reward density of SFT. When training it for math reasoning and as an internal chat assistant, we find that on-policy distillation can outperform other

Image 6: Image
Image 6: Image

5:31 PM · Oct 27, 2025

·

142.3K Views

31

51

767

292

Read 31 replies