Perplexity(@perplexity_ai)
Our reward design combines correctness, preference, and efficiency. Preference only counts when the...
6.0内容质量

TL;DR · AI 摘要
Perplexity解释了其奖励设计机制,强调正确性优先于偏好和效率。
核心要点
- 奖励设计结合正确性、偏好和效率。
- 偏好仅在答案正确时计入评分。
- 避免模型优化为“听起来更好但错误”的答案。
#AI#机器学习#奖励设计
打开原文Preference only counts when the answer is correct.
This keeps the model from optimizing for better-sounding wrong answers. https://t.co/VbJ1M4o26w" / X

Our reward design combines correctness, preference, and efficiency. Preference only counts when the answer is correct. This keeps the model from optimizing for better-sounding wrong answers.