We had guessed reward seeking might increase over the course of capabilities-focused RL training, bu...
OpenAI(@OpenAI)126 字 (约 1 分钟)
65
OpenAI提出新方法Contrastive SDF用于量化模型奖励寻求行为,与Apollo Research合作改进训练过程中的对齐检测。
入选理由:Contrastive SDF方法通过植入对比信念测量奖励寻求强度
精选推文#AI对齐#强化学习#OpenAI#模型评估英文