Anthropic(@AnthropicAI)
Research we co-authored on subliminal learning—how LLMs can pass on traits like preferences or misal...
7.5内容质量

TL;DR · AI 摘要
Anthropic与合作者在《Nature》发表论文,揭示大语言模型可通过数据中的隐藏信号传递偏好或不对齐等特质。
核心要点
- LLMs能通过看似无关的数据(如无意义数字)传递特定偏好
- 该现象被称为“潜意识学习”,可能影响模型对齐与安全性
- 研究已在《Nature》正式发表,此前预印本于2025年7月发布
#大语言模型#AI安全#潜意识学习#模型对齐#Nature
打开原文Read the paper: https://t.co/b1BYwcW9dH" / X
Anthropic on X: "Research we co-authored on subliminal learning—how LLMs can pass on traits like preferences or misalignment through hidden signals in data—was published today in @Nature. Read the paper: https://t.co/b1BYwcW9dH" / X
Don’t miss what’s happening

Research we co-authored on subliminal learning—how LLMs can pass on traits like preferences or misalignment through hidden signals in data—was published today in
. Read the paper: https://nature.com/articles/s4158 6-026-10319-8…
Quote

Owain Evans
@OwainEvans_UK
·
Apr 15
Our paper on Subliminal Learning was just published in Nature! Last July we released our preprint. It showed that LLMs can transmit traits (e.g. liking owls) through data that is unrelated to that trait (numbers that appear meaningless). What’s new?
·
202
415
2.6K
1.2K
Read 202 replies