Protecting people from harmful manipulation
TL;DR · AI 摘要
文章实际正文缺失,仅包含Google DeepMind官网的导航菜单与模型产品列表,未提供关于防范AI有害操纵的具体技术机制、研究进展或工程实践,信息密度极低,不具备工程师阅读价值。
核心要点
- 提供的文本仅为官网导航结构,缺失核心正文内容。
- 未涉及任何AI安全、模型对齐或防操纵的具体技术方案。
- 属于低信息密度的网页抓取残留,无法为工程实践提供参考。
Protecting People from Harmful Manipulation — Google DeepMind
Explore our next generation AI systems
Gemini

Specialized models

World models & embodied AI

Open models

Our latest AI breakthroughs and updates from the lab
Breakthroughs

Learn more
EvalsPublicationsResponsibility
Unlocking a new era of discovery with AI
Life sciences

Climate and sustainability

Our mission is to build AI responsibly to benefit humanity
Responsibility Ensuring AI safety through proactive security, even against evolving threatsNews Discover our latest AI breakthroughs, projects, and updatesCareers We’re looking for people who want to make a real, positive impact on the world
Education We work to make AI more accessible to the next generationOur National Partnerships for AI Working with governments worldwide to benefit people through frontier AIThe Podcast Join Professor Hannah Fry as she uncovers the extraordinary way AI is transforming our world
Models
Explore our next generation AI systems
Gemini

Specialized models

World models & embodied AI

Open models

Research
Our latest AI breakthroughs and updates from the lab
Breakthroughs

Learn more
EvalsPublicationsResponsibility
Science
Unlocking a new era of discovery with AI
Life sciences

Climate and sustainability

About
Our mission is to build AI responsibly to benefit humanity
Responsibility Ensuring AI safety through proactive security, even against evolving threatsNews Discover our latest AI breakthroughs, projects, and updatesCareers We’re looking for people who want to make a real, positive impact on the worldEducation We work to make AI more accessible to the next generationOur National Partnerships for AI Working with governments worldwide to benefit people through frontier AIThe Podcast Join Professor Hannah Fry as she uncovers the extraordinary way AI is transforming our world
Google AI Learn about all our AIGoogle DeepMind Explore the frontier of AIGoogle Labs Try our AI experimentsGoogle Research Explore our research
Products and apps
Gemini app Chat with GeminiGoogle AI Studio Build with our next-gen AI modelsGoogle Antigravity Our agentic development platform
March 26, 2026 Responsibility & Safety
Protecting people from harmful manipulation
Helen King
- [x]
Share
[](https://twitter.com/intent/tweet?url=https://deepmind.google/blog/protecting-people-from-harmful-manipulation/&text=Protecting%20people%20from%20harmful%20manipulation)[](https://www.facebook.com/sharer/sharer.php?u=https://deepmind.google/blog/protecting-people-from-harmful-manipulation/)[](https://www.linkedin.com/sharing/share-offsite/?url=https://deepmind.google/blog/protecting-people-from-harmful-manipulation/)[](mailto:?subject=Protecting%20people%20from%20harmful%20manipulation&body=https://deepmind.google/blog/protecting-people-from-harmful-manipulation/)Copied
As AI models get better at holding natural conversations, we must examine how these interactions affect people and society.
Building on a breadth of scientific research, today, we are releasing new findings on the potential for AI to be misused for harmful manipulation*, specifically, its ability to alter human thought and behavior in negative and deceptive ways. With this latest study, we have created the first empirically validated toolkit to measure this kind of AI manipulation in the real world, which we hope will help protect people and advance the field as a whole. We’re publicly releasing all materials necessary to run human participant studies using the same methodology. (_Note:_ _The behaviors observed during this study took place in a controlled lab setting, and do not necessarily predict real-world behaviors.)_
Why harmful manipulation matters
Consider two scenarios: One AI model gives you facts to make a well-informed healthcare decision that improves your well-being. Another AI model uses fear to pressure you to make an ill-informed decision that harms your health. The first educates and helps you; the second tricks and harms you.
These scenarios highlight the difference between two types of persuasion in human-AI interactions (also defined in earlier research):
- Beneficial (rational) persuasion: Using facts and evidence to help people make choices that align with their own interest
- Harmful manipulation: Exploiting emotional and cognitive vulnerabilities to trick people into making harmful choices
Our latest work helps us and the wider AI community better understand the risk of AI developing capabilities for harmful manipulation and build a scalable evaluation framework to measure this complex area. To do this effectively, we simulated misuse in high-stakes environments, explicitly prompting AI to try to negatively manipulate people's beliefs and behaviours on key topics.
Developing new evaluations for a complex challenge
Testing the outcomes of AI harmful manipulation
Testing for harmful manipulation is inherently difficult because it involves measuring subtle changes in how people think and act, varying heavily by topic, culture and context.
This is what motivated our latest research, which involved conducting nine studies involving over 10,000 participants across the UK, the US, and India. We focused on high-stakes areas such as finance, where we used simulated investment scenarios to test if AI could influence how people would behave in complex decision-making environments, and health, where we tracked if AI could influence which dietary supplements people preferred. Interestingly, the AI was least effective at harmfully manipulating participants on health-related topics.
Our findings show that success in one domain does not predict success in another, validating our targeted approach to testing for harmful manipulation in specific, high-stakes environments where AI could be misused.
How could AI manipulate?
In addition to tracking efficacy (whether the AI successfully changes minds), we also measured its propensity (how often it even _tries_ to use manipulative tactics). We tested propensity in two scenarios: when we explicitly told the model to be manipulative, and when we didn’t.
As detailed in our research, we counted manipulative tactics in experimental transcripts, confirming the AI models were most manipulative when explicitly instructed to be.
Our results also suggest that certain manipulative tactics may be more likely to result in harmful outcomes, though further research is required to understand these mechanisms in detail.
By measuring both efficacy and propensity, we can better understand how AI manipulation works and build more targeted mitigations.
Putting research into practice
As AI becomes a part of our everyday lives, we need to know it can’t be misused to harmfully manipulate people.
Beyond this latest study, we recently introduced an exploratory Harmful Manipulation Critical Capability Level (CCL) within our Frontier Safety Framework to help us track models with capabilities which could be misused to systematically change beliefs and behaviors in direct human-AI interactions in ways which could lead to severe harm.
These evaluations also serve as the foundation for how we test our models, including Gemini 3 Pro, for harmful manipulation. You can read more about this in this safety report. Like all our safety evaluations, this is an ongoing process. We will continue to refine our models and methodologies to keep pace with advancing AI.
Looking ahead
Understanding and mitigating harmful manipulation is a complex challenge. As model capabilities evolve, so too must our evaluation and mitigation techniques. For example, we’re currently exploring how to ethically evaluate the efficacy of harmful manipulation in even higher-stakes situations—like discussions involving deeply held personal beliefs—where users might be more susceptible to influence. Next, we will be expanding our research to investigate how audio, video, and image inputs as well as agentic capabilities, factor into AI manipulation.
We’ll continue to share findings and iterate based on feedback from the Frontier Model Forum and academic community. Our goal is to lead collective progress to prevent harmful manipulation, advancing AI models that prioritize safety and empower people.
_*Notes:__The scope of this particular research focuses exclusively on demonstrating general manipulation capabilities to help further the scientific study of evaluating harmful manipulation. This does not relate to testing safeguards around model outputs or manipulation in policy-violating and dangerous topics (e.g. terrorism and child safety) as this work is covered elsewhere and tested separately._
You can also read more about our harmful manipulation work in this interview with our researchers and in the Gemini 3 Pro Frontier Safety Report.
Acknowledgments
Canfer Akbulut, Rasmi Elasmar, Abhishek Roy, Anthony Payne, Priyanka Suresh, Lujain Ibrahim, Seliem El-Sayed, Charvi Rastogi, Ashyana Kachra, Will Hawkins, Kristian Lum, Laura Weidinger, William Isaac, Dawn Bloxwich, Lewis Ho, Eva Lu, Jenny Brennan, Mahmoud Hassan, Mark Graham
Follow us
[](https://x.com/googledeepmind)
[](https://www.instagram.com/googledeepmind)
[](https://www.youtube.com/@googledeepmind)
[](https://www.linkedin.com/company/googledeepmind/)
[](https://github.com/google-deepmind)
Sign up for updates on our latest innovations
I accept Google's Terms and Conditions and acknowledge that my information will be used in accordance with Google's Privacy Policy.
Sign up
Build AI responsibly to benefit humanity
Models
GeminiNano BananaGemini AudioGenieLyriaVeo
Research
EvalsBreakthroughsPublicationsResponsibility
Science
AlphaFoldAlphaGenomeWeatherNextAlphaEarth
Products
Gemini appGoogle AI StudioGoogle Antigravity
Learn more
AboutNewsCareersNational Partnerships for AIThe Podcast
[](https://www.google.com/?utm_source=ai.google&utm_medium=referral "Google")
Cookies management controls