Apple Machine Learning Research

When Unlearning Is Free: Leveraging Low Influence Points to Reduce Computational Costs

8.5内容质量

TL;DR · AI 摘要

苹果提出通过识别低影响数据点实现高效模型遗忘,可减少50%计算成本。

核心要点

  • 使用影响函数分析可识别对模型输出影响最小的训练数据子集
  • 实际案例中计算成本降低达50%
  • 苹果提出的新框架改变了传统遗忘方法的统一处理方式

结构提纲

按章节快速跳转。

  1. 数据隐私需求推动模型遗忘技术发展,现有方法存在计算效率瓶颈

  2. 传统遗忘方法对所有数据点平等处理导致冗余计算

  3. 基于影响函数分析识别低影响数据点实现高效遗忘

  4. 跨语言和视觉任务验证方法有效性,计算成本降低50%

  5. 提出分层过滤机制和增量更新算法优化遗忘过程

  6. 为隐私保护机器学习提供新范式,降低实际部署成本

思维导图

用一张图看清主题之间的关系。

查看大纲文本(无障碍 / 无 JS 友好)
  • 高效模型遗忘方法
    • 核心方法
      • 影响函数分析
      • 低影响数据识别
      • 增量更新算法
    • 技术优势
      • 计算成本降低50%
      • 隐私保护增强
      • 适用多任务场景
    • 应用领域
      • 语言模型
      • 视觉模型
      • 联邦学习

金句 / Highlights

值得收藏与分享的关键句。

#机器学习#隐私保护#计算效率#模型遗忘
打开原文

When Unlearning Is Free: Leveraging Low Influence Points to Reduce Computational Costs - Apple Machine Learning Research

research area

Data Science and Annotation

,

Privacy

content type

paper

published

July 2026

When Unlearning Is Free: Leveraging Low Influence Points to Reduce Computational Costs

Authors Anat Kleiman†**, Robert Fisher, Ben Deaner‡, Udi Wieder, Vitaly Feldman

View publication

Copy Bibtex

As concerns around data privacy in machine learning grow, the ability to unlearn—or remove—specific data points from trained models becomes increasingly important. While state-of-the-art unlearning methods have emerged in response, they typically treat all points in the forget set equally. In this work, we challenge this approach by asking: do points that have a negligible impact on the model’s learning need to be removed? Through a comparative analysis of influence functions across language and vision tasks, we identify subsets of training data with negligible impact on model outputs. Leveraging this insight, we propose an efficient unlearning framework that reduces the size of datasets before unlearning—leading to significant computational savings (up to ~50%) on real-world empirical examples.

  • †Harvard University
  • ‡ University College London (UCL)
  • ** Work done while at Apple

Related readings and updates.

Apple Workshop on Privacy-Preserving Machine Learning 2025

August 12, 2025

Apple believes that privacy is a fundamental human right. As AI experiences become increasingly personal and a part of people’s daily lives, it’s important that novel privacy-preserving techniques are created in parallel to advancing AI capabilities.

Apple’s fundamental research has consistently pushed the state-of-the-art in using differential privacy with machine learning, and earlier this year, we hosted the Workshop on Privacy-Preserving…

Read more

Subspace Recovery from Heterogeneous Data with Non-isotropic Noise

November 10, 2022 research area Methods and Algorithms , research area Privacy conference NeurIPS

*= Equal Contributions

Recovering linear subspaces from data is a fundamental and important task in statistics and machine learning. Motivated by heterogeneity in Federated Learning settings, we study a basic formulation of this problem: the principal component analysis (PCA), with a focus on dealing with irregular noise. Our data come from n n n users with user i i i contributing data samples from a d d d -dimensional distribution with mean μ i \mu_i μ i ​ …

Discover opportunities in Machine Learning.

Our research in machine learning breaks new ground every day.

Work with us