晨间信号Morning Signal
《科学》 第392卷 第6794期 · 2026年4月9日 · 中文解读

AI能“记住”不该记的数据,能否强制其遗忘?

AIs can ‘memorize’ data they shouldn’t. Can they be forced to forget? · P. Hall
约 9 分钟News在小程序里点播,20 到 60 分钟做好
这篇讲什么
研究人员开发出名为Hubble的开源工具,旨在帮助研究大型语言模型如何“记住”训练数据中的敏感信息,并探索使其“遗忘”的方法。
原文开头
New tool could help researchers probe how models “unlearn” sensitive training material PETER HALL To create a highly capable large language model (LLM), you need to train it on lots of data—books and articles and web pages. In theory, the model digests all this material in order to generate realistic yet brand-new text, but that doesn’t always pan out. Sometimes, LLMs spit out word-for-word copies of what they ingested, potentially violating copyright or exposing sensitive information such as credit card numbers and addresses. The problem is known as memorization—and it’s a big thorn in the side of artificial intelligence (AI) developers. …
摘自《科学》(Science)第392卷 第6794期 · 2026年4月9日,P. Hall。仅引用开头一小段供了解文章,版权归原刊所有,全文请阅读原刊。
晨间信号小程序码
微信扫码,在小程序里听完整版
不用登录先听一篇 · 或在微信搜索小程序 晨间信号
同期其他文章