Toward universal steering and monitoring of AI models · D. Beaglehole et al.
原文开头
Daniel Beaglehole1†, Adityanarayanan Radhakrishnan2,3†*, Enric Boix-Adserà4, Mikhail Belkin1,5* Artificial intelligence (AI) models contain much of human knowledge. Understanding the representation of this knowledge will lead to improvements in model capabilities and safeguards. Building on advances in feature learning, we developed an approach for extracting linear representations of semantic notions or concepts in AI models. We showed how these representations enabled model steering, through which we exposed vulnerabilities and improved model capabilities. We demonstrated that concept representations were transferable across languages and enabled multiconcept steering. Across hundreds of concepts, we found that larger models were more steerable and that steering improved model capabilities beyond prompting. …
摘自《科学》(Science)第391卷 第6787期 · 2026年2月19日,D. Beaglehole et al.。仅引用开头一小段供了解文章,版权归原刊所有,全文请阅读原刊。