Performance of a large language model on the reasoning tasks of a physician · P. G. Brodeur et al.
原文开头
Peter G. Brodeur1†, Thomas A. Buckley2†, Zahir Kanjee1, Ethan Goh3,4, Evelyn Bin Ling5, Priyank Jain6, Stephanie Cabral1,7, Raja-Elie Abdulnour8, Adrian D. Haimovich9, Jason A. Freed10, Andrew Olson11, Daniel J. Morgan12,13, Jason Hom5, Robert Gallo14,15, Liam G. McCoy1,16,17, Haadi Mombini18, Christopher Lucas1, Misha Fotoohi1, Matthew Gwiazdon1, Daniele Restifo1, Daniel Restrepo19, Eric Horvitz20,21, Jonathan Chen3,4,22‡,Arjun K. Manrai2‡*, Adam Rodman1‡* More than 65 years ago, complex clinical diagnostic reasoning cases were introduced as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. In this study, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases across five experiments with a baseline of hundreds of physicians. …
摘自《科学》(Science)第392卷 第6797期 · 2026年4月30日,P. G. Brodeur et al.。仅引用开头一小段供了解文章,版权归原刊所有,全文请阅读原刊。