Chen, Yingjian, Gao, Fan, Tong, Sherry T., Zhang, Haoyu, Feng, Aosong, Jin, Kevin W., Wu, Xing, Lu, Jinghui, Samad, Abdul, Faruqi, Akbar, Caraballo, Cesar, Brandão, Cibele, Dhruva, Gupta, Jeon, Eunji, Madera-Santiago, Gabriel, Lee, Geon, Itikawa, Hugo Toshio, Cho, Insook, Martins, Isabelli, Siddique, Isarar, Ahmed, Israr, Kwak, Jihyo, Veerakanjana, Kanyakorn, Cardoso, Luis Guilherme, Kim, Minjin, Ittichaiwong, Piyalitt, Dua, Renee, Gudiño-Rosales, Santiago, Chen, Xiujie, Lapalus, Zeo, Xu, Zixin, Yasunaga, Michihiro, Ying, Rex, Lim, Heuiseok, Kang, Jaewoo, Park, Chanjun, Jiang, Hang, Goh, Ethan, Kim, Hyunjae, Marrese-Taylor, Edison, Iwasawa, Yusuke, Matsuo, Yutaka, Chen, Qingyu, Li, Irene
Abstract
We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions. Each translation was evaluated and revised by two experts fluent in English and the corresponding target language. On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models. The strongest proprietary models were the most stable across languages, whereas many open-source and medically specialized models showed larger and less consistent gaps. Medical specialization alone did not ensure multilingual robustness. Furthermore, expert revision could either raise or lower measured performance, indicating that translation quality materially affects cross-language evaluation results.
Chinese Translation
我们提出了HealMed,这是一个经过专家审查的基准,用于对医学领域大型语言模型进行多语言评估。HealMed包含来自九个数据集的每种语言1,000个示例,涵盖三种任务格式:多项选择问答(MCQA)、自然语言推理(NLI)和开放式问答(QA)。该基准由来自九个国家和地区的23名医生和医学专家在两年内开发完成。每个翻译均由两位精通英语及相应目标语言的专家进行评估和修订。在HealMed上,低资源语言的表现下降最为明显,尽管不同语言和模型之间的差距大小差异显著。最强的专有模型在各语言间表现最为稳定,而许多开源和医学专业模型则显示出更大且不一致的差距。医学专业化本身并未确保多语言的稳健性。此外,专家的修订可能会提高或降低测得的表现,表明翻译质量对跨语言评估结果有实质性影响。