Home Resources Polish-English medical knowledge transfer: A new benchmark and results

Polish-English medical knowledge transfer: A new benchmark and results

Read full paper

Abstract

Large Language Models (LLMs) have demonstrated significant potential in specialized tasks, including medical problem-solving. However, most studies predominantly focus on Englishlanguage contexts. This study introduces a novel benchmark dataset based on Polish medical licensing and specialization exams (LEK, LDEK, PES). The dataset, sourced from publicly available materials provided by the Medical Examination Center and the Chief Medical Chamber, includes Polish medical exam questions, along with a subset of parallel PolishEnglish corpora professionally translated for foreign candidates. By structuring a benchmark
from these exam questions, we evaluate stateof-the-art LLMs, spanning general-purpose, domain-specific, and Polish-specific models, and compare their performance with that of human medical students and doctors. Our analysis shows that while models like GPT-4o achieve
near-human performance, challenges persist in cross-lingual translation and domain-specific understanding. These findings highlight disparities in model performance across languages and medical specialties, emphasizing the limitations and ethical considerations of deploying LLMs in clinical practice.

Authors: Łukasz Grzybowski, Jakub Pokrywka, Michal Ciesiółka, Jeremi I. Kaczmarek, Marek Kubis

References

  1. Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403.
  3. Andrew M. Bean, Karolina Korgul, Felix Krones, Robert McCraith, and Adam Mahdi. 2024. Do large language models have shared weaknesses in medical question answering?
  4. Martin Juan José Bucher and Marco Martini. 2024. Fine-tuned’small’llms (still) significantly outperform zero-shot generative ai models in text classification. arXiv preprint arXiv:2406.08660.
  5. Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3):1–45.
  6. Michelle Clark and Sharon Bailey. 2024. Chatbots in health care: Connecting patients to information. Canadian Journal of Health Technologies, 4(1). 9050
  7. Jan Clusmann, Fiona R Kolbinger, Hannah Sophie Muti, Zunamys I Carrero, Jan-Niklas Eckardt,
  8. Narmin Ghaffari Laleh, Chiara Maria Lavinia Löffler, Sophie-Caroline Schwarzkopf, Michaela Unger, Gregory P Veldhuizen, et al. 2023. The future landscape of large language models in medicine. Communications medicine, 3(1):141.
  9. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
  10. Clémentine Fourrier, Nathan Habib, Alina Lozovskaya, Konrad Szafer, and Thomas Wolf. 2024. Open llm leaderboard v2. https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard.
  11. Gemini, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
  12. Stefan Harrer. 2023. Attention is not all you need: the complicated case of ethically using large language models in healthcare and medicine. EBioMedicine, 90.
  13. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.
  14. Niclas Hertzberg and Anna Lokrantz. 2024. Medqaswe-a clinical question & answer dataset for swedish. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 11178–11186.
  15. Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L’elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b. ArXiv, abs/2310.06825.
  16. Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421.
  17. Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W Cohen, and Xinghua Lu. 2019. Pubmedqa: A dataset for biomedical research question answering. arXiv preprint arXiv:1909.06146.
  18. Yiqiao Jin, Mohit Chandra, Gaurav Verma, Yibo Hu, Munmun De Choudhury, and Srijan Kumar. 2024. Better to ask in english: Cross-lingual evaluation of large language models for healthcare queries. In Proceedings of the ACM Web Conference 2024, WWW’24, page 2627–2638, New York, NY, USA. Association for Computing Machinery.
  19. johnsnowlabs. 2024. Jsl-medllama-3-8b-v2.0. https://huggingface.co/johnsnowlabs/JSL-MedLlama-3-8B-v2.0. Accessed: 2024-11-02.
  20. Mert Karabacak and Konstantinos Margetis. 2023. Embracing large language models for medical applications: opportunities and challenges. Cureus, 15(5).
  21. Jungo Kasai, Yuhei Kasai, Keisuke Sakaguchi, Yutaro Yamada, and Dragomir Radev. 2023. Evaluating gpt-4 and chatgpt on japanese medical licensing examinations.
  22. Markus Kipp. 2024. From gpt-3.5 to gpt-4.o: A leap in ai’s medical exam performance. Information, 15(9).
  23. Jakub Kufel, Iga Paszkiewicz, Michał Bielówka, Wiktoria Bartnikowska, Michał Janik, Magdalena Stencel, Łukasz Czogalik, Katarzyna Gruszczynska, and Sylwia Mielcarska. 2023. Will chatgpt pass the polish specialty exam in radiology and diagnostic imaging? insights into strengths and limitations. Polish Journal of Radiology, 88:e430.
  24. Yanis Labrak, Adrien Bazoge, Emmanuel Morin, PierreAntoine Gourraud, Mickael Rouvier, and Richard Dufour. 2024. Biomistral: A collection of opensource pretrained large language models for medical domains.
  25. Peter Lee, Sebastien Bubeck, and Joseph Petro. 2023. Benefits, limits, and risks of gpt-4 as an ai chatbot for medicine. New England Journal of Medicine, 388(13):1233–1239.
  26. Hanzhou Li, John T Moon, Saptarshi Purkayastha, Leo Anthony Celi, Hari Trivedi, and Judy W Gichoya. 2023. Ethics of large language models in medicine and medical research. The Lancet Digital Health, 5(6):e333–e335.
  27. Jialin Liu, Changyu Wang, and Siru Liu. 2023. Utility of chatgpt in clinical practice. J Med Internet Res, 25:e48568.
  28. Junling Liu, Peilin Zhou, Yining Hua, Dading Chong, Zhongyu Tian, Andrew Liu, Helin Wang, Chenyu You, Zhenhua Guo, Lei Zhu, et al. 2024. Benchmarking large language models on cmexam-a comprehensive chinese medical exam dataset. Advances in Neural Information Processing Systems, 36.
  29. Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024. Large language models: A survey. arXiv preprint arXiv:2402.06196.
  30. Zabir Al Nazi and Wei Peng. 2024. Large language models in healthcare and medical domain: A review. In Informatics, volume 11, page 57. MDPI. 9051
  31. Harsha Nori, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, et al. 2023. Can generalist foundation models outcompete special-purpose tuning? case study in medicine. arXiv preprint arXiv:2311.16452.
  32. Krzysztof Ociepa, Łukasz Flis, Krzysztof Wróbel, Adrian Gwo´zdziej, and SpeakLeash Teamand Cyfronet Team. 2024. Introducing bielik-7b-v0.1: Polish language model. Accessed: 2024-11-02.
  33. OpenMeditron. 2024. Meditron3-70b. https://huggingface.co/OpenMeditron/Meditron3-70B. Accessed: 2024-11-02.
  34. Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Conference on health, inference, and learning, pages 248–260. PMLR.
  35. Ye-Jean Park, Abhinav Pillai, Jiawen Deng, Eddie Guo, Mehul Gupta, Mike Paget, and Christopher Naugler. 2024. Assessing the research landscape and clinical utility of large language models: A scoping review. BMC Medical Informatics and Decision Making, 24(1):72.
  36. Jakub Pokrywka, Jeremi Kaczmarek, and Edward Gorzelanczyk. 2024. ´ Gpt-4 passes most of the 297 written polish board certification examinations.
  37. Maciej Rosoł, Jakub S G ˛asior, Jonasz Łaba, Kacper Korzeniewski, and Marcel Młynczak. 2023. Evaluation of the performance of gpt-3.5 and gpt-4 on the polish medical final examination. Scientific Reports, 13(1):20512.
  38. Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R Pfohl, Heather Cole-Lewis, et al. 2025. Toward expert-level medical question answering with large language models. Nature Medicine, pages 1–8.
  39. S Suwała, P Szulc, A Dudek, A Białczyk, K Koperska, and R Junik. 2023. Chatgpt fails the internal medicine state specialization exam in poland: artificial intelligence still has much to learn. Pol Arch Intern Med, 133(11):16608.
  40. Qwen Team. 2024. Qwen2.5: A party of foundation models.
  41. Ehsan Ullah, Anil Parwani, Mirza Mansoor Baig, and Rajendra Singh. 2024. Challenges and barriers of using large language models (llm) such as chatgpt for diagnostic medicine with a focus on digital pathology–a recent scoping review. Diagnostic pathology, 19(1):43.
  42. Simona Wojcik, Anna Rulkiewicz, Piotr Pruszczyk, Wojciech Lisik, Marcin Pobozy, Iwona Pilchowska, and Justyna Domienik-Karłowicz. 2023. Beyond human understanding: Benchmarking language models for polish cariology expertise.
  43. T Wolf. 2019. Huggingface’s transformers: State-ofthe-art natural language processing. arXiv preprint arXiv:1910.03771.
  44. R. Yang, T. F. Tan, W. Lu, A. J. Thirunavukarasu, D. S. W. Ting, and N. Liu. 2023. Large language models in health care: Development, applications, and challenges. Health Care Science, 2(4):255–263.
  45. T. Zhou, D. Salman, and A. H. McGregor. 2024. Recent clinical practice guidelines for the management of low back pain: a global comparison. BMC Musculoskeletal Disorders, 25(1):344