
Read full paper
Abstract
Large language models (LLMs) in medicine are mainly evaluated using multiple-choice question answering (MCQA), which can overestimate real clinical ability due to guessing strategies and answer biases. To address these limitations, we introduce an expanded and more challenging benchmark based on Polish medical exams, adding over 15,000 questions, two new domains, and four structural modifications that reduce MCQA-specific artifacts and better test reasoning. We evaluate 21 LLMs and show that evaluation design strongly affects results. Under our harder setup, the best model (Qwen3.5-122B) drops by 28.4 and 31 pp on English and
Polish exams, respectively. Despite low evidence of data contamination, standard MCQA scores do not reliably reflect true medical competence. To facilitate further research, we make our benchmark publicly available
Authors: Antoni Lasik, Jakub Pokrywka, Łukasz Grzybowski, Jeremi Ignacy Kaczmarek, Gabriela Korzanska, Janusz Swieczkowski-Feiz, Oskar Pastuszek, Paulina Hoffman, Jakub Tomasz Dąbrowski, Wojciech Kusa
References
- AI@Meta. 2024. Llama 3 model card.
- AI@Meta. 2025. meta-llama/llama-3.3-70b-instruct. https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct. Instruction-tuned Llama 3.3 70B large language model.
- Iñigo Alonso, Maite Oronoz, and Rodrigo Agerri. 2024. Medexpqa: Multilingual benchmarking of large language models for medical question answering. Artificial intelligence in medicine, 155:102938.
- Rahul K Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero-Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, and 1 others. 2025. Healthbench: Evaluating large language models towards improved human health. arXiv preprint arXiv:2505.08775.
- Nishant Balepur, Abhilasha Ravichander, and Rachel Rudinger. 2024. Artifacts or abduction: How do LLMs answer multiple-choice questions without the question? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10308–10330, Bangkok, Thailand. Association for Computational Linguistics.
- Nishant Balepur, Rachel Rudinger, and Jordan Lee Boyd-Graber. 2025. Which of these best describes multiple choice evaluation with LLMs? a) forced B) flawed C) fixable D) all of the above. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3394–3418, Vienna, Austria. Association for Computational Linguistics.
- Nikhil Chandak, Shashwat Goel, Ameya Prabhu, Moritz Hardt, and Jonas Geiping. 2025. Answer matching outperforms multiple choice for language model evaluation. Preprint, arXiv:2507.02856.
- Michał Chojnicki, Katarzyna Kaczmarek-Majer, Paweł Burchardt, Yanwu Ren, and Marek Z Reformat. 2025. Pilot assessment of transparency of llm-based systems to support emergency rooms. In Proceedings of the Second Workshop on Explainable Artificial Intelligence for the medical domain-25-30 October.
- Jean-Philippe Corbeil, Amin Dada, Jean-Michel Attendu, Asma Ben Abacha, Alessandro Sordoni, Lucas Caccia, Francois Beaulieu, Thomas Lin, Jens Kleesiek, and Paul Vozila. 2025. A modular approach for clinical SLMs driven by synthetic data with pre-instruction tuning, model merging, and clinical-tasks alignment. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 19352– 19374, Vienna, Austria. Association for Computational Linguistics.
- DeepSeek-AI. 2026. Deepseek-v4: Towards highly efficient million-token context intelligence.
- Yella Diekmann, Chase Fensore, Rodrigo CarrilloLarco, Eduard Castejon Rosales, Sakshi Shiromani, Rima Pai, Megha Shah, and Joyce Ho. 2025. Llms as medical safety judges: Evaluating alignment with human annotation in patient-facing qa. In Proceedings of the 24th Workshop on Biomedical Language Processing, pages 217–224.
- Shahriar Golchin and Mihai Surdeanu. 2025. Data contamination quiz: A tool to detect and estimate contamination in large language models. Transactions of the Association for Computational Linguistics, 13:809–830.
- Łukasz Grzybowski, Jakub Pokrywka, Michał Ciesiółka, Jeremi Ignacy Kaczmarek, and Marek Kubis. 2025. Polish-english medical knowledge transfer: A new benchmark and results. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 9042–9063.
- Yu Gu, Jingjing Fu, Xiaodong Liu, Jeya Maria Jose Valanarasu, Noel CF Codella, Reuben Tan, Qianchu Liu, Ying Jin, Sheng Zhang, Jinyu Wang, and 1 others. 2025. The illusion of readiness: Stress testing large frontier models on multimodal medical benchmarks. arXiv preprint arXiv:2509.18234.
- Krzysztof Jassem, Michał Ciesiółka, Filip Gralinski, Piotr Jabłonski, Jakub Pokrywka, Marek Kubis, Monika Jabłonska, and Ryszard Staruch. 2025. Llmzsz {\L}: a comprehensive llm benchmark for polish. arXiv preprint arXiv:2501.02266.
- Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421.
- Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019. Pubmedqa: A dataset for biomedical research question answering. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), pages 2567–2577.
- Yiqiao Jin, Mohit Chandra, Gaurav Verma, Yibo Hu, Munmun De Choudhury, and Srijan Kumar. 2024. Better to ask in english: Cross-lingual evaluation of large language models for healthcare queries. In Proceedings of the ACM Web Conference 2024, pages 2627–2638.
- Jan Kocon, Maciej Piasecki, Arkadiusz Janz, Teddy Ferdinan, Łukasz Radlinski, Bartłomiej Koptyra, Marcin Oleksy, Stanisław Wo´zniak, Paweł Walkowiak, Konrad Wojtasik, Julia Moska, Tomasz Naskret, Bartosz Walkowiak, Mateusz Gniewkowski, Kamil Szyc, Dawid Motyka, Dawid Banach, Jonatan Dalasinski, Ewa Rudnicka, and 80 others. 2025. Pllum: A family of polish large language models. arXiv preprint arXiv:2511.03823.
- Yanis Labrak, Adrien Bazoge, Emmanuel Morin, PierreAntoine Gourraud, Mickael Rouvier, and Richard Dufour. 2024. Biomistral: A collection of opensource pretrained large language models for medical domains. Preprint, arXiv:2402.10373.
- Jan Nicikowski, Mikołaj Szczepanski, Miłosz Miedziaszczyk, and Bartosz Kudlinski. 2024. The potential of chatgpt in medicine: an example analysis of nephrology specialty exams in poland. Clinical kidney journal, 17(8):sfae193.
- Charles Nimo, Tobi Olatunji, Abraham Toluwase Owodunni, Tassallah Abdullahi, Emmanuel Ayodele, Mardhiyah Sanni, Ezinwanne C Aka, Folafunmi Omofoye, Foutse Yuehgoh, Timothy Faniran, and 1 others. 2025. Afrimed-qa: a pan-african, multispecialty, medical question-answering benchmark dataset. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1948–1973.
- Krzysztof Ociepa, Remigiusz Kinas, Krzysztof Wróbel, Adrian Gwo´L¸sdziej, and 1 others. 2025a. Bielik 11b v3: Multilingual large language model for european languages. arXiv preprint arXiv:2601.11579.
- Krzysztof Ociepa, Łukasz Flis, Remigiusz Kinas, Krzysztof Wróbel, and Adrian Gwo´zdziej. 2025b. Bielik v3 small: Technical report. Preprint, arXiv:2505.02550.
- Krzysztof Ociepa, Łukasz Flis, Krzysztof Wróbel, Adrian Gwo´zdziej, and Remigiusz Kinas. 2025c. Bielik 11b v2 technical report. Preprint, arXiv:2505.02410.
- OpenAI. 2025a. GPT-5 mini. https://chat.openai.com/.
- OpenAI. 2025b. gpt-oss-120b & gpt-oss-20b model card. Preprint, arXiv:2508.10925.
- OpenMeditron. Meditron3-70b. https://huggingface.co/OpenMeditron/Meditron3-70B. Large language model specialized in clinical medicine, based on Llama-3.1.
- Ankit Pal and Malaikannan Sankarasubbu. 2024. aaditya/llama3-openbiollm-70b. https://huggingface.co/aaditya/Llama3-OpenBioLLM-70B. Open source biomedical LLM fine-tuned from LLaMA-3 with 70B parameters.
- Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Conference on health, inference, and learning, pages 248–260. PMLR.
- Pouya Pezeshkpour and Estevam Hruschka. 2023. Large language models sensitivity to the order of options in multiple-choice questions, 2023. URL https://arxiv. org/abs/2308.11483.
- Jakub Pokrywka, Jeremi Kaczmarek, and Edward Gorzelanczyk. 2024. Gpt-4 passes most of the 297 written polish board certification examinations. arXiv preprint arXiv:2405.01589.
- Team Qwen. 2025. Qwen3 technical report. Preprint, arXiv:2505.09388.
- Yuval Reif and Roy Schwartz. 2024. Beyond performance: Quantifying and mitigating label bias in llms. arXiv preprint arXiv:2405.02743.
- Maciej Rosoł, Jakub S G ˛asior, Jonasz Łaba, Kacper Korzeniewski, and Marcel Młynczak. 2023. Evaluation of the performance of gpt-3.5 and gpt-4 on the polish medical final examination. Scientific Reports, 13(1):20512.
- Andrew Sellergren, Sahar Kazemzadeh, Tiam Jaroensri, Atilla Kiraly, Madeleine Traverse, Timo Kohlberger, Shawn Xu, Fayaz Jamil, Cían Hughes, Charles Lau, Justin Chen, Fereshteh Mahvar, Liron Yatziv, Tiffany Chen, Bram Sterling, Stefanie Anna Baby, Susanna Maria Baby, Jeremy Lai, Samuel Schmidgall, and 62 others. 2025. Medgemma technical report. Preprint, arXiv:2507.05201.
- Julia Siebielec, Michal Ordak, Agata Oskroba, Anna Dworakowska, and Magdalena Bujalska-Zadrozny. 2024. Assessment study of chatgpt-3.5’s performance on the final polish medical examination: Accuracy in answering 980 questions. In Healthcare, volume 12, page 1637. MDPI.
- Shrutika Singh, Anton Alyakin, Daniel Alexander Alber, Jaden Stryker, Ai Phuong S Tong, Karl Sangwon, Nicolas Goff, Mathew De La Paz, Miguel HernandezRovira, Ki Yun Park, and 1 others. 2025. The pitfalls of multiple-choice questions in generative ai and medical education. Scientific Reports, 15(1):42096.
- Szymon Suwala, Paulina Szulc, Aleksandra Dudek, Aleksandra Bialczyk, Kinga Koperska, and Roman Junik. 2023. Chatgpt fails the polish board certification examination in internal medicine: artificial intelligence still has much to learn. Pol Arch Int Med Pol Arch Med Wewnet, 133(11).
- Annalisa Szymanski, Noah Ziems, Heather A EicherMiller, Toby Jia-Jun Li, Meng Jiang, and Ronald A Metoyer. 2025. Limitations of the llm-as-a-judge approach for evaluating llm outputs in expert knowledge tasks. In Proceedings of the 30th International Conference on Intelligent User Interfaces, pages 952–966.
- Gemma Team. 2025. Gemma 3.
- Team Qwen. 2026. Qwen3.5: Towards native multimodal agents.
- Dorota Wójcik, Ola Adamiak, Gabriela Czerepak, Oskar Tokarczuk, and Leszek Szalewski. 2024. A comparative analysis of the performance of chatgpt4, gemini and claude for the polish medical final diploma exam and medical-dental verification exam. MedRxiv, pages 2024–07.
- Rui Yang, Ting Fang Tan, Wei Lu, Arun James Thirunavukarasu, Daniel Shu Wei Ting, and Nan Liu. 2023. Large language models in health care: Development, applications, and challenges. HealthCare Science, 2(4):255–263.
- Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2024. Large language models are not robust multiple choice selectors. In International Conference on Learning Rep





