Welcome to

Lancashire Online Knowledge

Image Credit Header image: Artwork by Professor Lubaina Himid, CBE. Photo: @Denise Swanson


Quality of Large Language Model-Generated MCQs Across Three Medical Disciplines: An Expert Rater-Based Comparison of Gemini, GPT-4 and Perplexity Pro

Mitra, Nilesh Kumar orcid iconORCID: 0000-0002-8487-4607, Vishnumukkala, Thirupathirao, Win, Thin Thin, Amirthalingam, Sasikala Devi orcid iconORCID: 0000-0002-9004-5573, Kwa, Siew Kim, Thangarasu, Gunasekar, Sumera, Afshan orcid iconORCID: 0000-0002-4108-2271 and Kandiah, Skantha (2026) Quality of Large Language Model-Generated MCQs Across Three Medical Disciplines: An Expert Rater-Based Comparison of Gemini, GPT-4 and Perplexity Pro. Advances in Medical Education and Practice, 17 . p. 618711.

[thumbnail of VOR]
Preview
PDF (VOR) - Published Version
Available under License Creative Commons Attribution Non-commercial.

4MB

Official URL: https://doi.org/10.2147/AMEP.S618711

Abstract

Background: The increasing number of student cohorts has compelled academics to create a larger number of Multiple-Choice Question (MCQ) items. Large language models (LLMs) can help educators generate assessment items across multiple disciplines. The study compared the ability of LLMs (Gemini Advanced, Perplexity Pro, and ChatGPT 4.0) to generate high-quality, clinical-scenario-based MCQ items across three disciplines in a medical program, using an 8-point quality rubric.

Materials and Methods: Learning Outcomes (LOs) from Anatomy and Pathology disciplines of a pre-clinical semester 4 module and Family Medicine discipline of a clinical semester 6 module of a medical undergraduate program were selected. Using a pre-determined descriptive prompt, 63 MCQ items were generated from three AI tools. The quality of item construction was assessed by external content experts who were blinded to item generation using an 8-criterion rubric and a 4-point Likert scale. Mean scores and ranks for MCQs under each LLM were analysed, and a Friedman test was conducted to compare them. Kendall’s W showed that all criteria except one demonstrated some effect and weak-to-fair inter-rater reliability.

Results: When measuring key problem-solving skills, Perplexity Pro-generated MCQ items received a higher percentage of "strongly agree” ratings in Anatomy (57.1%, n=63), Pathology (56.5%, n=63), and Family Medicine (71.4%, n=63). Perplexity Pro received a higher percentage of "strongly agree” ratings in Anatomy, at 47.6% (n = 63), when evaluating specific content. Gemini Advanced also scored highly, with 68.8% in Pathology and 81% in Family Medicine (n = 63). A comparative analysis of the higher mean scores and ranks across 24 quality criteria showed that Perplexity Pro, Gemini Advanced, and ChatGPT 4.0 achieved 13, 10, and 1 higher mean scores, respectively. There is no significant difference in mean scores and ranks among the three LLMs across different quality criteria.

Conclusion: MCQs created by Perplexity Pro and Gemini Advanced achieved comparatively higher percentages of "strongly agree” ratings across quality criteria for item construction. Assessing LLM-generated assessment items provides valuable insight into the quality of LLM-supported MCQ questions.


Repository Staff Only: item control page