Mitra, Nilesh Kumar
ORCID: 0000-0002-8487-4607, Vishnumukkala, Thirupathirao, Win, Thin Thin, Amirthalingam, Sasikala Devi
ORCID: 0000-0002-9004-5573, Kwa, Siew Kim, Thangarasu, Gunasekar, Sumera, Afshan
ORCID: 0000-0002-4108-2271 and Kandiah, Skantha
(2026)
Quality of Large Language Model-Generated MCQs Across Three Medical Disciplines: An Expert Rater-Based Comparison of Gemini, GPT-4 and Perplexity Pro.
Advances in Medical Education and Practice, 17
.
p. 618711.
Preview |
PDF (VOR)
- Published Version
Available under License Creative Commons Attribution Non-commercial. 4MB |
Official URL: https://doi.org/10.2147/AMEP.S618711
Abstract
Background: The increasing number of student cohorts has compelled academics to create a larger number of Multiple-Choice Question (MCQ) items. Large language models (LLMs) can help educators generate assessment items across multiple disciplines. The study compared the ability of LLMs (Gemini Advanced, Perplexity Pro, and ChatGPT 4.0) to generate high-quality, clinical-scenario-based MCQ items across three disciplines in a medical program, using an 8-point quality rubric.
Materials and Methods: Learning Outcomes (LOs) from Anatomy and Pathology disciplines of a pre-clinical semester 4 module and Family Medicine discipline of a clinical semester 6 module of a medical undergraduate program were selected. Using a pre-determined descriptive prompt, 63 MCQ items were generated from three AI tools. The quality of item construction was assessed by external content experts who were blinded to item generation using an 8-criterion rubric and a 4-point Likert scale. Mean scores and ranks for MCQs under each LLM were analysed, and a Friedman test was conducted to compare them. Kendall’s W showed that all criteria except one demonstrated some effect and weak-to-fair inter-rater reliability.
Results: When measuring key problem-solving skills, Perplexity Pro-generated MCQ items received a higher percentage of "strongly agree” ratings in Anatomy (57.1%, n=63), Pathology (56.5%, n=63), and Family Medicine (71.4%, n=63). Perplexity Pro received a higher percentage of "strongly agree” ratings in Anatomy, at 47.6% (n = 63), when evaluating specific content. Gemini Advanced also scored highly, with 68.8% in Pathology and 81% in Family Medicine (n = 63). A comparative analysis of the higher mean scores and ranks across 24 quality criteria showed that Perplexity Pro, Gemini Advanced, and ChatGPT 4.0 achieved 13, 10, and 1 higher mean scores, respectively. There is no significant difference in mean scores and ranks among the three LLMs across different quality criteria.
Conclusion: MCQs created by Perplexity Pro and Gemini Advanced achieved comparatively higher percentages of "strongly agree” ratings across quality criteria for item construction. Assessing LLM-generated assessment items provides valuable insight into the quality of LLM-supported MCQ questions.
Repository Staff Only: item control page
Lists
Lists