Performance Of Large Language Models On Nursing Licensure Examinations: A Systematic Review And Meta-Analysis
Loading...
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Nurse Education Today
Abstract
Objectives: This systematic review and meta-analysis assessed the performance of large language models (LLMs)
in nursing licensure examinations. Despite the increasing use of LLMs in healthcare education, their capabilities
in nursing licensure examinations remain uncertain. This study provides evidence on the accuracy and limita
tions of LLMs to help guide their integration into nursing education and licensure.
Design: The systematic review and meta-analysis adhered to PRISMA 2020 guidelines.
Data sources: PubMed, CINAHL, PsycINFO, EMCARE, and ERIC were searched from April to June 2025.
Eligibility criteria: Studies were eligible if they evaluated LLMs (e.g., GPT-4, ChatGPT, Qwen-2.5) using multiple-
choice nursing licensure questions under exam-like conditions and reported quantitative accuracy. Open-ended
items were excluded from the meta-analysis due to incompatible scoring methods, but were narratively
synthesised.
Review methods: Two reviewers independently screened, extracted data, and appraised the risk of bias. A random-
effects meta-analysis estimated pooled accuracy; subgroup and meta-regression analyses explored heterogeneity.
Results: Twelve studies assessed 13,870 MCQs across seven exam systems and ten LLMs. Pooled accuracy was
69.6% (95% CI: 65.6–73.6%) with substantial heterogeneity (I
2
= 98%). GPT-4 outperformed GPT-3.5 (77.2%
vs. 60.4%); domain-customised and newer models reached 93.6%. LLMs excelled in general medicine and
pharmacology but underperformed in ethics and psychosocial integrity. Accuracy did not differ significantly by
exam system (p = 0.14), question difficulty (p = 0.90) or format (p = 0.96). In meta-regression, Custom GPT (p =
0.0006) and Qwen 2.5 (p = 0.026) were the only significant predictors of higher accuracy; no exam system,
question format, or difficulty level reached significance. Methodological variability and underreporting of model
parameters were common.
Conclusions: LLMs show promise for low-stakes educational applications, such as formative assessments within
hybrid teaching models; however, they are unsuitable for unmoderated, high-stakes licensure decisions due to
inconsistent performance. Regulatory guidelines, equitable access, and nursing-specific model development are
needed to ensure fairness and validity. Research must prioritise standardised frameworks, error analysis, and
broader geographic representation to address these limitations.
Description
Research Article
Citation
Odoom, A., Kasim, A., Kobiah, E., Diebieri, M., Boateng, E. A., Gyamfi, S., & Hales, C. (2026). Performance of Large Language Models in Nursing Licensure Examinations: A Systematic Review and Meta-Analysis.
