Cognitive Workload and Mental Burden in Health Care Professionals Interacting With AI: Systematic Review and Meta-Analysis.
Authors
Affiliations (4)
Affiliations (4)
- Institute for Liver and Digestive Diseases, Hallym University, Chuncheon, Republic of Korea.
- Institute of New Frontier Research, Hallym University College of Medicine, Chuncheon, Republic of Korea.
- Department of Internal Medicine, Hallym University College of Medicine, Sakju-ro 77, Chuncheon, 24253, Republic of Korea, 82 82 33 240 582.
- Department of Anesthesiology and Pain Medicine, Hallym University College of Medicine, Chuncheon, Republic of Korea.
Abstract
AI adoption in health care has accelerated rapidly, with ambient documentation tools, diagnostic imaging AI, and clinical decision support systems (CDSSs) entering routine practice. However, the cognitive demands placed on clinicians supervising these systems remain understudied. Specifically, the concept of verification burden requires closer examination. Consequently, institutional decision-makers lack a structured, certainty-graded evidence base regarding the true impact of AI on clinician workload and burnout. This study aimed to systematically review evidence on cognitive workload and burnout in health care professionals that use AI-powered clinical tools, quantify pooled effects under a conservative inferential framework, and assess certainty of evidence by AI category. The study was registered in PROSPERO (CRD420261284298) and reported per PRISMA 2020 and PRISMA-S guidelines. We searched MEDLINE, Embase, Web of Science, and Cochrane CENTRAL (January 2015-2026) for studies measuring cognitive workload or burnout using validated instruments (NASA Task Load Index [NASA-TLX] and Professional Fulfillment Index [PFI]) among health care professionals using clinical AI. Risk of bias was assessed using ROB 2.0 and ROBINS-I; certainty was rated using GRADE. Meta-analyses applied Hartung-Knapp-Sidik-Jonkman adjustment with restricted maximum likelihood estimation, incorporating prediction intervals (PIs). We included 21 studies representing 2885 health care professionals across 7 countries. The synthesis demonstrated that the cognitive impact of clinical AI varies according to its specific application. Pooled analyses of ambient AI documentation showed statistically significant reductions in NASA-TLX temporal demand (SMD -1.46, 95% CI -2.81 to -0.11; k=2; I2=31.1%) and effort (SMD -1.29, 95% CI -2.16 to -0.42; k=2; I2=0%), PFI work exhaustion (MD -0.35, 95% CI -0.58 to -0.12; k=3; I2=0%; 95% PI -1.03 to 0.33), and burnout prevalence (OR 0.47, 95% CI 0.25-0.86; k=3; I2=0%; 95% PI 0.06-3.82). Two pools favored ambient AI but did not reach significance at k=2: NASA-TLX mental demand (SMD -1.29, 95% CI -3.64 to 1.07) and documentation time (SMD -0.24, 95% CI -1.10 to 0.61). Diagnostic imaging AI and CDSS showed mixed or paradoxically increased workload. GRADE certainty was moderate for cognitive workload reduction with ambient AI, low for burnout reduction with ambient AI, and very low for imaging AI and CDSS outcomes. This review combines validated workload instruments, meta-analysis, and PIs in health care AI, delivering a GRADE certainty assessment across 5 AI categories that prior accuracy- or efficiency-focused reviews have not provided. Ambient AI documentation was associated with reduced cognitive workload and burnout, but only in voluntary early-adopter cohorts and based on few studies; the conservative CIs were wide and, where estimable, PIs crossed the null. Findings inform institutional pilots with prospective workload measurement, regulatory human-factors evaluation of AI medical devices, and human-centered AI design. Net benefit on the health care workforce remains an open empirical question.