Safeguarding biomedical AI: a critical scoping review of privacy-enhancing technologies, hybrid approaches, and deployment models.
Authors
Affiliations (3)
Affiliations (3)
- Department of Biomedical Engineering, Wake Forest School of Medicine, Winston-Salem, NC, United States.
- National Cancer Institute, Shady Grove, Rockville, MD, United States.
- Department of Cancer Biology, Wake Forest School of Medicine, Winston-Salem, NC, United States.
Abstract
Biomedical artificial intelligence (AI) requires the integration of privacy-enhancing technologies (PETs) to safeguard sensitive clinical, imaging, and genomic data while preserving analytical utility. This review critically and systematically maps applications of PETs across the biomedical AI lifecycle in accordance with PRISMA-ScR guidelines and evaluates their technical trade-offs, deployment feasibility, and residual risks. We systematically searched PubMed, IEEE Xplore, ACM Digital Library, and Scopus for studies published between 2015 and 2025. Eligible studies addressed differential privacy, federated learning, secure multiparty computation, homomorphic encryption, or hybrid approaches in biomedical AI. Data were charted on PET type, modality, lifecycle stage, utility metrics, privacy parameters, and deployment considerations. A critical appraisal rubric assessed threat-model adequacy, methodological clarity, reproducibility, privacy-utility transparency, and deployment realism. Additionally, we hand-searched major venues (USENIX Security, NeurIPS, AAAI) and screened Google Scholar for grey literature, applying de-duplication across sources. We identified 87 studies spanning clinical decision support, genomics, and medical imaging. From 25,761 initial records, 3,754 underwent title/abstract screening and 1,968 underwent full-text assessment. PETs demonstrated distinct strengths and limitations: differential privacy provided provable guarantees but reduced performance on imbalanced data; federated learning improved data access but remained vulnerable to gradient leakage; and cryptographic methods ensured confidentiality at high computational cost. Synthetic data generation supported privacy-conscious data sharing and benchmarking but remained sensitive to disclosure risk, fidelity loss, and subgroup representation. Hybrid and emerging approaches, including trusted execution environments, zero-knowledge proofs, and privacy-preserving transformer architectures, mitigated composability gaps yet lacked full end-to-end assurance. Case studies at hospital and biobank scale illustrated practical feasibility and infrastructure demands. Situating PETs within technical and operational contexts clarifies their capabilities, limitations, and deployment challenges. Residual risks persist, including fairness concerns, inference-time leakage, and overreliance on PETs as compliance proxies. Sustained technical innovation and institutional governance remain essential for the trustworthy integration of PETs in biomedical AI.