Federated artificial intelligence monitoring service (FAMOS): an in silico feasibility study.
Authors
Affiliations (8)
Affiliations (8)
- Newton's Tree, London, W1B 1NT, United Kingdom.
- Radiology, Barts Health NHS Trust, London, E1 2ES, United Kingdom.
- King's College London, London, SE1 7EH, United Kingdom.
- Radiology.Leeds Teaching Hospitals NHS Trust, Leeds, LS9 7TF, United Kingdom.
- University of Leeds, LeedsLS2 9JT, United Kingdom.
- University of Bristol, Bristol, BS8 1QU, United Kingdom.
- NHS Greater Glasgow and Clyde, Glasgow, G3 8SJ, United Kingdom.
- Health Tech Innovation and Translation Lab, University of Glasgow, Glasgow, G3 8SJ, United Kingdom.
Abstract
To evaluate feasibility of a federated AI monitoring service (FAMOS) for post-deployment surveillance of third-party AI applications used in chest X-ray (CXR) interpretation. FAMOS was deployed at 2 NHS Trusts using a federated architecture enabling local data processing while maintaining data governance compliance. De-identified CXRs from patients aged >18 years were retrospectively identified, along with relevant patient attributes (age, sex, inpatient status, image orientation, season, artefact). Chest X-rays were processed by 2 AI applications to simulate real-world deployment. FAMOS analyzed input data, AI inference values, and longitudinal human-AI agreement as proxy indicators of data, prediction, and behavioral drift. Input monitoring used image feature embeddings analyzed with principal component analysis and Hotelling's <i>T</i>² statistics, while risk-adjusted cumulative sum charts were applied to sequentially monitor AI inference values. human-AI agreement was evaluated at 10 time points over 3 months. Input monitoring identified a small proportion of outlier examinations, predominantly associated with modifiable image quality issues. Artificial intelligence inference values remained stable across most findings for both vendors, with limited drift events detected. Human-AI agreement patterns differed between sites, remaining stable at 1 site while increasing over time at another, suggesting evolving automation bias. Real-time, federated monitoring of deployed radiology AI systems is technically and operationally feasible within clinical environments. Multi-domain platform-based monitoring provides scalable, independent oversight and may function as an early-warning system supporting identification of emerging risks following deployment. This study introduces a federated, multi-domain monitoring framework integrating sociotechnical indicators for continuous post-deployment surveillance of clinical radiology AI tools.