Detectability and healthcare implications of generative AI-synthesized chest radiographs: a blinded radiologist reader study.
Authors
Affiliations (4)
Affiliations (4)
- Department of Radiology, The Second Affiliated Hospital of Xinjiang Medical University, Ürümqi, Xinjiang, China.
- Department of Radiology, The Second Xiangya Hospital of Central South University, Changsha, Hunan, China.
- Department of Radiology, The First Hospital of Lanzhou University, Lanzhou, Gansu, China.
- Department of Radiology, Pengan County People's Hospital, Nanchong, Sichuan, China.
Abstract
Generative artificial intelligence (GenAI) is increasingly explored for medical image synthesis, medical education, and dataset augmentation; however, the detectability and reader-perceived visual authenticity of synthetic chest radiographs generated by accessible multimodal models remain insufficiently understood. This study evaluated synthetic disease-specific chest radiographs generated by gpt-image-2 and gemini-3-pro-image-preview, hereafter referred to as the GPT-image model and the Gemini-image model, respectively, using two generation strategies: text-only generation and image-conditioned generation based on age- and sex-matched normal conditioning radiographs. We included 320 real disease-positive frontal chest radiographs covering cardiomegaly, pneumothorax, pleural effusion, and pneumonia, together with 320 matched normal conditioning radiographs. Synthetic and real disease-positive radiographs were randomized and independently assessed by four radiologists in a blinded single-image reader study, followed by paired evaluation of previously undetected image-conditioned synthetic radiographs. Image-level similarity was assessed using the structural similarity index measure (SSIM), learned perceptual image patch similarity (LPIPS), and Fréchet Inception Distance (FID). Image-conditioned generation had significantly lower AI detection rates than text-only generation (34.0% vs. 56.1%) and consistently lower FID values. Synthetic radiographs generated by the GPT-image model were less frequently detected than those generated by the Gemini-image model and showed higher image-level structural and perceptual similarity to the conditioning radiographs. Paired comparison improved detection of previously undetected image-conditioned synthetic radiographs. These findings indicate that image-conditioned GenAI can produce synthetic chest radiographs that may appear visually authentic to radiologists, highlighting the need for transparent labeling, provenance tracking, expert review, and controlled integration of synthetic medical images into healthcare, education, and research workflows.