Reference-Label Construction and Annotation Reporting in CT- and MRI-Based Adipose Tissue Segmentation: A Systematic Review and Evidence Map.
Authors
Affiliations (2)
Affiliations (2)
- ITMO University, Saint Petersburg, Russia. [email protected].
- Federal State Budgetary Institution "V.A. Almazov National Medical Research Centre" of the Ministry of Health of the Russian Federation, 197341, Saint Petersburg, Russia.
Abstract
CT/MRI adipose tissue segmentation underpins body-composition and medical-imaging AI research, yet reference labels are often described only as "manual," "expert," or "reference" segmentations, obscuring heterogeneity in ground-truth construction. To systematically review how reference labels, annotation procedures, validation standards, and reproducibility-relevant reporting are defined in CT- and MRI-based adipose tissue segmentation studies. The review protocol was prospectively registered in OSF Registries before the formal multi-database search. We searched PubMed/MEDLINE, Embase, Scopus, Web of Science, IEEE Xplore, ACM Digital Library, SpringerLink/MICCAI/LNCS, and SPIE Digital Library. Of 25,774 records, 8619 remained after deduplication, and 204 studies were included. The final corpus comprised 204 studies organized into 122 Tier 1, 64 Tier 2, and 18 Tier 3 synthesis groups. CT-family imaging accounted for 106 studies and MRI-only imaging for 67. Manual, expert, or reference labels were identified in 178 studies (87.3%), while automated or model-based computational methods were used in 181 (88.7%). External or multicenter validation was reported in 37 studies (18.1%), and reader- or operator-level reproducibility assessment in 64 (31.4%). Data availability was explicitly reported or partly available in 68 studies (33.3%), code, software, or model availability in 57 (27.9%), and label, mask, contour, or annotation availability in 3 (1.5%). Mean reporting completeness was 9.44/12; 149 studies (73.0%) had high reporting completeness, and 55 (27.0%) had moderate reporting completeness. Reference labels should be treated as provenance-bearing data objects rather than self-evident ground truth. Reproducible medical-imaging AI evaluation requires explicit reporting of compartment definitions, boundary rules, annotator roles, correction/adjudication, validation, and label availability. We propose a domain-specific annotation-chain reporting checklist that complements existing medical-imaging AI guidance by operationalizing adipose-compartment definitions, anatomical boundary rules, correction workflows, reference-label provenance, and label availability. For algorithm development, transparent annotation-chain reporting is also necessary to interpret training targets, benchmark performance, label-efficient learning, and cross-dataset validation.