Statistically Rigorous and Interpretable Radiomics for Benign-Malignant Classification on the BUSI-WHU Breast Ultrasound Dataset.
Authors
Affiliations (2)
Affiliations (2)
- Department of Statistics, Mawlana Bhashani Science and Technology University, Tangail, Bangladesh. [email protected].
- Department of Statistics, Mawlana Bhashani Science and Technology University, Tangail, Bangladesh.
Abstract
Radiomics has shown promise for breast ultrasound lesion classification, yet reported performance is often affected by patient-level leakage, feature selection bias, dependence between lesions from the same patient, and image formats lacking pixel-spacing metadata. The BUSI-WHU dataset allows these issues to be addressed together. We analyzed 927 breast lesions from 816 patients using 939 handcrafted radiomic features from expert-annotated images. Pixel spacing was recovered from dataset-supplied scale factors and images resampled to a common isotropic resolution. Features were assessed by Mann-Whitney U tests with Benjamini-Hochberg correction and Cliff's delta, with a one-lesion-per-patient sensitivity analysis. Six classifiers, including a penalized logistic regression baseline, were evaluated under ten-fold patient-grouped, class-stratified cross-validation with correlation filtering and Boruta selection inside each training fold. Confidence intervals and paired comparisons used patient-level bootstrap of out-of-fold predictions, and SHAP was computed on held-out patients. Of the 939 features, 587 differed significantly after correction; most remained significant under one-lesion-per-patient sampling; and 48 were selected consistently across folds. Pooled out-of-fold ROC-AUCs ranged from 0.828 to 0.834; no pairwise difference survived correction, and the logistic regression performed comparably to the best ensemble. Discrimination was unchanged when absolute size features were removed. Calibration did not follow discrimination: the two highest-ranked models were systematically overconfident, whereas the random forest was well calibrated despite equivalent discrimination. SHAP attributions were dominated by filtered texture. Radiomic features gave moderate discrimination under calibrated, leakage-resistant, patient-clustered evaluation. Classifier differences were minimal, no ensemble improved on a transparent linear model, and predicted probabilities would need recalibration before use.