Back to all papers

Statistically Rigorous and Interpretable Radiomics for Benign-Malignant Classification on the BUSI-WHU Breast Ultrasound Dataset.

September 15, 2026pubmed logopapers

Authors

Antu PS,Rashid MM,Ali MH

Affiliations (2)

  • Department of Statistics, Mawlana Bhashani Science and Technology University, Tangail, Bangladesh. [email protected].
  • Department of Statistics, Mawlana Bhashani Science and Technology University, Tangail, Bangladesh.

Abstract

Radiomics has shown promise for breast ultrasound lesion classification, yet reported performance is often affected by patient-level leakage, feature selection bias, dependence between lesions from the same patient, and image formats lacking pixel-spacing metadata. The BUSI-WHU dataset allows these issues to be addressed together. We analyzed 927 breast lesions from 816 patients using 939 handcrafted radiomic features from expert-annotated images. Pixel spacing was recovered from dataset-supplied scale factors and images resampled to a common isotropic resolution. Features were assessed by Mann-Whitney U tests with Benjamini-Hochberg correction and Cliff's delta, with a one-lesion-per-patient sensitivity analysis. Six classifiers, including a penalized logistic regression baseline, were evaluated under ten-fold patient-grouped, class-stratified cross-validation with correlation filtering and Boruta selection inside each training fold. Confidence intervals and paired comparisons used patient-level bootstrap of out-of-fold predictions, and SHAP was computed on held-out patients. Of the 939 features, 587 differed significantly after correction; most remained significant under one-lesion-per-patient sampling; and 48 were selected consistently across folds. Pooled out-of-fold ROC-AUCs ranged from 0.828 to 0.834; no pairwise difference survived correction, and the logistic regression performed comparably to the best ensemble. Discrimination was unchanged when absolute size features were removed. Calibration did not follow discrimination: the two highest-ranked models were systematically overconfident, whereas the random forest was well calibrated despite equivalent discrimination. SHAP attributions were dominated by filtered texture. Radiomic features gave moderate discrimination under calibrated, leakage-resistant, patient-clustered evaluation. Classifier differences were minimal, no ensemble improved on a transparent linear model, and predicted probabilities would need recalibration before use.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.