Back to all papers

CNN Sample-Size Effects Across Biomedical Datasets: A Reliability Pattern in Overfitting, Ranking, and Monotonicity.

August 25, 2026pubmed logopapers

Authors

Sgarro GA,Mendikowski M,Santoro D,Grilli L

Affiliations (4)

  • Department of Economics, University of Foggia, 71121 Foggia, Italy.
  • German Research Center for Artificial Intelligence (DFKI), 23562 Lübeck, Germany.
  • Computer Science Department, University of Hamburg, 22527 Hamburg, Germany.
  • Department of Economics, Statistics and Business, Faculty of Technological and Innovation Sciences, Universitas Mercatorum, 00186 Rome, Italy.

Abstract

Convolutional neural networks (CNNs) are widely used for biomedical image classification, yet it remains unclear under which conditions training on reduced subsets of available data can provide reliable guidance during model development, how much training data is required to achieve stable and comparable performance across CNN architectures, and whether increasing the training set size always leads to improved generalization or can sometimes result in degraded performance. We study this question across four biomedical datasets (breast mammography, pediatric chest X-ray, brain tumor MRI, and skin lesion dermoscopy) using the full grid of 39 CNN architectures (1-3 convolutional layers, 16/32/64 filters) from our companion architectural study, training each configuration from scratch on seven proportions of the training data (5%, 10%, 20%, 40%, 60%, 80%, and 100%) over 5 independent runs per configuration, with the test set held at a fixed size across all sample-size conditions to ensure a like-for-like comparison of generalization performance. The analysis investigates three complementary aspects of sample-size sensitivity: the stabilization of the training-test generalization gap as training-set size increases, the reliability of architecture rankings obtained from reduced training fractions as a proxy for the full-dataset ranking, and the monotonicity of test performance with respect to training-set size. Taken together, the results point to a rough four-band pattern of reliability across the sampled fractions-unstable below 20% of the training set, of uncertain overfitting status between 20% and 60%, comparatively stable between 60% and 80%, and potentially counterproductive beyond 80%-while showing that this pattern is itself dataset-dependent and offers no guarantee on architecture ranking, arguing against reduced-fraction screening as a reliable shortcut for CNN architecture selection in biomedical imaging. All code and datasets are publicly released for reproducibility.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAISlice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.