Back to all papers

H-QDCT: hierarchical quantum DCT for structural-textural feature fusion in medical imaging.

July 31, 2026pubmed logopapers

Authors

Aragones XF,Ballester MÁG

Affiliations (5)

  • BCN Medtech, Department of Engineering, Universitat Pompeu Fabra, Barcelona, Spain. [email protected].
  • Parc Tecnològic TecnoCampus Mataró-Maresme, Universitat Pompeu Fabra, Mataró, Spain. [email protected].
  • BCN Medtech, Department of Engineering, Universitat Pompeu Fabra, Barcelona, Spain.
  • Barcelona Supercomputing Center, Barcelona, Spain.
  • ICREA, Barcelona, Spain.

Abstract

Deep learning models for medical image classification typically rely on millions of parameters or extensive pretraining, limiting deployment in resource-constrained settings. We investigate whether a hybrid quantum-classical architecture can achieve competitive diagnostic performance with extreme parameter efficiency and superior stability compared to lightweight classical counterparts. We propose the hierarchical quantum discrete cosine transform (H-QDCT), a lightweight pipeline combining classical patch-based preprocessing with quantum frequency-domain feature extraction. Images are partitioned into patches, amplitude-embedded into quantum states, and transformed via a quantum DCT that reorganizes information by spatial frequency. A hierarchical variational ansatz processes structural (low-frequency) and textural (high-frequency) components in disjoint qubits subspaces before global fusion. We evaluate H-QDCT on six clinically binarized MedMNIST datasets under grayscale constraints against both heavy and lightweight classical baselines. H-QDCT achieves competitive performance with only 1726 parameters, exceeding a 99.98% reduction compared to ResNet-18 (11.2M parameters). On PneumoniaMNIST, the model attains an AUC of 0.90, approaching the 0.93 of pretrained ResNet-18. Crucially, H-QDCT demonstrates superior robustness in data-scarce regimes: While the lightweight vision transformer (Tiny-ViT) collapsed to the majority class on DermaMNIST (F1 <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo>=</mo> <mn>0.00</mn></mrow> </math> ) and BloodMNIST (F1 <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo>=</mo> <mn>0.06</mn></mrow> </math> ), H-QDCT retained learnability on BloodMNIST (F1 up to 0.22) with stable convergence. Beyond binary screening, the architecture extends natively to multi-class differential diagnosis (from <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>K</mi> <mo>=</mo> <mn>2</mn></mrow> </math> to <math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>K</mi> <mo>=</mo> <mn>9</mn></mrow> </math> ) without mode collapse, and on PneumoniaMNIST and PathMNIST its compact encoder surpasses a from-scratch ViT-B/16 ( <math xmlns="http://www.w3.org/1998/Math/MathML"><mo>∼</mo></math> 86 M parameters) in macro-F1 while using under 1,000 trainable parameters. The architecture operates entirely on grayscale inputs, confirming its reliance on morphological semantics rather than color bias. H-QDCT demonstrates that quantum spectral feature extraction achieves clinically meaningful classification with orders-of-magnitude parameter reduction. By mitigating the convergence instability inherent in small-scale classical transformers, H-QDCT establishes a robust design principle for compact, frequency-aware medical diagnostics on near-term quantum hardware.

Topics

Journal Article

Ready to Sharpen Your Edge?

Subscribe to join 11k+ peers who rely on RadAI Slice. Get the essential weekly briefing that empowers you to navigate the future of radiology.

We respect your privacy. Unsubscribe at any time.