Classification of malignant/benign groups in lung cancer by machine learning and investigation of feature significance of parameters.
Authors
Affiliations (11)
Affiliations (11)
- Physiotherapy and Rehabilitation Master Program, Istanbul Bilgi University, Institute of Graduate Programs, Istanbul, Türkiye.
- Faculty of Management Sciences, Department of Management Information Systems, Bogazici University, Istanbul, Türkiye.
- Nuclear Physics Master Program, Istanbul University, Institute of Science and Technology, Istanbul, Türkiye.
- Vocational School, Physiotherapy Program, Istanbul Galata University, Istanbul, Türkiye.
- Department of Nuclear Medicine, Yedikule Chest Diseases Hospital, Istanbul, Türkiye.
- Faculty of Science, Department of Nuclear Physics, Istanbul University, Istanbul, Türkiye.
- Department of Nuclear Medicine, Istanbul Training and Research Hospital, Istanbul, Türkiye.
- Faculty of Medicine, Department of Chest Diseases, Altınbaş University, Istanbul, Türkiye.
- Faculty of Health Sciences, Physiotherapy and Rehabilitation Program, Istanbul Bilgi University, Istanbul, Türkiye.
- Vocational School of Health Sciences, Radiotherapy Program, Altınbaş University, Istanbul, Türkiye.
- Department of Arts & Sciences, Maine Maritime Academy, Castine, ME, USA.
Abstract
This study aimed to apply machine learning (ML) models to enhance lung cancer (LC) classification, distinguishing malignant from benign tumors, using data from 73 patients. The dataset included PET/CT biomarkers and demographic factors, with SMOTE applied to address class imbalance. Three models were evaluated using 10-fold cross-validation to compare Random Forest (RF), Decision Tree (DT), Extra Trees Classifier (ETC), and XGBoost algorithms based on accuracy and AUC. ETC performed best in Model 1 (86 % accuracy, AUC 0.95), RF in Model 2 (94 % accuracy, AUC 0.97), and XGBoost in Model 3 (94 % accuracy, AUC 0.98). XGBoost consistently outperformed others, particularly in Model 3, which included age and smoking. Feature importance analysis highlighted SUVmax as the most predictive variable, with smoking having a moderate influence and gender being minimal. Integrating clinical and lifestyle data with PET/CT parameters significantly improved LC classification. XGBoost emerged as the most effective model, demonstrating that comprehensive models enhance diagnostic accuracy beyond traditional metrics.