Evaluating Cervical Cancer Risk Using Machine Learning
Haseki Tip Bulteni, cilt.63, sa.4, ss.188-194, 2025 (ESCI, Scopus, TRDizin)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 63 Sayı: 4
- Basım Tarihi: 2025
- Doi Numarası: 10.4274/haseki.galenos.2025.35220
- Dergi Adı: Haseki Tip Bulteni
- Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus, CINAHL, EMBASE, Directory of Open Access Journals, TR DİZİN (ULAKBİM)
- Sayfa Sayıları: ss.188-194
- Anahtar Kelimeler: Cervical cancer, risk factors, machine learning, gynecology, diagnosis
- Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
- Sağlık Bilimleri Üniversitesi Adresli: Evet
Özet
Aim: Cervical cancer development is influenced by a complex interaction of socio-demographic, behavioral, and clinical factors, which can be systematically analyzed using large datasets. Therefore, this study aimed to evaluate the effectiveness of machine learning (ML) models applied to the University of California, Irvine (UCI), cervical cancer risk factors dataset in predicting cervical health outcomes and supporting early detection strategies. Methods: This study was designed as a retrospective data analysis covering a random sampling of patients between 2012 and 2013 who attended the gynecology service at Hospital Universitario de Caracas in Caracas, Venezuela. The publicly available UCI cervical cancer risk factors dataset was utilized for the analysis. A correlation heatmap was generated to explore the relationships among various risk factors. To address the class imbalance present in the dataset, the synthetic minority over-sampling technique (SMOTE) was applied. Subsequently, different ML classifiers were trained and evaluated to predict cervical cancer outcomes with improved accuracy. Results: The correlation analysis revealed strong correlations among smoking-related measures and diagnostic variables, indicating internal consistency. After applying SMOTE, the dataset achieved a balanced distribution of healthy and diseased individuals. The ensemble classifiers demonstrated high accuracy, up to 97%, and precision, with random forest and light gradient boosting machine performing particularly well. However, the recall for cancer detection was lower: 0.80, indicating potential missed diagnoses. Conclusion: The findings support the integration of ML in clinical diagnostics for cervical cancer, highlighting its potential for improving early detection and patient outcomes while also emphasizing the need for ongoing refinement in model performance.