Evaluation of 2024 Turkish Medical Oncology Board Exam with ChatGPT
Journal of Oncological Science, cilt.11, sa.1, ss.36-39, 2025 (Scopus, TRDizin)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 11 Sayı: 1
- Basım Tarihi: 2025
- Doi Numarası: 10.37047/jos.galenos.2025.2024.106532
- Dergi Adı: Journal of Oncological Science
- Derginin Tarandığı İndeksler: Scopus, TR DİZİN (ULAKBİM)
- Sayfa Sayıları: ss.36-39
- Anahtar Kelimeler: clinical reasoning, large language model, medical board exams, Medical oncology
- Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
- Sağlık Bilimleri Üniversitesi Adresli: Evet
Özet
Objective: This study aims to assess ChatGPT-4’s performance on the Turkish Medical Oncology Board Exam questions, highlighting its potential uses and limitations in medical specialty evaluations. Material and Methods: ChatGPT-4 was presented with each question from the 2024 Turkish Medical Oncology Proficiency Exam. Answers were determined to be correct or incorrect by comparison with the official answer key. Results: The overall accuracy of ChatGPT-4.0 in this study was 64% out of 100 questions. For the fact-based questions (45 items), which require knowledge of specific information, such as molecules and side effects, ChatGPT-4o demonstrated an accuracy of 75.5%, with 34 correct responses. However, in the case-based questions (55 items) that require clinical judgment, its accuracy dropped to 54.5% (correct responses of 30). All these results highlight strengths of ChatGPT-4o on fact-driven questions but expose its limitations in scenarios needing nuanced decision-making. Conclusion: Oncological clinical decision-making necessitates a nuanced approach that extends beyond standardized guidelines, integrating individual patient variables such as medical history, comorbidities, and therapeutic responses. While artificial intelligence (AI) systems demonstrate proficiency in processing guideline-driven data, they exhibit limitations in contextual clinical judgment requiring physician expertise. This study observed ChatGPT-4’s superior performance on knowledge-based assessments (75.5% accuracy), attributable to its training on the American Society of Clinical Oncology/ the European Society for Medical Oncology frameworks. However, its accuracy declined significantly in case-based evaluations (54.5%), highlighting challenges in personalized care integration. These findings underscore the indispensable role of clinician judgment in navigating complex, individualized treatment landscapes. Enhancing AI’s clinical utility requires training on real-world patient data, though ethical constraints-particularly General Data Protection Regulation compliance-limit access to such datasets. Institution-specific AI tools leveraging anonymized records may bridge this gap, pending technological and regulatory advancements.