Evaluation of 2024 Turkish Medical Oncology Board Exam with ChatGPT


Creative Commons License

Kahraman E. G., ÜNAL O. Ü.

Journal of Oncological Science, cilt.11, sa.1, ss.36-39, 2025 (Scopus, TRDizin)

Özet

Objective: This study aims to assess ChatGPT-4’s performance on the Turkish Medical Oncology Board Exam questions, highlighting its potential uses and limitations in medical specialty evaluations. Material and Methods: ChatGPT-4 was presented with each question from the 2024 Turkish Medical Oncology Proficiency Exam. Answers were determined to be correct or incorrect by comparison with the official answer key. Results: The overall accuracy of ChatGPT-4.0 in this study was 64% out of 100 questions. For the fact-based questions (45 items), which require knowledge of specific information, such as molecules and side effects, ChatGPT-4o demonstrated an accuracy of 75.5%, with 34 correct responses. However, in the case-based questions (55 items) that require clinical judgment, its accuracy dropped to 54.5% (correct responses of 30). All these results highlight strengths of ChatGPT-4o on fact-driven questions but expose its limitations in scenarios needing nuanced decision-making. Conclusion: Oncological clinical decision-making necessitates a nuanced approach that extends beyond standardized guidelines, integrating individual patient variables such as medical history, comorbidities, and therapeutic responses. While artificial intelligence (AI) systems demonstrate proficiency in processing guideline-driven data, they exhibit limitations in contextual clinical judgment requiring physician expertise. This study observed ChatGPT-4’s superior performance on knowledge-based assessments (75.5% accuracy), attributable to its training on the American Society of Clinical Oncology/ the European Society for Medical Oncology frameworks. However, its accuracy declined significantly in case-based evaluations (54.5%), highlighting challenges in personalized care integration. These findings underscore the indispensable role of clinician judgment in navigating complex, individualized treatment landscapes. Enhancing AI’s clinical utility requires training on real-world patient data, though ethical constraints-particularly General Data Protection Regulation compliance-limit access to such datasets. Institution-specific AI tools leveraging anonymized records may bridge this gap, pending technological and regulatory advancements.