Can Gpt-4o Accurately Diagnose Trauma X-Rays? A Comparative Study with Expert Evaluations


Öztürk A., Günay S., Ateş S., Yiğit (Yavuz Yigit) Y.

Journal of Emergency Medicine, cilt.73, ss.71-79, 2025 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 73
  • Basım Tarihi: 2025
  • Doi Numarası: 10.1016/j.jemermed.2024.12.010
  • Dergi Adı: Journal of Emergency Medicine
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, EMBASE, MEDLINE
  • Sayfa Sayıları: ss.71-79
  • Anahtar Kelimeler: artificial intelligence, ChatGPT, GPT-4o, trauma, X-ray
  • Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Background: The latest artificial intelligence (AI) model, GPT-4o, introduced by OpenAI, can process visual data, presenting a novel opportunity for radiographic evaluation in trauma patients. Objective: This study aimed to assess the efficacy of GPT-4o in interpreting radiographs for traumatic bone pathologies and to compare its performance with that of emergency medicine and orthopedic specialists. Methods: The study involved 10 emergency medicine specialists, 10 orthopedic specialists, and the GPT-4o AI model, evaluating 25 cases of traumatic bone pathologies of the upper and lower extremities selected from the Radiopaedia website. Participants were asked to identify fractures or dislocations in the radiographs within 45 minutes. GPT-4o was instructed to perform the same task in 10 different chat sessions. Results: Emergency medicine specialists and orthopedic specialists demonstrated an average accuracy of 82.8% and 87.2%, respectively, in radiograph interpretation. In contrast, GPT-4o achieved an accuracy of only 11.2%. Statistical analysis revealed significant differences among the three groups (p < 0.001), with GPT-4o performing significantly worse than both groups of specialists. Conclusion: GPT-4o's ability to interpret radiographs of traumatic bone pathologies is currently limited and significantly inferior to that of trained specialists. These findings underscore the ongoing need for human expertise in trauma diagnosis and highlight the challenges of applying AI to complex medical imaging tasks.