Task-specific versus general-purpose AI models in ECG analysis: A comparative study with emergency medicine specialists


ALTINBİLEK E., Az A., SÖĞÜT Ö., Dogan Y., Akdemir T., BELEN E., ...Daha Fazla

American Journal of Emergency Medicine, cilt.95, ss.220-226, 2025 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 95
  • Basım Tarihi: 2025
  • Doi Numarası: 10.1016/j.ajem.2025.06.068
  • Dergi Adı: American Journal of Emergency Medicine
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, EMBASE, MEDLINE
  • Sayfa Sayıları: ss.220-226
  • Anahtar Kelimeler: Electrocardiogram, Artificial intelligence, Emergency medicine, Diagnostic accuracy, Domain-specific AI
  • Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Purpose: To evaluate and compare the diagnostic accuracy of three Artificial intelligence (AI) models—GPT-4o, Canva-GPT, and ECG Reader-GPT—against emergency medicine specialists (EMSs) in electrocardiogram (ECG) interpretation using a standardized and validated test set. Methods: In this prospective diagnostic accuracy study, 50 ECG questions were selected from the reference text 150 ECG Cases. Thirty EMSs completed the test once; each AI model was evaluated on the same test set daily over 30 consecutive days. Diagnostic accuracy was compared across predefined ECG subcategories and clinical case types. Results: EMSs achieved the highest overall diagnostic accuracy (median: 41.5; IQR: 37.0–43.0), followed closely by ECG Reader-GPT (median: 39.5; IQR: 39.0–41.0), with no statistically significant difference between them (p = 0.530). ECG Reader-GPT significantly outperformed both GPT-4o and Canva-GPT across all case categories (p < 0.001). Subgroup analysis revealed that ECG Reader-GPT performed comparably to EMSs in identifying ischemic syndromes, channelopathies and genetic syndromes, and normal ECGs (all p < 0.05); it surpassed them in interpreting rhythm disorders (p = 0.007). Conclusion: ECG Reader-GPT, a customized AI model for ECG interpretation, demonstrated diagnostic accuracy comparable to experienced EMSs and significantly outperformed general-purpose models across all ECG subcategories. These findings highlight the value of domain specialization in developing clinically effective AI tools.