Evaluation of ChatGPT-4o’s Responses to Questions about Myasthenia Gravis in English and Turkish ChatGPT-4o’nun Miyastenia Gravis Hakkındaki İngilizce ve Türkçe Sorulara Verdiği Yanıtların Değerlendirilmesi


Creative Commons License

İNAN B., KARADAŞ Ö., ODABAŞI Z.

Medical Journal of Bakirkoy, cilt.21, sa.3, ss.310-315, 2025 (ESCI, Scopus, TRDizin)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 21 Sayı: 3
  • Basım Tarihi: 2025
  • Doi Numarası: 10.4274/bmj.galenos.2025.2025.5-9
  • Dergi Adı: Medical Journal of Bakirkoy
  • Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus, CINAHL, EMBASE, TR DİZİN (ULAKBİM)
  • Sayfa Sayıları: ss.310-315
  • Anahtar Kelimeler: Artificial intelligence, ChatGPT-4o, large language models, myasthenia gravis, neurology
  • Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
  • Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Objective: Large language models, such as Chat Generative Pre-Trained Transformer 4o (ChatGPT-4o), are increasingly used by both patients and medical professionals to access health-related information. Myasthenia gravis (MG) is a chronic autoimmune neuromuscular disorder requiring long-term treatment. Therefore, timely access to accurate medical information about MG is important. This study aimed to evaluate the accuracy, completeness, clarity, appropriateness for the target audience, risk of misinformation or harm, and readability of ChatGPT-4o-generated responses to queries about MG from patients and neurology residents, in both English and Turkish. Methods: We developed four sets of 20 questions, frequently asked by patients and neurology residents about MG in both English and Turkish, covering pathophysiology and symptoms, diagnosis, treatment, prognosis, and daily management. ChatGPT-4o responses were generated in separate sessions on March 29, 2025. Two neurologists independently evaluated the responses using a 5-point Likert scale across five domains. Readability was assessed using the Flesch Reading Ease score, Flesch-Kincaid grade level, and Gunning-Fog index for English, and the Ateşman readability index for Turkish. Results: Scores for accuracy, clarity, appropriateness, and risk of misinformation or harm were consistently above 4 in both languages, with clarity rated as 5 in all responses. Completeness received the lowest scores (3.5-5.0), particularly in Turkish responses to resident-directed questions. Readability was higher in Turkish. English responses to resident queries were extremely difficult to read, while patient-directed ones remained in the “difficult” to “very difficult” range. Several discrepancies were observed in specific contents between English and Turkish outputs, such as differences in differential diagnosis lists, treatment options, contraindicated medications, and thymectomy indications. Conclusion: ChatGPT-4o produced high-quality responses overall to MG-related queries in both languages. However, language-specific inconsistencies and content omissions highlight the need for further model refinement, particularly in multilingual and professional-use contexts.