Comparative assessment of ChatGPT and Gemini answers to common chronic obstructive pulmonary disease questions: An expert panel evaluation by pulmonologists


Güçsav M. O., Serçe Unat D., Akçay O., Unat Ö. S., Ayrancı A., Erbaycu A. E.

Chronic Respiratory Disease, cilt.23, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 23
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1177/14799731261443321
  • Dergi Adı: Chronic Respiratory Disease
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, EMBASE, MEDLINE, Directory of Open Access Journals, Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest)
  • Anahtar Kelimeler: generative artificial intelligence, chronic obstructive pulmonary disease, telemedicine, hallucinations, health education
  • Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Background: AI-based chatbots are increasingly used as sources of health information. However, their reliability in delivering accurate and scientifically sound responses to patient questions remains uncertain, especially in chronic diseases such as chronic obstructive pulmonary disease (COPD). This study aims to compare the reliability of ChatGPT-4o and Gemini 2.5 Flash in providing patient-centered medical information on COPD. Methods: A total of 34 common public questions about COPD were submitted to ChatGPT-4o and Gemini 2.5 Flash. Responses were evaluated blindly by four pulmonologists across three domains: accuracy, clarity, and scientific adequacy. The mean scores and word counts were analyzed and compared via nonparametric tests. Results: Gemini 2.5 Flash outperforms ChatGPT-4o in terms of scientific adequacy (mean score: 4.69 ± 0.31 vs. 4.34 ± 0.45, p<0.001). No significant difference was found in accuracy or clarity. The Gemini 2.5 Flash also generated significantly longer responses, particularly in the treatment and prognosis domains (p<0.001). Both models provided generally acceptable answers, but ChatGPT-4o′s responses were shorter and occasionally less complete. Conclusions: While both models delivered largely accurate and understandable content, Gemini 2.5 Flash tended to produce more detailed responses and received higher scientific adequacy ratings; however, this difference should be interpreted in light of the substantial imbalance in response length. These tools may support patient education however, the findings reflect a comparison between AI systems only and should be interpreted within this scope.