Evaluating the effectiveness of chatbots and traditional resources in patient education on dry eye disease
Clinical and Experimental Optometry, cilt.109, sa.2, ss.182-186, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 109 Sayı: 2
- Basım Tarihi: 2026
- Doi Numarası: 10.1080/08164622.2025.2517750
- Dergi Adı: Clinical and Experimental Optometry
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, EMBASE, MEDLINE, Academic Search Ultimate (EBSCO), Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest)
- Sayfa Sayıları: ss.182-186
- Anahtar Kelimeler: Optic coherence tomography, tear meniscus, tear meniscus depth, tear meniscus height, vitamin D insufficiency
- Sağlık Bilimleri Üniversitesi Adresli: Evet
Özet
Clinical Relevance: Artificial intelligence chatbots demonstrate potential as valuable educational resources for patients with dry eye disease, offering complementary information to established medical platforms. Background: The increasing prevalence of dry eye disease necessitates reliable and comprehensible patient information resources. This study evaluates and compares the quality of information provided by contemporary AI chatbots with established ophthalmological sources. Methods: Three leading AI chatbots (ChatGPT-3.5, Gemini, and Llama) and the American Academy of Ophthalmology (AAO) website were systematically evaluated using 20 common patient questions about dry eye disease. Responses were assessed for accuracy using the Structure of Observed Learning Outcome (SOLO) taxonomy, understandability and actionability using the Patient Education Materials Assessment Tool (PEMAT), and linguistic accessibility using Flesch-Kincaid readability metrics. Results: Gemini demonstrated superior understandability with a mean PEMAT-U score of 73.4 ± 11.4, significantly higher than ChatGPT (65.4 ± 10.6), Llama (63.4 ± 10.3), and AAO (52.5 ± 19.3) (p < 0.001). No significant differences were observed in actionability scores (p = 0.120). The AAO website exhibited the highest reading ease score (50.4 ± 17.9, p = 0.015). For accuracy assessment, ChatGPT achieved the highest mean SOLO score (3.4 ± 0.7), followed closely by Gemini (3.3 ± 0.8), with no significant performance differences detected among chatbots (p = 0.574). No instances of incorrect or potentially harmful information were identified across any evaluated source. Conclusion: While AI chatbots demonstrate promising capabilities for patient education in dry eye disease, particularly in providing comprehensive and understandable information, their higher linguistic complexity presents a potential accessibility barrier. Future development should focus on enhancing readability while maintaining comprehensive content, positioning chatbots as valuable complements to–rather than replacements for–professional medical consultation.