Do AI Chatbots Tell the Truth About Dentin Hypersensitivity? A Comparative Evaluation of Quality, Accuracy, and Readability
Cumhuriyet Dental Journal, cilt.29, sa.1, ss.138-147, 2026 (Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 29 Sayı: 1
- Basım Tarihi: 2026
- Doi Numarası: 10.7126/cumudj.1848545
- Dergi Adı: Cumhuriyet Dental Journal
- Derginin Tarandığı İndeksler: Scopus, Directory of Open Access Journals
- Sayfa Sayıları: ss.138-147
- Anahtar Kelimeler: Artificial Intelligence, CLEAR, CLEAR, Dentin hassasiyeti, dentin hypersensitivity, DISCERN, DISCERN, modified Global Quality Score, modifiye Global Kalite Skoru, okunabilirlik, readability, yapay zekâ
- Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
- Sağlık Bilimleri Üniversitesi Adresli: Evet
Özet
Objectives: Dentin hypersensitivity (DH) is a common dental complaint, and many patients now seek information from Al chatbots. Yet, the accuracy, reliability, and readability of chatbot-generated DH content remain uncertain. Materials and Methods: A consensus-based DH question set was presented to three AI chatbots (ChatGPT-4o, DeepSeek, Copilot) in independent, standardized sessions. Three blinded periodontologists evaluated the responses using CLEAR, mGQS, accuracy scores, DISCERN, and readability metrics 仆RE, FKGL). Non-parametric tests compared inter-model differences. Results: Inter-group comparisons revealed statistically significant variations in FKGL (p = 0.025), DISCERN (p = 0.004), and the length of generated responses (p < 0.001). Copilot yielded the highest reliability and quality, DeepSeek produced the most readable content, and ChatGPT showed the most significant variability. Copilot also had the highest proportion of fully accurate, high-quality answers, whereas low-accuracy output occurred only with ChatGPT. Strong correlations were noted among accuracy, completeness, and overall quality. Conclusions: AI chatbots can generate clinically relevant DH information, but performance varies. Copilot showed the best balance of accuracy and reliability, DeepSeek provided the most accessible language, and ChatGPT demonstrated inconsistent results. Clinician oversight remains essential when using AI-generated content for patient education.