Patient Facing AI Chatbots in Digital Orthodontic Education: A Comparative Evaluation of Understandability, Actionability, Quality, and Safety of Dietary Advice
Healthcare (Switzerland), cilt.14, sa.14, 2026 (SCI-Expanded, SSCI, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 14 Sayı: 14
- Basım Tarihi: 2026
- Doi Numarası: 10.3390/healthcare14142192
- Dergi Adı: Healthcare (Switzerland)
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Social Sciences Citation Index (SSCI), Scopus, CINAHL, Health Research Premium Collection (ProQuest)
- Anahtar Kelimeler: artificial intelligence, digital health, patient education, large language models, ChatGPT, Gemini, orthodontics, health literacy, actionability, AI safety
- Sağlık Bilimleri Üniversitesi Adresli: Evet
Özet
Background/Objectives: Patient-facing artificial intelligence chatbots are increasingly used as informal digital health education tools. In orthodontics, eating and drinking advice may directly affect appliance integrity, oral hygiene, enamel demineralization, caries risk, and clear aligner use. This study compared the understandability, actionability, overall quality, and potential harmfulness of responses generated by free and paid versions of ChatGPT and Gemini to patient-oriented orthodontic dietary questions. Methods: This cross-sectional comparative observational study evaluated 160 responses generated from 40 Turkish patient-oriented orthodontic eating and drinking questions across five clinically relevant categories. Each question was submitted separately to ChatGPT free version, ChatGPT Plus, Gemini free version, and Gemini Pro on 1 May 2026, using newly opened independent chat sessions without prompt engineering, follow-up prompts, response regeneration, or manual editing. Responses were anonymized, randomly coded, and independently evaluated by two specialist dentists. PEMAT-P actionability was defined as the primary outcome. PEMAT-P understandability, Global Quality Score, and potentially harmful advice classification were secondary outcomes. Repeated-measures comparisons were performed using Friedman tests and Bonferroni-adjusted Wilcoxon signed-rank tests. Results: PEMAT-P understandability was high across all groups, with median scores of 100.0 in every group. Significant group differences were found for understandability, actionability, and Global Quality Score. ChatGPT Plus achieved the highest actionability score and Global Quality Score and produced no responses classified as potentially harmful. Potentially harmful responses were identified in ChatGPT free version, Gemini free version, and Gemini Pro. For the primary outcome, PEMAT-P actionability, the overall group difference was statistically significant with a small effect size (χ2 = 20.527, p < 0.001, Kendall’s W = 0.171), while the largest effect was observed for GQS with a moderate effect size (χ2 = 46.520, p < 0.001, Kendall’s W = 0.388). Conclusions: All chatbot groups generated highly understandable responses; however, actionability, overall quality, and safety varied across systems. ChatGPT Plus showed the strongest overall performance under the specific interface, subscription, language, and date conditions tested; however, this finding should be interpreted as a time-specific benchmark rather than evidence that paid chatbot systems are intrinsically safer or more clinically reliable. Structured evaluation, transparent reporting, digital health equity considerations, and professional oversight remain necessary before AI-generated orthodontic dietary advice can be integrated into routine patient education.