Patient-Facing AI Chatbot Treatment-Direction Advice in Orthodontic Health Communication: A Scenario-Based Comparison with Expert Consensus
Healthcare (Switzerland), cilt.14, sa.16, 2026 (SCI-Expanded, SSCI, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 14 Sayı: 16
- Basım Tarihi: 2026
- Doi Numarası: 10.3390/healthcare14162565
- Dergi Adı: Healthcare (Switzerland)
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Social Sciences Citation Index (SSCI), Scopus, CINAHL, Health Research Premium Collection (ProQuest)
- Anahtar Kelimeler: artificial intelligence, large language model, chatbot, strategic health communication, AI-mediated health communication, orthodontics, patient education, treatment-direction advice, algorithmic ethics, trust and transparency
- Sağlık Bilimleri Üniversitesi Adresli: Evet
Özet
Background/Objectives: AI chatbots may shape patient expectations before professional consultation. This scenario-based first-response study evaluated whether four user-facing chatbots provided orthodontic treatment-direction advice concordant with an expert benchmark and whether responses contained safety, referral, or overconfidence concerns. Methods: Forty fictional Turkish patient-oriented scenarios across eight categories were independently coded by three orthodontists as clear aligners, fixed appliances, both options, examination required, or advanced specialist/surgical evaluation required. Each scenario was submitted once to ChatGPT, Claude, Copilot, and Gemini on 20 May 2026. Two independent non-author orthodontists coded 160 archived first responses using a predefined framework, with adjudication before analysis. Results: Inter-expert agreement was moderate (Fleiss kappa = 0.491; Gwet AC1 = 0.528). Under the majority benchmark, exact concordance was 82.5% for ChatGPT, 67.5% for Claude, 42.5% for Copilot, and 37.5% for Gemini (Cochran Q = 34.105, p < 0.001). The overall difference remained significant in the 17 unanimous scenarios (Q = 11.455, p = 0.010), but a post hoc alternative-reference analysis that adopted the dissenting expert code in the 23 non-unanimous scenarios attenuated the rates to 57.5%, 52.5%, 52.5%, and 42.5%, respectively (Q = 4.222, p = 0.238). Coded safety-concern rates ranged from 15.0% to 62.5%. Conclusions: The sampled first responses differed in treatment direction and safety coding, but estimates were sensitive to the expert reference definition. Under the tested single-date, single-language, and single-run conditions, the findings represent a conditional snapshot rather than a time-invariant ranking of model capability. Patient-facing chatbots should support nondirective pre-consultation education and referral, not autonomous appliance selection.