Guideline-based, but not error-free: Multilingual risks in AI-powered patient counseling on gallstones
International Journal of Medical Informatics, cilt.212, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 212
- Basım Tarihi: 2026
- Doi Numarası: 10.1016/j.ijmedinf.2026.106341
- Dergi Adı: International Journal of Medical Informatics
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, BIOSIS, CINAHL, Compendex, EMBASE, INSPEC, MEDLINE
- Anahtar Kelimeler: Large language models, Digital health, Gallstone disease, Language bias, ChatGPT, Guideline concordance, Gemini
- Sağlık Bilimleri Üniversitesi Adresli: Evet
Özet
Background: Patients increasingly use large language models (LLMs) for health information, yet the guideline concordance and safety of patient-facing outputs—particularly across languages—remain uncertain. We evaluated three widely used LLM platforms (web interfaces) and their underlying default models for gallstone-related counseling in Turkish and English. Methods: In this cross-sectional content analysis, 14 real-world, guideline-mappable patient questions were developed in Turkish and translated into semantically equivalent English. Each question was submitted once to ChatGPT (ChatGPT-4o mini), Gemini (Gemini 3-flash), and Perplexity (Sonar family; default free-tier routing at the time of testing) in both languages under standardized conditions, yielding 84 responses. Two blinded hepatobiliary surgeons independently rated each response using a prespecified 3-point guideline concordance scale (0–2) mapped to EASL 2016 gallstone guidelines and Tokyo Guidelines 2018 for acute cholecystitis; disagreements were adjudicated by a third surgeon. Within-model language differences were assessed with Wilcoxon signed-rank tests; between-model comparisons used Friedman tests. Full correctness (score = 2) was analyzed using Cochran's Q with McNemar post-hoc tests. Error types and response length were also examined. Results: In English, model performance differed significantly, with ChatGPT and Gemini outperforming Perplexity (p < 0.01), while Turkish differences were not statistically significant. ChatGPT performed better in English than Turkish (p = 0.008). Error profiles were language-dependent: Turkish outputs more often showed under-explanation, whereas English outputs more frequently amplified risk. Perplexity demonstrated the highest overall error burden. . Conclusion: LLM responses to gallstone questions are often guideline-aligned but remain model- and language-sensitive, with clinically relevant safety risks. Multilingual evaluation standards are needed, and unsupervised reliance on LLMs for patient guidance—especially in low-resource languages—should be discouraged.