Large Language Models and Male Circumcision: A Reliability Assessment
Haseki Tip Bulteni, cilt.63, sa.3, ss.123-127, 2025 (ESCI, Scopus, TRDizin)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 63 Sayı: 3
- Basım Tarihi: 2025
- Doi Numarası: 10.4274/haseki.galenos.2025.79663
- Dergi Adı: Haseki Tip Bulteni
- Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus, CINAHL, EMBASE, Directory of Open Access Journals, TR DİZİN (ULAKBİM)
- Sayfa Sayıları: ss.123-127
- Anahtar Kelimeler: Male circumcision, large language models, patient education
- Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
- Sağlık Bilimleri Üniversitesi Adresli: Evet
Özet
Aim: Male circumcision remains routine in some countries for neonatal or religious reasons; however, it continues to be the subject of ongoing debate concerning its health benefits, potential risks, and implications for bodily autonomy. This study aims to evaluate the reliability of patient-facing content generated by four widely used large language models (LLMs) on various aspects of male circumcision. Methods: A search regarding LLMs was conducted using 20 standardized questions on 10 May 2025. Responses from ChatGPT, Copilot, Gemini, and Perplexity were evaluated by three independent experts. Inter-rater reliability was assessed with the intraclass correlation coefficient, and model performance differences were analyzed using Kruskal-Wallis tests with Bonferroni correction. Results: Inter-rater reliability was strong, with an intraclass correlation coefficient of 0.79 (p<0.001). Perplexity demonstrated statistically significant lower performance compared to ChatGPT, Copilot, and Gemini when evaluated across the thematic domains (p<0.001). Similarly, Perplexity performed statistically significantly worse than the other models across the criteria of clarity, structure, utility, and factual accuracy (p<0.001). Conclusion: Gemini and Copilot were the top performers across both thematic domains and evaluation criteria, highlighting substantial differences among LLMs in their ability to provide accurate and well-structured medical information regarding male circumcision. While ChatGPT shows promise for patient guidance, the inconsistent performance of models such as Perplexity highlights the need for cautious implementation and continuous oversight in healthcare communication.