Comparison of large language models in the management of pediatric ureteropelvic junction obstruction: a comparative analysis of 125 clinical scenarios


Baloğlu İ. H., Karlı G., Çekmece A. E., Özgür M. Ö., Albayrak A. T., Günay K. C., ...Daha Fazla

Pediatric Surgery International, cilt.42, sa.1, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 42 Sayı: 1
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1007/s00383-026-06547-8
  • Dergi Adı: Pediatric Surgery International
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, EMBASE, MEDLINE, Academic Search Ultimate (EBSCO), Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest)
  • Anahtar Kelimeler: Artificial intelligence, Large language models, Ureteropelvic junction obstruction, Pediatric urology
  • Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Purpose: To evaluate and compare the clinical accuracy, reliability, comprehensiveness and readability of three prominent Large Language Models (LLMs) (ChatGPT, Gemini, and Copilot) in the management of pediatric ureteropelvic junction obstruction. Methods: A total of 125 unique clinical scenarios representing various stages and complexities of pediatric ureteropelvic junction obstruction(UPJO) were developed. Responses from ChatGPT (GPT-5.3 version, OpenAI), Gemini(Google), and Copilot GPT 5 (Microsoft) were independently evaluated by two expert pediatric urologists across four domains: Clinical Accuracy, Completeness, Reliability, and Readability (scored 1–5). Initial inter-rater discrepancies were resolved through a consensus-building process. Statistical analysis included Kruskal-Wallis tests for performance comparison and weighted Cohen’s kappa for inter-rater agreement. Results: ChatGPT demonstrated significantly higher scores in clinical accuracy (4.13 +/- 1.31) and reliability (4.08 +/- 1.30) compared to its counterparts, showing the closest alignment with international guidelines. Conversely, Copilot achieved the highest readability score (4.13 +/- 0.65) but exhibited a “readability-accuracy paradox,” where professional formatting masked frequent clinical inaccuracies (3.49 +/- 1.67). Gemini provided comprehensive content but was hindered by structural deficits and the lowest readability score (2.88 ± 0.83). The expert consensus process successfully improved inter-rater agreement (κ = 0.532) to (κ = 0.586). Conclusion: While LLMs show promise as decision-support tools in pediatric surgery, their performance is inconsistent. ChatGPT is currently the most robust model for guideline-based management of UPJO. However, the “deceptive confidence” of models like Copilot poses a risk of misinformation. Future integration should explore multimodal capabilities, including the analysis of imaging and ongoing validation against standardized reporting frameworks.