Alignment between AI clinical decision tools and multidisciplinary tumor board decisions in prostate cancer
International Urology and Nephrology, cilt.58, sa.10, ss.4131-4138, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 58 Sayı: 10
- Basım Tarihi: 2026
- Doi Numarası: 10.1007/s11255-026-05178-1
- Dergi Adı: International Urology and Nephrology
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, BIOSIS, EMBASE, MEDLINE, Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest), Pharma Collection (ProQuest)
- Sayfa Sayıları: ss.4131-4138
- Anahtar Kelimeler: Prostate cancer, Artificial intelligence, Large language models, Multidisciplinary tumor board, Clinical decision support
- Sağlık Bilimleri Üniversitesi Adresli: Evet
Özet
Purpose: This study aimed to evaluate the concordance between treatment recommendations generated by LLMs and decisions made by a multidisciplinary uro-oncology tumor board. Methods: Forty-eight consecutive prostate cancer cases previously discussed at a multidisciplinary tumor board were retrospectively analyzed. For each case, treatment recommendations were generated using five LLM platforms (ChatGPT-4o, ChatGPT, Perplexity, Copilot, and DeepSeek) based on standardized clinical summaries. Four independent urology specialists evaluated the concordance between LLM recommendations and tumor board decisions using a 5-point Likert scale. Differences among models were assessed using the Friedman test followed by Bonferroni-corrected Wilcoxon signed-rank tests. Inter-rater agreement was calculated using the intraclass correlation coefficient. Results: Significant differences in concordance were observed among the evaluated AI platforms (χ2 = 32.16, p < 0.001). Perplexity and ChatGPT-4o demonstrated the highest alignment with tumor board decisions, each achieving a median Likert score of 4.75, whereas Copilot showed the lowest concordance (median 3.00). DeepSeek and ChatGPT demonstrated intermediate performance. Post hoc analyses revealed that Perplexity significantly outperformed several lower-performing platforms; however, no statistically significant difference was observed between Perplexity and ChatGPT-4o (p = 0.149). Expert evaluations showed strong inter-rater agreement (ICC = 0.82). Conclusion: Large language models can demonstrate substantial concordance with multidisciplinary tumor board decisions in prostate cancer management. However, variability among models and the risk of hallucinated information indicate that LLMs should function as clinical decision-support tools under expert supervision rather than as autonomous decision-makers.