Effect of ChatGPT-Assisted Reflective Reasoning on Guideline-Concordant Procedural Decision-Making Among Early-Career Interventional Radiologists


Yasar Y., Demir M., Canturk A., Ozyilmaz S., Turgan A. H., Agackaya Y.

Academic Radiology, cilt.33, sa.4, ss.1577-1582, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 33 Sayı: 4
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1016/j.acra.2026.01.017
  • Dergi Adı: Academic Radiology
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, EMBASE, MEDLINE
  • Sayfa Sayıları: ss.1577-1582
  • Anahtar Kelimeler: Interventional radiology, Clinical decision support, Large language models, ChatGPT, Simulation-based training
  • Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Rationale and Objectives This study aims to evaluate the effect of ChatGPT-assisted reflective reasoning on guideline-concordant procedural decision-making among early-career interventional radiologists using standardized clinical scenarios based on the American College of Radiology Appropriateness Criteria. Materials and Methods This prospective simulation-based study included 128 scenarios across common interventional radiology indications. Two expert interventional radiologists served as the reference standard. Three early-career radiologists completed all scenarios twice: first independently (pre-ChatGPT) and, after a two-month washout period, with access to ChatGPT-generated reasoning before recording final decisions (post-ChatGPT). Guideline concordance was assessed using a three-tier scoring system (appropriate = 2, may be appropriate = 1, inappropriate = 0) and a binary score reflecting avoidance of inappropriate decisions. Predifferences and postdifferences were analyzed with Wilcoxon signed-rank and McNemar tests. Agreement with experts was measured using Cohen’s kappa. Results ChatGPT-assisted reflective reasoning significantly improved guideline-concordant decision-making. The mean detailed compliance score increased from 1.697 to 1.900, and minimal compliance enhanced from 90.89% to 98.70%. A total of 30 scenario-level corrections shifted from inappropriate to guideline-concordant selections (McNemar χ² = 27.03; p < 0.0001). Detailed compliance improved significantly for all radiologists (p < 0.01). Weighted Cohen’s kappa increased from 0.08–0.13 to 0.21–0.30, indicating better agreement with expert consensus. Performance variability decreased, narrowing the gap between early-career radiologists and experts. Conclusion ChatGPT-assisted reflective reasoning enhanced guideline alignment and reduced inappropriate procedural selections among early-career interventional radiologists. These findings support the role of large language models as cognitive support tools during early clinical practice and warrant prospective evaluation in real-world settings.