Comparative Performance and Utility of Large Language Models in Generating Psychological Screening Checklists for Bariatric Surgery


Çalışkan Y. K., BAŞAK F.

Bariatric Surgical Practice and Patient Care, 2026 (SCI-Expanded, SSCI, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1177/2168023x261481848
  • Dergi Adı: Bariatric Surgical Practice and Patient Care
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Social Sciences Citation Index (SSCI), Scopus, CINAHL, EMBASE, Health Research Premium Collection (ProQuest)
  • Anahtar Kelimeler: large language models (LLMs), artificial intelligence, psychological assessment, bariatric surgery, checklist, preoperative evaluation
  • Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Background: Psychological assessment is essential for bariatric surgery candidacy but remains inconsistent and time-consuming. This study compared three large language models (LLMs) in generating structured psychological screening checklists to support standardized preoperative evaluation. Methods: Models A, B, and C generated checklists for five standardized bariatric vignettes. Three blinded psychologists independently rated the outputs on 5-point Likert scales for clinical relevance, completeness, specificity, and usability. Statistical analysis included Jaccard similarity to quantify content overlap and analysis of variance with Tukey post hoc tests to compare expert ratings. Results: Model A produced the most extensive checklists (mean = 35.6 items, standard deviation = 4.2), while Models B and C were more concise (28.4 and 26.7 items, respectively). Model A achieved the highest completeness scores, whereas Model B was rated highest for specificity and usability. Interrater agreement was good to excellent (intraclass correlation coefficient = 0.79–0.87). Moderate content overlap (Jaccard similarity = 0.45–0.63) suggested complementary model strengths. Between-model differences were significant for completeness, F(2, 6) = 11.5, p = 0.008, and specificity, F(2, 6) = 9.8, p = 0.013. Conclusions: LLMs differ significantly in checklist breadth and focus. Strategic model selection can enhance preoperative workflows by balancing comprehensiveness with targeted, actionable assessment. However, deployment should remain clinician-supervised because outputs may vary with model updates, prompt phrasing, and local workflow constraints.