Pre-deployment safety and governance assessment of LLM-based clinical decision support systems: A health technology assessment-oriented evaluation framework


Karataş S. Ş., Öner S. F.

Health Policy and Technology, cilt.15, sa.9, 2026 (SSCI, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 15 Sayı: 9
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1016/j.hlpt.2026.101281
  • Dergi Adı: Health Policy and Technology
  • Derginin Tarandığı İndeksler: Social Sciences Citation Index (SSCI), Scopus, EMBASE
  • Anahtar Kelimeler: Large language models, Clinical decision support, Patient safety, Risk assessment, Health technology assessment
  • Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Background: The clinical integration of large language models (LLMs) as decision support components has raised substantial concerns regarding patient safety, governance, and risk management. Despite increasing interest in their use, there is a lack of robust, domain-adaptable evaluation frameworks to characterize the clinical readiness and safety-related output behavior of LLM-based decision support systems prior to real-world deployment. Objective: This study aimed to develop and apply a safety-focused evaluation framework for assessing LLMs as components of clinical decision support systems, with particular emphasis on anticipated harm severity, concordance with predefined guideline-informed reference frameworks, and reliability in safety-critical medical tasks. Methods: A framework-based evaluation pipeline was designed incorporating standardized clinical scenarios, blinded dual-expert assessment, harm-severity classification, factual hallucination assessment, and guideline concordance scoring. One hundred high-risk anesthesiology scenarios were used as a stress-test domain to evaluate five consumer-facing LLM systems—ChatGPT Plus, Gemini Pro, Claude, Perplexity, and Copilot—yielding a total of 500 model responses. Model outputs were generated using standardized prompts and evaluated independently by two experienced anesthesiologists, with inter-rater reliability assessed using Cohen's kappa statistics. Results: Across 500 evaluated responses, the systems demonstrated substantial variability in observed safety characteristics, guideline concordance, and anticipated harm severity. While some models produced responses aligned with predefined guideline-informed reference frameworks, others generated incomplete, misleading, or potentially harmful recommendations, particularly in scenarios involving acute instability and complex decision-making. Inter-rater agreement varied across systems, indicating that some higher-risk or borderline classifications required cautious interpretation. Conclusions: This study presents a structured, safety-first evaluation framework for characterizing sampled output behavior and safety-related response patterns of consumer-facing LLM systems prior to clinical integration. Rather than ranking systems based on performance alone, the framework emphasizes risk characterization, governance, and patient safety. The proposed approach may be conceptually adaptable to other clinical domains, although such transferability requires empirical validation. From a health policy perspective, these findings highlight the need for transparent safety evaluation, governance oversight, and ongoing re-assessment before institutional adoption of LLM-based decision-support tools. Further validation in larger, repeated, and externally reviewed evaluations is required before definitive safety thresholds or governance benchmarks can be established.