Evaluating the Educational Quality of Benign Prostate Surgery Videos on YouTube: A Human and AI-Based Comparative Analysis YouTube’daki Benign Prostat Cerrahisi Videolarının Eğitsel Kalitesinin Değerlendirilmesi: İnsan ve Yapay Zekâ Tabanlı Karşılaştırmalı Bir Analiz
Medeniyet Medical Journal, cilt.41, sa.1, ss.51-58, 2026 (ESCI, Scopus, TRDizin)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 41 Sayı: 1
- Basım Tarihi: 2026
- Doi Numarası: 10.4274/mmj.galenos.2026.38586
- Dergi Adı: Medeniyet Medical Journal
- Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus, TR DİZİN (ULAKBİM)
- Sayfa Sayıları: ss.51-58
- Anahtar Kelimeler: Benign prostate obstruction, ChatGPT, DISCERN, Global Quality Score (GQS), urology, YouTube
- Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
- Sağlık Bilimleri Üniversitesi Adresli: Evet
Özet
Objective: YouTube has become an increasingly popular platform for surgical education, yet the quality and reliability of its medical content remain uncertain. With the rise of artificial intelligence (AI), models such as ChatGPT offer new possibilities for automated educational content evaluation. This study aimed to compare human and AI-based assessments of the educational quality and reliability of YouTube videos on benign prostate surgery. Methods: A total of 100 videos of holmium laser enucleation of the prostate, transurethral resection of the prostate (TURP), transvesical prostatectomy, and thulium fiber laser enucleation of the prostate were analyzed. Two urology specialists and ChatGPT-5 (in two independent runs) evaluated each video using the Global Quality Score (GQS) and the modified DISCERN tool. Popularity metrics (views, likes, subscribers, duration) were also recorded. Non-parametric statistical tests and Spearman correlation analyses were applied. Results: Human raters assigned significantly higher DISCERN and GQS scores than both AI runs (p<0.01). TURP videos consistently received lower scores across all evaluators. No significant quality differences were found among video sources. Both AI runs showed strong internal consistency (ρ=0.62-0.75) and reproduced human rating patterns, though with lower mean values. Engagement metrics showed weak or no correlation with quality. Conclusions: AI models can provide consistent, scalable quality assessments but still underestimate educational value compared with human experts. Hybrid AI-expert evaluation may enhance the reliability of the appraisal of online surgical videos.