Interobserver Variability in Hysterosalpingography Interpretation: Radiologists vs Gynecologists
International Journal of Women's Health, cilt.17, ss.3429-3435, 2025 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 17
- Basım Tarihi: 2025
- Doi Numarası: 10.2147/ijwh.s555065
- Dergi Adı: International Journal of Women's Health
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, EMBASE, Directory of Open Access Journals
- Sayfa Sayıları: ss.3429-3435
- Anahtar Kelimeler: hysterosalpingography, interobserver variability, radiologists, gynecologists, infertility, diagnostic imaging, Fleiss' kappa, standardization, artificial intelligence
- Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
- Sağlık Bilimleri Üniversitesi Adresli: Evet
Özet
Objective: To evaluate interobserver variability in the interpretation of hysterosalpingography (HSG) examinations among radiologists, gynecologists, and between the two specialties, highlighting areas of diagnostic agreement and discrepancy. Materials and Methods: In this prospective, multicenter study, 12 specialists (6 radiologists and 6 gynecologists) independently reviewed HSG images from 100 patients, evaluating 10 predefined diagnostic categories and overall image quality. Fleiss’ kappa (κ) was used to assess interobserver agreement within and between groups. Results: Interobserver agreement ranged from poor to moderate across most diagnostic categories. The highest agreement was observed for tubal occlusion (κ = 0.508 radiologists vs gynecologists; κ = 0.536 among radiologists; κ = 0.460 among gynecologists), followed by uterine anomalies and hydrosalpinx. Poor agreement was noted for subjective parameters such as image quality and additional findings (κ values < 0.1), with some negative kappa scores indicating agreement below chance. Conclusion: Significant interobserver variability exists in HSG interpretation, particularly between radiologists and gynecologists. Structured findings yielded higher agreement, while subjective assessments showed poor reproducibility. These findings underscore the need for standardized reporting guidelines, interdisciplinary collaboration, and the potential integration of AI-assisted interpretation to enhance diagnostic consistency and improve patient outcomes.