Large language models in methodological quality evaluation of radiomics research based on METRICS: ChatGPT vs NotebookLM vs radiologist
European Journal of Radiology, cilt.184, 2025 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 184
- Basım Tarihi: 2025
- Doi Numarası: 10.1016/j.ejrad.2025.111960
- Dergi Adı: European Journal of Radiology
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, EMBASE, MEDLINE
- Anahtar Kelimeler: Artificial intelligence, Large language models, Machine learning, Radiomics, Texture analysis
- Sağlık Bilimleri Üniversitesi Adresli: Hayır
Özet
Objectives: This study aimed to evaluate the effectiveness of large language models (LLM) in assessing the methodological quality of radiomics research, using METhodological RadiomICs Score (METRICS) tool. Methods: This study included open access radiomic research articles published in 2024 across various journals and a preprint repository, all under the Creative Commons Attribution License. Each study was independently evaluated using METRICS by two LLMs, ChatGPT-4 and NotebookLM, and a consensus assessment performed by two radiologists with expertise in radiomics research. Results: A total of 48 open access articles were included in this study. ChatGPT-4, NotebookLM, and human readers achieved median scores of 79.5 %, 61.6 %, and 69.0 %, respectively, with a statistically significant difference across these evaluations (p < 0.05). Pairwise comparisons indicated no statistically significant difference for NotebookLM vs human experts (p > 0.05), in contrast to other pairs (p < 0.05). Intraclass correlation coefficient (ICC) for ChatGPT-4 and human experts was 0.563 (95 % CI: 0.050–––0.795), corresponding to poor to good agreement. The ICC for ChatGPT-4 and NotebookLM and for human experts and NotebookLM were 0.391 (95 % CI: −0.031–––0.665) and 0.555 (95 % CI: 0.326–––0.723), respectively, indicating poor to moderate agreement. LLMs completed the tasks in a significantly shorter time (p < 0.05). In item-wise reliability analysis, ChatGPT-4 generally demonstrated higher consistency than NotebookLM. Conclusion: LLMs hold promise for automatically evaluating the quality of radiomics research using METRICS, a new tool that is relatively more complex yet comprehensive compared to its counterparts. However, substantial improvements are needed for full alignment with human experts.