Benchmarking Artificial Intelligence Models for Citation Accuracy in Neuro-Ophthalmological Disorders Research: A Comparative Analysis of Four Models
Journal of Hospital Librarianship, cilt.25, sa.3-4, ss.140-148, 2025 (Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 25 Sayı: 3-4
- Basım Tarihi: 2025
- Doi Numarası: 10.1080/15323269.2025.2576895
- Dergi Adı: Journal of Hospital Librarianship
- Derginin Tarandığı İndeksler: Scopus, IBZ Online, CINAHL, Library, Information Science & Technology Abstracts (LISTA), Public Affairs Index
- Sayfa Sayıları: ss.140-148
- Anahtar Kelimeler: AI models, artificial intelligence, citation accuracy, neuro-ophthalmology, PubMed citations
- Sağlık Bilimleri Üniversitesi Adresli: Evet
Özet
This study evaluated the accuracy of four artificial intelligence models—ChatGPT, Copilot, DeepSeek, and Gemini—in generating PubMed citations for neuro-ophthalmology research. Using thirty-five standardized clinical paragraphs from The Review of Ophthalmology (4th edition), each model produced references formatted in AMA 11 style. Accuracy was determined by checking publication correctness, DOI matching, and citation relevance, with expert reviewers classifying outputs as Fully Cited, Partially Cited, or Not Cited. Inter-rater reliability was measured using Cohen’s kappa. Among the models, DeepSeek demonstrated the highest accuracy (75.0%), followed by Copilot (60.5%), ChatGPT (31.4%), and Gemini (3.0%). Common issues included DOI mismatches and irrelevant references, with Gemini generating 32 incorrect citations. Expert evaluation confirmed DeepSeek’s superiority, producing 15 fully cited references compared to Copilot’s 7 and 4 each for ChatGPT and Gemini. Reviewer agreement was substantial (κ = 0.70). The findings suggest that while domain-specific AI models can aid citation generation, frequent inaccuracies and hallucinations highlight the necessity of human oversight. A hybrid approach combining AI with expert review may provide more reliable outcomes.