Exploring Artificial Intelligence’s Role in Citation Generation for Ocular Inflammation and Uveal Diseases Research: A Comparative Evaluation Across Four Models


Civelekler M., ÇITIRIK M.

Ocular Immunology and Inflammation, cilt.34, sa.2, ss.397-402, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 34 Sayı: 2
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1080/09273948.2026.2615858
  • Dergi Adı: Ocular Immunology and Inflammation
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, EMBASE, MEDLINE, Academic Search Ultimate (EBSCO), Biomedical Reference Collection: Corporate Edition (EBSCO)
  • Sayfa Sayıları: ss.397-402
  • Anahtar Kelimeler: Artificial intelligence, citation accuracy, ocular inflammation, uveal diseases, AI Hallucinations
  • Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Purpose: This study evaluated four artificial intelligence (AI) models—ChatGPT, Copilot, DeepSeek, and Gemini—for their ability to generate PubMed citations related to ocular inflammation and uveal disease. The aim was to assess their performance in a specialized clinical context and determine whether these tools can support accurate academic referencing. Methods: Thirty-five clinical paragraphs from The Review of Ophthalmology (4th edition) were provided to each model, which was instructed to generate AMA 11-style PubMed citations. Outputs were examined for accuracy, DOI matching, and clinical relevance. Expert reviewers classified each citation as Fully Cited, Partially Cited, or Not Cited. Statistical differences among the models were assessed using ANOVA with post hoc analysis. Results: DeepSeek, a domain-specific model, outperformed other models with an accuracy of 65.7% (p < 0.001). Copilot and ChatGPT achieved moderate accuracy rates of 42.9% and 37.1% (p = 0.042), respectively, whereas Gemini performed the worst, with an accuracy of 5.7% (p < 0.001). These results were statistically significant, as confirmed by ANOVA and post hoc analysis. Additional errors included incorrect citations, DOI mismatches, and incomplete reference lists. Expert validation also showed that DeepSeek produced the highest number of fully accurate citations, while the remaining models generated more partial or uncited references. Conclusion: AI tools can assist with citation generation, but their reliability varies significantly. Domain-specific systems perform better, yet inconsistencies such as partial citations and hallucinated details highlight the continued need for expert oversight. Accurate academic referencing still depends on combining AI-generated material with careful human review.