Using ChatGPT-4 in visual field test assessment


Gumus Akgun G., Altan C., Balci A. S., ALAGÖZ N., Çakır I., YAŞAR T.

Clinical and Experimental Optometry, cilt.108, sa.8, ss.1031-1036, 2025 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 108 Sayı: 8
  • Basım Tarihi: 2025
  • Doi Numarası: 10.1080/08164622.2025.2463518
  • Dergi Adı: Clinical and Experimental Optometry
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, MEDLINE
  • Sayfa Sayıları: ss.1031-1036
  • Anahtar Kelimeler: Artificial intelligence, ChatGPT, large language model, visual field test
  • Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Clinical relevance: Visual field testing is essential in the diagnosis and management of various ophthalmic diseases, particularly glaucoma. Integrating ChatGPT-4 into the interpretation of these tests has the potential to aid clinical decision making and improve efficiency and accessibility in clinical practice. Background: This study aims to evaluate the capability of ChatGPT-4 in interpreting visual field tests. Method: A total of 30 patient visual field printouts, either with or without defects, were included in this study. The performance of ChatGPT-4 in identifying test name, pattern, reliability indices, total deviation map, pattern deviation map and greyscale map was evaluated and compared with that of 2 experienced glaucoma consultants. The study also focused on the ability of ChatGPT to categorise tests as ‘normal’ or suggest diagnosis by interpreting tests accurately. Results: The results showed that ChatGPT-4 was highly accurate in identifying test names (100%), patterns (90%) and global visual field indices (96.7%). It also accurately classified tests as reliable or unreliable (93.3%).The model provided 66.7% and 30% accurate and adequate answers in interpreting deviation and greyscale maps, respectively. In addition, in 33.3% of tests, it was able to accurately interpret the visual field test and classify it as ‘normal’ or suggest a diagnosis. Conclusion: The study highlights the potential of large language models like ChatGPT-4 in assessing visual field tests. ChatGPT-4 could interpret numeric data on tests accurately. However, it was inadequate in interpreting deviation and greyscale maps and suggesting a diagnosis according to the defects.