Readability and Appropriateness of Responses Generated by ChatGPT 3.5, ChatGPT 4.0, Gemini, and Microsoft Copilot for FAQs in Refractive Surgery


Creative Commons License

Aydın F. O., Aksoy B. K., Ceylan A., Akbaş Y. B., Ermiş S., KEPEZ YILDIZ B., ...Daha Fazla

Turkish Journal of Ophthalmology, cilt.54, sa.6, ss.313-317, 2024 (Scopus, TRDizin)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 54 Sayı: 6
  • Basım Tarihi: 2024
  • Doi Numarası: 10.4274/tjo.galenos.2024.28234
  • Dergi Adı: Turkish Journal of Ophthalmology
  • Derginin Tarandığı İndeksler: Scopus, TR DİZİN (ULAKBİM)
  • Sayfa Sayıları: ss.313-317
  • Anahtar Kelimeler: Artificial intelligence, chatbots, refractive surgery FAQs, ChatGPT, Gemini, Copilot
  • Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
  • Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Objectives: To assess the appropriateness and readability of large language model (LLM) chatbots’ answers to frequently asked questions about refractive surgery. Materials and Methods: Four commonly used LLM chatbots were asked 40 questions frequently asked by patients about refractive surgery. The appropriateness of the answers was evaluated by 2 experienced refractive surgeons. Readability was evaluated with 5 different indexes. Results: Based on the responses generated by the LLM chatbots, 45% (n=18) of the answers given by ChatGPT 3.5 were correct, while this rate was 52.5% (n=21) for ChatGPT 4.0, 87.5% (n=35) for Gemini, and 60% (n=24) for Copilot. In terms of readability, it was observed that all LLM chatbots were very difficult to read and required a university degree. Conclusion: These LLM chatbots, which are finding a place in our daily lives, can occasionally provide inappropriate answers. Although all were difficult to read, Gemini was the most successful LLM chatbot in terms of generating appropriate answers and was relatively better in terms of readability.