Comparative quality and readability of explanatory responses generated by ChatGPT-4, Claude, Gemini, and Copilot for undergraduate medical biochemistry questions
Turkish Journal of Biochemistry, 2026 (SCI-Expanded, Scopus, TRDizin)
- Yayın Türü: Makale / Tam Makale
- Basım Tarihi: 2026
- Doi Numarası: 10.1515/tjb-2026-0027
- Dergi Adı: Turkish Journal of Biochemistry
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Applied Science & Technology Source, EMBASE, Food Science & Technology Abstracts, Directory of Open Access Journals, TR DİZİN (ULAKBİM)
- Anahtar Kelimeler: artificial intelligence, medical education, undergraduate, biochemistry, natural language processing, readability
- Sağlık Bilimleri Üniversitesi Adresli: Evet
Özet
This study aimed to compare the educational quality and readability of explanatory answers generated by ChatGPT-4, Claude, Gemini, and Copilot for open-ended undergraduate medical biochemistry questions. Twenty open-ended questions were prepared by two Medical Biochemistry specialists. Identical prompts were submitted once to each model, yielding 80 responses. Two blinded experts evaluated the responses using an 11-item binary analytic rubric and the Global Quality Scale. Readability was assessed using the Flesch Reading Ease Score. Group differences were analyzed using Friedman tests with Holm-adjusted post-hoc comparisons, and inter-rater agreement was assessed using ICC(3,1). Analytic rubric scores differed significantly among models (p<0.001). Claude achieved the highest median score [10.0 (9.13-11.0)], followed by Gemini [8.0 (6.38-9.88)], ChatGPT-4 [7.0 (5.63-9.88)], and Copilot [7.0 (6.50-9.00)]. Global Quality Scale ratings were also highest for Claude and Gemini (p<0.001). Readability differed significantly among models (p<0.001), with Claude showing the lowest reading ease values and all models falling within the "very difficult"range. Claude and Gemini generated higher-quality medical biochemistry explanations than ChatGPT-4 and Copilot. However, readability scores reflected surface-level linguistic complexity and should not be interpreted as direct indicators of student comprehension.