Diagnostic Performance of a General-Purpose Artificial Intelligence Model in Radiographic Staging of Coxarthrosis According to the Tönnis Classification: A Multiclass Analysis


Özbek İ. C., AKGÜL Ö., Temel M. H.

Indian Journal of Orthopaedics, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1007/s43465-026-01919-7
  • Dergi Adı: Indian Journal of Orthopaedics
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, Academic Search Ultimate (EBSCO), Health Research Premium Collection (ProQuest)
  • Anahtar Kelimeler: Hip osteoarthritis, T & ouml;nnis classification, Artificial intelligence, Radiographic staging, Diagnostic performance
  • Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Objective: Recent advances in artificial intelligence have enabled the use of general-purpose multimodal models in medical image analysis. This study evaluated the diagnostic performance of ChatGPT-5o, a vision–language model, for staging hip osteoarthritis according to the Tönnis classification using anteroposterior pelvic radiographs. Methods: In this cross-sectional diagnostic study,400 anteroposterior pelvic radiographs obtained for hip pain or suspected osteoarthritis were analyzed using stratified sampling to ensure equal representation of each Tönnis stage. Reference staging was determined independently by two musculoskeletal imaging specialists and finalized by consensus. ChatGPT-5o assessed each radiograph in isolated sessions using a standardized prompt without access to clinical data or additional imaging. Diagnostic performance was evaluated using accuracy, sensitivity, specificity, precision, macro-F1 score, Cohen’s κ coefficient, and one-vs-rest receiver operating characteristic (ROC) analysis. Results: Overall multiclass accuracy was 39.3% (157/400) with a macro-F1 score of 0.392. The unweighted Cohen’s κ was 0.190, indicating slight agreement, whereas ordinal weighted κ values were higher (linear κ = 0.329; quadratic κ = 0.453). Sensitivity varied across stages and was lowest in Stage 1 (21.0%). Sensitivity and specificity were 60.0 and 87.7% for Stage 0, respectively. In Stage 3, specificity reached 80.3% with a sensitivity of 42.0%. Most errors occurred between adjacent stages, demonstrating a tendency toward stage compression. ROC analysis showed AUC values of 0.775 for Stage 0 and 0.707 for Stage 3, while Stage 1 discrimination was near chance (AUC = 0.522). The micro-average AUC was 0.656. Conclusion: ChatGPT-5o demonstrated low-to-moderate diagnostic performance in Tönnis staging of hip osteoarthritis. General-purpose vision–language models currently appear insufficient as standalone diagnostic tools for coxarthrosis staging but may have supportive roles when combined with task-specific AI systems and expert clinical evaluation.