Investigation of an Open AI Model in the Analysis of STN Microelectrode Recordings: Consistency With Clinicians and Potential for DBS Targeting STN Analizinde Açık Bir Yapay Zeka Modelinin Araştırılması Mikroelektrot Kayıtları: Klinisyenlerle Tutarlılık ve DBS Hedefleme Potansiyeli


Creative Commons License

Haşimoğlu O., Altınkaya A., Hanoğlu T., Karaçoban T. Ö., Geylan N. B., Demir F., ...Daha Fazla

Medical Journal of Bakirkoy, cilt.21, sa.3, ss.288-295, 2025 (ESCI, Scopus, TRDizin)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 21 Sayı: 3
  • Basım Tarihi: 2025
  • Doi Numarası: 10.4274/bmj.galenos.2025.2025.2-1
  • Dergi Adı: Medical Journal of Bakirkoy
  • Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus, CINAHL, EMBASE, TR DİZİN (ULAKBİM)
  • Sayfa Sayıları: ss.288-295
  • Anahtar Kelimeler: Deep brain stimulation, subthalamic nucleus, microelectrode recordings, artificial intelligence, machine learning, ChatGPT
  • Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
  • Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Objective: Deep brain stimulation (DBS) of the subthalamic nucleus (STN) requires precise electrode placement, often assisted by microelectrode recording (MER). However, the interpretation of MER remains highly subjective, varying among clinicians based on experience. This study evaluates the ability of an artificial intelligence (AI) model (ChatGPT 4.0) to classify STN MER recordings and evaluate its consistency with experienced and less experienced clinicians. Methods: A total of 32 STN MER recordings were independently evaluated by two experienced clinicians, two less experienced clinicians, and the AI model. Classifications were assigned to artifact, thalamus, silent, STN, suspicious STN, substantia nigra, and N/A (no recording), categories. Fleiss’ Kappa was used to assess inter-rater consistency, while Cohen’s Kappa measured agreement between generative pre-trained transformer (GPT) and each clinician. Additionally, precision and recall were calculated for each category. Results: The overall Fleiss’ Kappa among all evaluators was 0.544, with higher agreement among experienced clinicians (0.738) compared to less experienced ones (0.631). GPT showed low agreement with both groups, with Cohen’s Kappa values ranging from 0.341 to 0.375. GPT demonstrated the highest accuracy in detecting STN (73.47%), but its performance was significantly lower for other categories. Within-category consistency (14.28%) indicated variability in transition zones, with a misclassification rate of 45.87% compared to the majority opinion of clinicians. Conclusion: While GPT exhibited partial consistency with clinicians in identifying the STN, its reliability in classifying transition zones and adjacent structures was low. For AI to serve as a reliable tool in STN targeting, further refinement of its algorithms and expanded training datasets is necessary. Although GPT is not yet suitable for clinical decision making, its potential for future DBS applications is promising.