Large language models in artificial intelligence to answer patient questions in spine surgery: an evaluation of current evidence

0Citations
Citations of this article
4Readers
Mendeley users who have this article in their library.
Get full text

Abstract

BACKGROUND: Large language models (LLMs) are increasingly being explored in healthcare, particularly for enhancing patient education. In spine surgery, LLMs have the potential to enhance communication and support patients through perioperative care. However, concerns remain regarding the accuracy, readability, and overall reliability of these tools in delivering patient-facing information. This review aimed to understand the current use of LLMs in answering patient questions in spine surgery. METHODS: A structured search of PubMed and Google Scholar was conducted using terms focused on LLMs and neurosurgery. Studies were only included if they tested LLMs’ ability in answering patient questions related to spine surgery. Exclusion criteria included non-peer-reviewed articles, studies that did not evaluate chatbot performance, or those using LLMs for non-educational purposes. RESULTS: LLMs were tested across a variety of spine-related topics, including scoliosis, lumbar and cervical fusion, endoscopic procedures, and spinal cord stimulation. Studies consistently reported moderate to high accuracy ratings. Readability scores remained a limitation, with most responses written at a college reading level. Empathy and clarity varied by model and condition, with some studies showing improved ratings when assessed by non-medical reviewers. Methodological variability across studies introduced inconsistencies and limited comparability. CONCLUSIONS: LLMs show promising utility for patient education in spine surgery for addressing frequently asked questions. However, challenges in readability, accuracy, and standardization limit their current clinical adoption. Moving forward, studies must incorporate standardized evaluation tools, address high rate of content hallucination, and focus on chatbot performance in personalized scenarios. Cross-disciplinary collaboration is essential to ensure safe, accessible integration into neurosurgical care pathways.

Cite

CITATION STYLE

APA

Patel, J., Tirmizi, Z., Waheed, A. A., Aydin, S., Gajjar, A. A., Muhammad, N., … Deng, H. (2026, June 1). Large language models in artificial intelligence to answer patient questions in spine surgery: an evaluation of current evidence. Journal of Neurosurgical Sciences. Edizioni Minerva Medica. https://doi.org/10.23736/S0390-5616.26.06678-6

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free