|
February 2026 TESTING LLM FUNCTIONALITY FOR SPINE CONSULTSAt the recent CSRS Annual Meeting, Srikanth Divi, MD, presented his work evaluating whether large language models (LLMs) can function as a digital spine consultant, accurately synthesizing complex clinical data and supporting surgical triage for cervical spine consults.
Using a decade of real-world inpatient consult data with assessments and plans masked, the authors benchmarked multiple LLMs against clinically relevant decision-making metrics. Study Design The study evaluated three LLMs:
The study gave each LLM clinical tasks:
The study graded LLM performance on:
Results ChatGPT‑5 consistently outperformed DeepSeek and Llama-4-Maverick across all primary endpoints. It achieved the highest mean accuracy scores with statistically significant separation from the other models. ChatGPT‑5 provided the correct operative recommendation in roughly 78% of cases, compared with DeepSeek (around 53%) and Llama-4-Maverick (around 43%). In addition, ChatGPT‑5 demonstrated more nuanced and surgeon‑level rationale, particularly for nonoperative management decisions. What Worked Well
Key Limitations
Takeaway While not a replacement for surgeon expertise, LLMs may help improve consistency and support decision‑making when thoughtfully implemented. This study showed that they have the ability to support triage and consult efficiency, not replace surgeon judgment. What’s Next The study team plans to:
|
Srikanth Divi, MD, is an assistant professor of Orthopaedic Surgery and Neurological Surgery at Northwestern Medicine
Refer a PatientNorthwestern Medicine welcomes the opportunity to collaborate with you in caring for your patients.
|
