
Loading, please wait...

Loading, please wait...

Large language models (LLMs) are increasingly recognized as potential decision-support tools across various medical fields. Recent research has specifically explored the role of LLMs in anesthesia, focusing on their ability to assign American Society of Anaesthesiologists (ASA) physical status classifications and generate appropriate anesthetic protocols. A retrospective study conducted at Atatürk University Veterinary Hospital analyzed 225 feline and canine cases to benchmark the performance of ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Pro.
Correctly identifying a patient's ASA status is critical for risk stratification and perioperative planning. In this study, ChatGPT-5 demonstrated the highest accuracy in ASA classification at 53.3%. This was followed by ChatGPT-4o at 46.7% and Gemini 2.5 Pro at 30.7%. Interestingly, the models performed significantly better in higher-risk categories (ASA 3‒5). Conversely, healthy patients in the ASA 1 category were frequently misclassified. This discrepancy often stemmed from an overestimation of risk, where the AI models flagged minor clinical findings as more severe than an experienced anesthesiologist would.
Beyond risk scoring, the researchers evaluated the clinical adequacy of the generated anesthetic protocols using a four-point scale. ChatGPT-5 emerged as the superior model, consistently producing the most clinically sufficient protocols compared to its counterparts. However, while the models could synthesize complex data into a structured format, the study emphasized that model-to-model variability remains a concern. Furthermore, clinicians observed that the models occasionally suggested drug combinations that required careful expert adjustment.
While the study utilized veterinary cases, the findings offer valuable insights for human medicine. The ability of LLMs to process unstructured clinical data and provide rapid decision support is promising. Nevertheless, the high rate of misclassification in healthy patients underscores a current limitation: the tendency for AI to prioritize caution over clinical nuance. Experts suggest that while these tools can streamline documentation and protocol drafting, they cannot yet function autonomously.
LLMs can analyze patient history and clinical data to suggest an ASA physical status score, helping clinicians categorize preoperative risks more consistently, though they currently require expert validation.
In the Atatürk University study, ChatGPT-5 achieved the highest accuracy (53.3%) and generated the most clinically relevant anesthetic protocols compared to ChatGPT-4o and Gemini 2.5 Pro.
No. The study concludes that performance varies significantly across different models and cases, making expert oversight essential for ensuring patient safety and protocol accuracy.
Disclaimer: This content is for informational and educational purposes only. It is not intended to provide medical advice or to be a substitute for professional medical judgment, diagnosis, or treatment. Refer to the latest local and national guidelines for clinical practice.
References

Read summarized clinical updates, watch expert medical content, and earn CME certifications right from your smartphone.


A study assesses ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Pro in determining ASA status and anesthesia protocols for feline and canine cases....
3 months ago

A new study reveals that lipid-related metabolic dysregulation, marked by elevated TG/HDL-C ratio and glymphatic changes, independently impacts survival in idiopathic normal pressure hydrocephalus.
Today

A premature neonate developed upper limb compartment syndrome after uterine rupture extruded the arm through a scar defect. Conservative management with continuous monitoring yielded complete functional recovery and normal limb growth at 10-year follow-up, highlighting non-operative safety in selected cases.
Today

A meta-analysis of 13 propensity score-matched studies shows ViV-TAVR delivers lower early mortality and reduced bleeding compared to redo-SAVR for degenerated bioprosthetic aortic valves, though long-term hemodynamics warrant careful anatomical and patient-centered evaluation.
Today

Endoscopic posterior cervical fusion combines minimally invasive decompression, joint preparation, and rigid screw-rod fixation for atlantoaxial pathologies. Early clinical findings demonstrate solid bony union, excellent symptom relief, and minimal soft-tissue morbidity without significant vascular compromise.
Yesterday

The All-India Food Processors' Association has approached the Supreme Court to oppose FSSAI's proposed per-100g benchmark for front-of-pack warning labels, advocating instead for a per-serving threshold. We explore the regulatory showdown, nutritional evidence, and implications for clinical lifestyle counseling.
Today