
Loading, please wait...

Loading, please wait...

Large language models (LLMs) are increasingly recognized as potential decision-support tools across various medical fields. Recent research has specifically explored the role of LLMs in anesthesia, focusing on their ability to assign American Society of Anaesthesiologists (ASA) physical status classifications and generate appropriate anesthetic protocols. A retrospective study conducted at Atatürk University Veterinary Hospital analyzed 225 feline and canine cases to benchmark the performance of ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Pro.
Correctly identifying a patient's ASA status is critical for risk stratification and perioperative planning. In this study, ChatGPT-5 demonstrated the highest accuracy in ASA classification at 53.3%. This was followed by ChatGPT-4o at 46.7% and Gemini 2.5 Pro at 30.7%. Interestingly, the models performed significantly better in higher-risk categories (ASA 3‒5). Conversely, healthy patients in the ASA 1 category were frequently misclassified. This discrepancy often stemmed from an overestimation of risk, where the AI models flagged minor clinical findings as more severe than an experienced anesthesiologist would.
Beyond risk scoring, the researchers evaluated the clinical adequacy of the generated anesthetic protocols using a four-point scale. ChatGPT-5 emerged as the superior model, consistently producing the most clinically sufficient protocols compared to its counterparts. However, while the models could synthesize complex data into a structured format, the study emphasized that model-to-model variability remains a concern. Furthermore, clinicians observed that the models occasionally suggested drug combinations that required careful expert adjustment.
While the study utilized veterinary cases, the findings offer valuable insights for human medicine. The ability of LLMs to process unstructured clinical data and provide rapid decision support is promising. Nevertheless, the high rate of misclassification in healthy patients underscores a current limitation: the tendency for AI to prioritize caution over clinical nuance. Experts suggest that while these tools can streamline documentation and protocol drafting, they cannot yet function autonomously.
LLMs can analyze patient history and clinical data to suggest an ASA physical status score, helping clinicians categorize preoperative risks more consistently, though they currently require expert validation.
In the Atatürk University study, ChatGPT-5 achieved the highest accuracy (53.3%) and generated the most clinically relevant anesthetic protocols compared to ChatGPT-4o and Gemini 2.5 Pro.
No. The study concludes that performance varies significantly across different models and cases, making expert oversight essential for ensuring patient safety and protocol accuracy.
Disclaimer: This content is for informational and educational purposes only. It is not intended to provide medical advice or to be a substitute for professional medical judgment, diagnosis, or treatment. Refer to the latest local and national guidelines for clinical practice.
References

Read summarized clinical updates, watch expert medical content, and earn CME certifications right from your smartphone.


A study assesses ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Pro in determining ASA status and anesthesia protocols for feline and canine cases....
2 months ago

Andhra Pradesh reported 10 new Covid-19 cases, taking the state tally to 49 while deaths remain at four. With 24 patients hospitalized and 16 under home isolation, the Health Department has intensified monitoring. Medical professionals should review regional distribution, diagnostic protocols, and management plans.
Today

An 11-year Swedish registry study of 618 uterine sarcoma patients found that minimally invasive surgery yielded survival comparable to open surgery in early stages. However, adjuvant chemotherapy conferred no survival benefit in localized or advanced disease, highlighting stage and histology as key outcomes.
3 days back

A cross-sectional study evaluates post-intensive care syndrome in cardiac patients 2-4 weeks post-ICU discharge, highlighting cognitive, psychological, and functional impairments and the need for structured multidisciplinary rehabilitation.
3 days back

Anterior cruciate ligament reconstruction failure lacks uniform definition. A narrative review proposes an integrative framework incorporating objective and subjective instability, persistent pain, restricted motion, graft rupture, and secondary meniscal injury to standardize clinical reporting.
3 days back

With World Obesity Atlas data warning that over 41 million Indian children are overweight or obese, ICMR and NIN have unveiled a 10-point policy roadmap. The initiative calls for mandatory front-of-pack labeling, HFSS taxes, strict marketing bans, and healthier school environments to curb non-communicable diseases.
Today