
Loading, please wait...

Loading, please wait...

The clinical management of unruptured intracranial aneurysms represents one of the most complex balancing acts in modern neurovascular medicine. Today, patients increasingly consult artificial intelligence platforms before meeting their neurosurgeons or interventional neuroradiologists. Consequently, individuals frequently enter specialist clinics with preformed expectations about whether they require endovascular coiling, surgical clipping, or watchful waiting. However, whether frontier large language models deliver consistent, medically sound guidance remains an essential question for clinicians worldwide.
Recent advances in generative artificial intelligence have introduced powerful frontier models into public healthcare communication. Specifically, patients now turn to systems like ChatGPT, Gemini, and Claude to decipher radiological reports and assess aneurysm rupture risks. While these digital tools demonstrate impressive fluency, they operate without clinical intuition. Therefore, evaluating how these models evaluate real-world clinical vignettes is crucial. When digital recommendations diverge from evidence-based consensus, they generate significant patient anxiety and complicate shared decision-making. In addition, neurovascular teams must proactively understand these algorithmic variations to address misleading digital advice effectively during clinical consultations. Furthermore, understanding algorithmic nuances empowers clinicians to reassure apprehensive individuals, ensuring treatment discussions remain anchored in sound clinical evidence.
To quantify inter-model variability, a landmark investigation evaluated sixty-seven consecutive aneurysm cases referred to a specialised neurovascular tertiary service. The researchers retrospectively analysed patient records spanning an entire calendar year. De-identified clinical vignettes captured detailed patient demographics, vascular comorbidities, aneurysm dimensions, morphology, and anatomical locations. Subsequently, the investigators submitted each standardized clinical scenario to three state-of-the-art frontier models: Claude Opus 4.6, ChatGPT-5.4, and Gemini 3 Pro Thinking. Each model processed every scenario across five independent iterations to examine internal consistency.
Importantly, the researchers established rigorous clinical benchmarks to contextualize algorithmic recommendations. They anchored the outputs of each model against the final consensus of a neurovascular multidisciplinary team (MDT). Furthermore, they compared results against the Unruptured Intracranial Aneurysm Treatment Score (UIATS). The UIATS represents an internationally validated scoring system that weighs patient factors, aneurysm morphology, and procedural risks. Within-model reproducibility was calculated using Fleiss' kappa, while inter-model agreement and directional propensities were evaluated using Cohen's kappa and McNemar's test. Consequently, this rigorous design enabled precise statistical comparisons between artificial intelligence outputs and validated human expertise across varying clinical scenarios.
The experimental findings revealed remarkable internal consistency alongside surprising inter-model divergence. Within-model reproducibility across the five independent runs proved almost perfect for all three platforms, yielding Fleiss' kappa values between 0.837 and 0.860. Therefore, clinicians can expect individual models to provide stable recommendations when presented with identical clinical prompts. However, pairwise comparisons revealed striking asymmetry when models evaluated the same clinical case.
Specifically, ChatGPT-5.4 and Gemini 3 Pro Thinking exhibited near-identical therapeutic behavior, achieving a high Cohen's kappa of 0.850 with a narrow 95% confidence interval. In sharp contrast, Claude Opus 4.6 diverged substantially from both platforms, recording significantly lower agreement metrics of 0.688 with ChatGPT and 0.667 with Gemini. Overall, recommendations were non-unanimous in 19.4% of the evaluated cases. Remarkably, Claude was the sole outlier in eight of those thirteen split decisions. In seven of those eight instances, Claude advocated for conservative observation while the other platforms recommended active intervention. Thus, inter-model agreement is not uniform across frontier artificial intelligence platforms. Additionally, these observable discrepancies underscore that proprietary architectures and alignment methodologies shape clinical recommendations in fundamentally divergent ways.
The study demonstrated significant directional biases when comparing model recommendations against human clinical benchmarks. When evaluated against real-world neurovascular multidisciplinary teams, Gemini and ChatGPT exhibited a significant pro-treatment propensity. McNemar's testing confirmed statistically significant leanings toward procedural intervention for both Gemini (p = 0.0022) and ChatGPT (p = 0.0153). Consequently, these models were far more likely to recommend surgical clipping or endovascular therapy than human neurovascular specialists reviewing identical clinical data.
Conversely, Claude demonstrated an opposite, highly conservative therapeutic disposition. In direct pairwise comparisons, Gemini proved significantly more pro-treatment than Claude (p = 0.0117). Furthermore, when researchers anchored Claude against the standardized UIATS criteria, the model was significantly more conservative than the established score (p = 0.0162). While aggressive models risk promoting unnecessary invasive interventions, overly cautious algorithms might delay vital treatment for unstable aneurysms. These conflicting algorithmic biases highlight the hazard of treating frontier artificial intelligence systems as homogeneous clinical advisors. Therefore, clinicians must interpret digital advice with extreme caution. In addition, specialists should remain mindful that differing platform algorithms produce widely disparate clinical impressions for the exact same patient.
These divergent algorithmic tendencies carry immediate, practical implications for neurosurgical and neurological practice worldwide. Because patients frequently test different AI tools at home, their pre-consultation expectations will vary depending on which platform they query. A patient consulting ChatGPT or Gemini may arrive convinced that urgent surgery is mandatory. Conversely, a patient querying Claude might assume that serial radiological surveillance is completely safe. Therefore, neurovascular specialists must proactively inquire about pre-visit digital research during consultations.
Additionally, clinicians must clearly articulate why multidisciplinary team deliberation remains superior to isolated digital advice. Real-world decision-making accounts for intricate nuances, including microcatheter navigability, patient frailty, and institutional surgical outcomes. Automated models cannot reliably weigh these dynamic intraoperative and anatomical factors. Furthermore, future investigations should test diverse prompt architectures, patient affective framing, and multi-center validation across international healthcare settings. Calibrating algorithms against real-world cohorts ensures safer clinical decision-support development. Ultimately, multidisciplinary consensus remains the gold standard of care. Consequently, clinicians should view artificial intelligence as an exploratory adjuvant rather than a surrogate for nuanced, individualized specialist evaluation. Such balanced approaches protect clinical integrity and improve overall patient outcomes.
Frontier models demonstrate strong internal consistency across multiple runs. However, their treatment recommendations diverge significantly from one another. Gemini and ChatGPT frequently recommend active surgical or endovascular interventions. In contrast, Claude adopts a more conservative, observational stance. Consequently, clinicians must not rely on these automated tools without independent multidisciplinary evaluation.
Claude Opus 4.6 demonstrated a markedly conservative therapeutic philosophy compared to ChatGPT-5.4 and Gemini 3 Pro Thinking. Specifically, Claude recommended observational monitoring in cases where the other models favored intervention. In addition, Claude proved more conservative than standardized UIATS guidelines. These variations likely stem from differences in pre-training data, safety thresholds, and alignment protocols.
Clinicians must actively ask patients whether they consulted artificial intelligence tools prior to their appointment. Because models often display pro-treatment tendencies, patients may anticipate invasive surgery prematurely. Therefore, specialists should explain the nuanced risks of intervention versus rupture. Transparent communication contextualizes algorithmic recommendations against established multidisciplinary standards, addressing patient anxiety and correcting misconceptions.
Disclaimer: This content is for informational and educational purposes only... Refer to the latest local and national guidelines for clinical practice.
References
Narayan A et al. Inter-Model Variability of ChatGPT, Gemini and Claude in Treatment Recommendations for Unruptured Intracranial Aneurysms. Neurosurg Rev. 2026 Sep 16. doi: 10.1007/s10143-026-04498-1. PMID: 42747656.
Etminan N, Brown RD Jr, Beseoglu K, et al. The unruptured intracranial aneurysm treatment score: a multidisciplinary consensus. Neurology. 2015;85(10):881-889.
European Stroke Organisation. European Stroke Organisation (ESO) guidelines on management of unruptured intracranial aneurysms. Eur Stroke J. 2022;7(3):LXI-LXXXVIII.

Read summarized clinical updates, watch expert medical content, and earn CME certifications right from your smartphone.


A benchmarking study reveals significant inter-model variability among ChatGPT, Gemini, and Claude for unruptured intracranial aneurysm management. While ChatGPT and Gemini favor intervention, Claude leans conservative. Clinicians must anticipate AI-driven patient expectations.
Today

A rare case of acute coronary obstruction during TAVR caused by suspected calcified amorphous tumor embolization underscores the critical role of intracardiac imaging, emergency ECMO, and the paradox of enlarged-aperture valve geometry.
Today

Prelacteal feeding persists in South Asian communities despite known clinical risks. This review examines determinants from Chitral, pathophysiological hazards, and pediatric strategies to eliminate harmful feeds and support exclusive breastfeeding.
Today

Transcatheter mitral valve replacement in patients with degenerated bioprostheses carries a substantial risk of LVOT obstruction. Combining leaflet modification via the BATMAN technique with mechanical circulatory support offers a viable, life-saving solution for anatomically challenging surgical candidates.
Today

A comparative cohort study demonstrates that patients undergoing total hip arthroplasty after prior hip arthroscopy achieve comparable functional recovery, low complication rates, and equivalent implant survivorship when conversion occurs after six months.
Today

A cross-sectional study of 320 adults evaluated the link between habitual fasting duration and dietary quality using Nutrition Quotient scores, highlighting the clinical necessity of balancing time-restricted eating with nutrient density and meal quality in routine metabolic practice.
Today