
Loading, please wait...

Loading, please wait...

Artificial intelligence is rapidly expanding across clinical ophthalmology, offering novel opportunities to assist with diagnostic synthesis and treatment planning. In surgical subspecialties, algorithmic tools now provide automated procedural recommendations. However, evaluating the validity and clinical safety of these tools remains essential before clinical integration. Recent research explores how foundation models perform when assisting with glaucoma surgical decision-making in real-world scenarios. Glaucoma management demands nuanced clinical reasoning, particularly when medical therapies fail and surgical intervention becomes necessary. Therefore, clinicians must understand the capabilities and limitations of generative artificial intelligence when guiding procedural choices.
Ophthalmic surgeons routinely evaluate multifaceted variables before selecting an appropriate operative intervention. Factors such as target intraocular pressure, previous ocular surgeries, conjunctival mobility, and lens status heavily influence procedural selection. When clinicians introduce artificial intelligence into this diagnostic equation, they expect precise, guideline-concordant guidance. Consequently, recent investigations have focused on how large language models navigate these intricate clinical determinations.
Leading conversational models, including OpenAI's ChatGPT, Microsoft Copilot, and Google Gemini, demonstrate remarkable fluency across complex medical domains. However, conversational fluency does not inherently guarantee surgical safety or sound therapeutic rationale. While early studies highlighted the potential of language models in drafting patient education materials, procedural planning requires far greater precision. Glaucoma surgical decision-making involves substantial risk, as improper intervention choices can lead to irreversible vision loss or severe surgical complications.
Furthermore, glaucoma represents a spectrum ranging from straightforward open-angle presentations to highly refractory secondary glaucomas. Because clinical guidelines continuously evolve with new microinvasive glaucoma surgery devices and tube shunts, models must synthesize dynamic evidence. Researchers therefore sought to evaluate whether public large language models could consistently generate rational, safe, and feasible surgical recommendations when presented with authentic patient records.
To rigorously evaluate model performance, investigators conducted a retrospective study utilizing authentic patient records from a tertiary eye care hospital. The research team converted real-world clinical encounters into standardized vignettes representing diverse patient profiles. Subsequently, researchers stratified these clinical scenarios into two distinct cohorts: primary glaucoma cases and complex glaucoma cases. This stratification allowed the team to assess whether structural complexity altered model accuracy or therapeutic reasoning.
Each language model—ChatGPT, Microsoft Copilot, and Google Gemini—received identical standardized clinical prompts. The prompts required each model to recommend a single surgical approach and articulate a comprehensive clinical rationale. Following output generation, a panel of fellowship-trained glaucoma specialists conducted a masked evaluation. To ensure rigorous benchmarking, the specialists evaluated each recommendation using six predefined parameters: appropriateness, rationale quality, procedural specificity, guideline adherence, technical feasibility, and potential safety risks.
The evaluators scored each criterion using a normalized scale from zero to one hundred points. Additionally, the study design minimized bias through masked scoring and standardized prompt engineering. By employing genuine clinical histories rather than theoretical vignettes, the investigators effectively simulated real-world clinical consultation. Thus, the methodology provided an objective framework to assess algorithmic decision support across varying levels of ophthalmic disease severity.
The overall median quality scores across all evaluated clinical cases showed remarkable parity among the three platforms. ChatGPT achieved a median score of 80.7, Microsoft Copilot recorded 80.7, and Google Gemini scored 82.7. Statistical analysis revealed no significant difference in general baseline performance among the models across the entire dataset. In standard, uncomplicated open-angle glaucoma scenarios, all three models frequently recommended appropriate first-line interventions, such as trabeculectomy or selective laser trabeculoplasty.
However, case complexity exerted a profound impact on model performance. When evaluating complex glaucoma scenarios, substantial disparities emerged among the platforms. For Google Gemini, scores for procedural appropriateness and rationale quality decreased significantly compared to its baseline results. Moreover, Gemini demonstrated a statistically significant increase in safety risk scores during complex case evaluations. This sharp decline revealed an inability to properly weigh complicating clinical factors, such as extensive conjunctival scarring or prior surgical failure.
In contrast, ChatGPT and Microsoft Copilot demonstrated superior stability across different clinical categories. Nevertheless, both ChatGPT and Copilot experienced a statistically significant reduction in rationale quality when analyzing complex presentations. While they selected reasonable procedures more consistently, their underlying explanations lacked depth. Consequently, these findings highlight that increasing disease complexity significantly strains algorithmic clinical reasoning.
The increase in safety risk scores among complex cases represents the most critical finding for practicing ophthalmologists. In advanced glaucoma management, choosing an inappropriate intervention can precipitate severe complications, including hypotony maculopathy, bleb failure, or accelerated endothelial cell loss. Although language models frequently memorize textbook indications, they often overlook subtle contraindications present in real-world patient records.
For instance, models occasionally failed to account for prior incisional surgeries or concurrent corneal disease when recommending filtering procedures. In complex cases requiring glaucoma drainage devices or cyclodestructive procedures, the models sometimes defaulted to conventional trabeculectomy without adequate risk mitigation. Furthermore, the models struggled to integrate past pharmacological failures and structural visual field progression into a cohesive surgical strategy.
Additionally, artificial intelligence systems remain vulnerable to hallucination and unjustified confidence. Even when models generated suboptimal surgical plans, their generated rationales appeared grammatically sound and authoritative. This authoritative tone creates a false sense of security for inexperienced clinicians or trainees who might rely on automated suggestions. Therefore, current foundation models lack the nuanced clinical judgment necessary to manage high-risk surgical scenarios independently. Without stringent human oversight, uncritical reliance on AI recommendations introduces unacceptable risks into ophthalmic surgical planning.
The findings of this benchmarking study carry practical implications for clinicians navigating the integration of artificial intelligence into ophthalmic workflows. First, while large language models can serve as useful educational aids and administrative assistants, they cannot replace specialized surgical judgment. Clinicians must recognize that overall median performance scores do not reflect clinical dependability in challenging individual cases.
Second, ophthalmic trainees and junior surgeons must exercise caution when utilizing conversational AI platforms for clinical problem-solving. Because models perform best in standard primary glaucoma scenarios, they may reinforce basic algorithmic knowledge effectively. However, in complex scenarios involving secondary glaucomas, previous surgical failures, or multiorgan pathology, expert clinical evaluation remains irreplaceable. Glaucoma specialists integrate physical examination findings, gonioscopic anatomy, and patient lifestyle factors that text-only prompts cannot capture fully.
Ultimately, artificial intelligence should function as a complementary tool rather than an autonomous decision-maker in ophthalmic surgery. Future developments must focus on domain-specific fine-tuning, multimodal image integration, and specialized safety guardrails. Until validated ophthalmic foundation models undergo prospective clinical trials, surgeons must maintain strict oversight over all AI-assisted procedural planning. By treating automated suggestions with disciplined skepticism, ophthalmologists can harness technological advances while safeguarding patient visual outcomes.
No, large language models cannot independently select surgical procedures for glaucoma patients. Although models achieve respectable scores in standard presentations, their reasoning declines sharply in complex cases. Selecting an operative strategy requires physical examination, gonioscopic assessment, and nuanced risk stratification that language models cannot execute. Therefore, licensed ophthalmic surgeons must make all surgical decisions independently to ensure patient safety and avoid vision-threatening complications.
AI models struggle with complex glaucoma cases because these scenarios involve multiple overlapping variables, prior surgical interventions, and atypical anatomical constraints. Standard foundation models rely heavily on broad training data and common textbook scenarios. Consequently, they often fail to synthesize nuanced contraindications, such as extensive conjunctival scarring or neovascular pathology, leading to compromised rationale quality and elevated procedural safety risks during complex decision-making.
Ophthalmologists can safely integrate large language models by using them primarily for administrative assistance, literature summarization, and preliminary educational drafting. Clinicians should never rely on automated platforms for definitive surgical planning without rigorous validation. When evaluating AI-generated treatment suggestions, surgeons must critically cross-examine the underlying clinical rationale against current professional guidelines and independently verify that all patient-specific risk factors have been addressed.
Disclaimer: This content is for informational and educational purposes only and does not constitute medical advice, diagnosis, or treatment. Healthcare professionals should exercise independent clinical judgment. Refer to the latest local and national guidelines for clinical practice.
References

Read summarized clinical updates, watch expert medical content, and earn CME certifications right from your smartphone.


A benchmarking study evaluated ChatGPT, Copilot, and Gemini in glaucoma surgical decision-making. While overall median scores were comparable (~81-83), AI performance dropped significantly in complex glaucoma cases, revealing heightened safety risks and rationale deficits.
Today

A systematic review reveals that microplastics in bottled water cause multi-organ toxicity via oxidative stress, inflammation, and mitochondrial dysfunction, impacting reproductive, hepatic, and vascular systems.
Today

A multicenter Italian registry study evaluated 153 pregnancies in women with multiple sclerosis exposed to anti-CD20 monoclonal antibodies, demonstrating excellent maternal disease control and reassuring fetal safety without heightened risk of major congenital anomalies.
Today

Inadvertent left common carotid artery occlusion during TEVAR demands rapid diagnosis and immediate bailout revascularization to prevent stroke. This case analysis highlights duplex ultrasound detection and direct-access chimney stenting.
Today

A long-term study evaluated progression from knee cartilage biopsy to second-stage MACI. Only 31% of patients underwent implantation at 4.3 years, while 60% of non-implanted patients improved after index chondroplasty. Lower BMI and larger chondral defect size significantly predicted progression.
Today

A 49-year-old man with uncontrolled type 2 diabetes developed a severe MSSA thigh abscess after inserting a continuous glucose monitor on his upper thigh. This case highlights the risks of off-label device placement and the critical role of interdisciplinary care in preventing cutaneous complications.
Today