
Loading, please wait...

Loading, please wait...

Hepatopancreatobiliary malignancies demand complex clinical choices. Multidisciplinary team conferences establish the benchmark for cancer treatment decisions. Recently, investigators have evaluated LLMs in oncology MDT workflows to support clinical documentation and planning. However, clinical specialists require dependable systems before integrating artificial intelligence into high-stakes tumor boards. In particular, clinicians must know whether an artificial intelligence model produces consistent recommendations when given identical patient summaries. This vital operational metric, known as response stability, remains largely unexamined in multidisciplinary oncology settings.
To address this critical gap, researchers conducted a retrospective comparative study at a tertiary cancer center. The investigators systematically evaluated four advanced artificial intelligence architectures using real-world clinical records. Consequently, the team evaluated whether contemporary systems provide reproducible treatment choices across independent sessions. Furthermore, they compared algorithmic recommendations directly against actual human tumor board consensus. Understanding these computational limitations helps clinicians integrate artificial intelligence safely into oncology practice.
The investigators analyzed consecutive hepatopancreatobiliary cancer cases discussed at a single academic center between September 2024 and August 2025. In total, the cohort included 107 complex clinical cases. The team extracted standardized clinical summaries from preconference documentation. This methodology ensured that models reviewed the exact information presented to human specialists. Subsequently, the researchers evaluated four leading commercial systems: GPT-4o, GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5 through standard interfaces.
To evaluate output consistency, the researchers submitted identical patient queries four separate times in distinct sessions. The models selected from predefined institutional treatment options. Furthermore, the investigators quantified stability through discordance rates relative to the initial response. They also calculated reference-free Fleiss kappa values across all four iterations. In addition, the team measured concordance against human tumor board decisions using initial and modal responses. Therefore, this rigorous framework enabled thorough comparisons across distinct algorithmic architectures.
The study demonstrated statistically significant differences in stability across the four evaluated models. Notably, Gemini 3 Pro showed the highest stability, recording an average discordance rate of only 12.8 percent across repeated runs. In addition, this model achieved the highest reference-free agreement with a Fleiss kappa score of 0.737. In contrast, GPT-4o exhibited the highest instability among all tested systems. Specifically, GPT-4o registered an average discordance rate of 30.2 percent and a Fleiss kappa of 0.430.
These pronounced variations highlight a crucial challenge for clinical artificial intelligence adoption. A clinical tool that provides contradictory recommendations for the same patient undermines physician trust. Therefore, response stability represents a fundamental safety threshold for decision support tools. When oncologists consult clinical algorithms, they expect reproducible logic rather than random output variance. Consequently, healthcare organizations must benchmark output stability before adopting generative tools in multidisciplinary cancer workflows.
When comparing artificial intelligence recommendations with multidisciplinary team consensus, the models achieved moderate concordance. Alignment ranged between 48.6 percent and 72.9 percent using initial model responses. However, assessing the modal response increased concordance to between 66.3 percent and 74.5 percent across the cohort. Interestingly, the top-performing model varied depending on whether researchers analyzed initial queries or modal consensus across repetitions.
Furthermore, model accuracy diverged substantially across different therapeutic modalities. The models struggled to replicate human decisions regarding surgical resections, achieving class-wise F1-scores between only 0.400 and 0.520. In contrast, the systems showed much higher concordance for systemic medical oncology. Specifically, class-wise F1-scores for chemotherapy recommendations reached between 0.621 and 0.836. Surgical decisions involve delicate anatomical assessments that brief text summaries rarely convey. Chemotherapy protocols, conversely, follow structured clinical guidelines. As a result, language models align better with medical oncologists than hepatobiliary surgeons.
Complete discordance between artificial intelligence suggestions and human decisions occurred in 17 cases, representing nearly sixteen percent of the study cohort. Interestingly, complete discordance never occurred in anatomically unresectable disease. Multivariable regression identified three independent clinical factors associated with complete discordance. First, recurrent or on-treatment cancer increased the odds of complete disagreement over fivefold. Second, primary pancreatic tumor location carried a sevenfold higher risk of discordance. Third, cases with low agreement among human clinicians increased discordance more than tenfold.
These findings illustrate the clinical boundaries of contemporary language models. When patients present with recurrent disease or borderline pancreatic tumors, clinical decisions demand extensive multidisciplinary nuance. Similarly, cases that create debate among human specialists also destabilize artificial intelligence reasoning. Consequently, algorithms falter precisely where expert human judgment matters most. Therefore, artificial intelligence cannot replace multidisciplinary discussions in challenging oncological scenarios. Instead, high-risk cases mandate direct specialist evaluations to ensure optimal patient outcomes.
These study findings provide vital insights for oncologists evaluating artificial intelligence solutions. Large language models offer valuable utility for data extraction, literature synthesis, and preliminary guideline checks. However, notable response instability highlights why autonomous decision-making remains unsafe. Relying on a single query risks misleading recommendations due to model variance. Therefore, clinical developers should incorporate multi-run consensus checks before presenting suggestions to clinicians.
In addition, health systems must implement automated flags for complex clinical presentations. When encountering pancreatic neoplasms or recurrent malignancies, software systems should recommend immediate human tumor board review. In resource-constrained settings, artificial intelligence tools can assist regional teams with meeting preparation and administrative summaries. Nevertheless, multidisciplinary cancer care requires clinical intuition that language models currently lack. By establishing strict validation standards, oncology teams can safely explore artificial intelligence while protecting patient care.
Response stability measures whether an artificial intelligence model delivers consistent recommendations when presented with identical clinical patient records. In high-stakes oncology settings, inconsistent responses create dangerous clinical uncertainty and undermine physician confidence. If a model recommends surgery on one attempt and palliative chemotherapy on the next, clinicians cannot trust its utility. Therefore, measuring stability across multiple repeated queries is essential before implementing any artificial intelligence system into clinical tumor boards.
Surgical oncology decisions depend on complex anatomical evaluations, vascular involvement, and individualized patient performance scores that standard text summaries cannot fully capture. Conversely, medical oncology choices follow standardized systemic chemotherapy protocols grounded in published national guidelines. Large language models process structured guideline texts effectively but struggle with nuanced anatomical reasoning and dynamic operative risk assessments. Consequently, the models demonstrated significantly lower class-wise agreement for surgical resections than for systemic oncological treatments.
The study identified recurrent or active disease, primary pancreatic tumor location, and low human consensus as the strongest predictors of discordance. Recurrent cancers involve heterogeneous treatment histories that lack straightforward algorithmic protocols. Similarly, pancreatic tumors present challenging vascular borders that require extensive debate among experienced surgeons and radiologists. Furthermore, when human multidisciplinary experts themselves experience difficulty reaching consensus, artificial intelligence models frequently fail to generate reliable or concordant recommendations.
Disclaimer: This content is for informational and educational purposes only and does not constitute formal clinical advice. Healthcare professionals must exercise independent clinical judgment when managing patient care. Refer to the latest local and national guidelines for clinical practice.
References

Read summarized clinical updates, watch expert medical content, and earn CME certifications right from your smartphone.


A feasibility study evaluates LLM response stability and concordance in hepatopancreatobiliary oncology tumor boards. While models achieved up to 74.5% concordance, stability varied widely, with high discordance in recurrent cases, pancreatic tumors, and scenarios where human specialists debated choices.
Today

Union Health Minister JP Nadda highlighted patient safety, clinical hygiene, and institutional quality at ABVIMS and Dr RML Hospital on World Patient Safety Day 2026. The national strategy emphasizes NQAS certification, Ayushman Bharat expansions, PMBJP generic access, and safe lifelong management of chronic diseases.
Today

Prelacteal feeding persists in South Asian communities despite known clinical risks. This review examines determinants from Chitral, pathophysiological hazards, and pediatric strategies to eliminate harmful feeds and support exclusive breastfeeding.
Today

A randomized controlled trial demonstrates that theory-driven personalized mobile messaging reduces the risk of consecutive walking goal failures by 33 percent, highlighting the vital role of digital health tools in sustaining long-term exercise habits and chronic disease prevention.
Today

A comparative cohort study demonstrates that patients undergoing total hip arthroplasty after prior hip arthroscopy achieve comparable functional recovery, low complication rates, and equivalent implant survivorship when conversion occurs after six months.
Today

A cross-sectional study of 320 adults evaluated the link between habitual fasting duration and dietary quality using Nutrition Quotient scores, highlighting the clinical necessity of balancing time-restricted eating with nutrient density and meal quality in routine metabolic practice.
Today