
Loading, please wait...

Loading, please wait...

The field of drug safety monitoring is currently witnessing a transformative shift as artificial intelligence begins to automate complex evaluative tasks. Traditionally, pharmacovigilance has relied on manual expert review to determine if a specific drug administration caused an adverse clinical event. However, as the volume of medical data grows, the manual process has become increasingly resource-intensive and prone to subjectivity. Recent research has focused on the utility of LLMs in pharmacovigilance to streamline the World Health Organization-Uppsala Monitoring Centre (WHO-UMC) causality assessment framework. This methodology is the global standard for evaluating the relationship between drugs and adverse reactions, yet it requires nuanced clinical judgment that was once thought to be exclusively human.
The integration of Large Language Models (LLMs) like GPT-5.4 and Gemini 2.5 into these safety workflows offers a potential solution to the scalability challenges faced by regulatory agencies and pharmaceutical companies. By utilizing advanced natural language processing, these models can analyze semi-structured data from sources like the FDA Adverse Event Reporting System (FAERS). This capability is particularly relevant for maintaining public health in rapidly evolving medical landscapes. As we explore the findings of recent comparative studies, it becomes clear that while AI cannot yet replace human experts, its role as a supportive triage tool is becoming more definitive. For clinicians and pharmacists, understanding these technological advancements is essential for the future of patient safety and drug monitoring.
To rigorously evaluate how modern AI handles drug safety data, researchers constructed a curated dataset consisting of 55 distinct cases derived from FAERS. These cases represented 337 unique drug-level assessments, ranging from simple scenarios to complex cases involving up to eleven different suspected medications. The goal was to mirror real-world clinical complexity where multiple comorbidities and polypharmacy often cloud the causality of an adverse drug reaction. Before testing the models, human experts established a ground truth through independent assessments, achieving high inter-expert agreement. This provided a reliable benchmark to measure the accuracy of LLMs in pharmacovigilance against gold-standard clinical judgment.
The study employed various sophisticated prompting strategies to optimize the performance of the AI models. These techniques included standard prompting, chain-of-thought (CoT), few-shot learning, and even complex tree-of-thought strategies. By varying the way information was presented to models like Gemini 2.5 Flash and Pro or GPT-5.4, the investigators sought to identify which reasoning path most closely mimics a pharmacist\'s analytical process. The data was formatted into a standardized semi-structured style to preserve the essential elements required for a WHO-UMC assessment, such as temporal relationships, de-challenge results, and alternative causes. This structured approach ensured that the AI had access to all necessary clinical evidence while testing its ability to synthesize that information into a valid causality category.
The results of the comparative study revealed that performance varied significantly depending on the specific model and the reasoning strategy applied. Among the tested platforms, Gemini 2.5 Flash emerged as a top performer, particularly when paired with chain-of-thought (CoT) prompting. This specific combination achieved a Cohen kappa of 0.641, indicating substantial agreement with human experts. Furthermore, when using CoT with self-consistency, the model reached a peak accuracy of 0.804. These findings suggest that LLMs in pharmacovigilance are capable of reaching a level of consistency that approaches substantial agreement, which is a significant milestone for automated medical reasoning systems.
Interestingly, the study noted that while the gains from using advanced prompting strategies like tree-of-thought were present, they were often modest compared to simpler CoT methods. The internal consistency of the models remained remarkably high, with Fleiss kappa scores ranging up to 0.915, reflecting almost perfect agreement across repeated inferences. This indicates that once an LLM identifies a reasoning path for a case, it remains highly stable in its conclusion. For medical educators and safety officers, this consistency is a double-edged sword; while it ensures predictable outputs, it also means that any inherent bias or logical error in the model\'s training could be systematically repeated unless caught by human oversight during preliminary case triage.
One of the most revealing aspects of the performance data was the discrepancy in accuracy across different causality categories. The LLMs performed exceptionally well in identifying cases that were classified as Certain or Unlikely, with F-scores reaching nearly 0.90 for the latter. In these extremes, the clinical evidence is often more definitive, making it easier for the AI to apply the WHO-UMC criteria successfully. However, the models struggled significantly with the Possible category, where the F-score plummeted to approximately 0.293. This specific category represents the most challenging aspect of pharmacovigilance, where a temporal relationship exists but alternative causes or underlying diseases could also explain the event.
The difficulty in assessing Possible causality reflects the limits of current LLMs in pharmacovigilance when dealing with clinical ambiguity. Intermediate causality requires an intricate weighing of probabilities and a deep understanding of pathophysiology that current models may lack. Unlike human experts who can draw upon years of clinical intuition to differentiate a drug-induced reaction from a disease manifestation, AI models tend to be more literal in their interpretation of text. This finding underscores why LLMs are currently viewed as supportive tools rather than autonomous decision-makers. They are excellent at filtering out the noise and identifying clear-cut cases but require human intervention for the complex middle-ground assessments that form the bulk of drug safety work.
For the healthcare landscape in India, the advancement of LLMs in pharmacovigilance holds significant potential for the Pharmacovigilance Programme of India (PvPI). As one of the world\'s largest producers and consumers of generic medications, India generates a massive volume of adverse drug reaction reports every year. The ability to utilize AI for preliminary case triage could alleviate the burden on national monitoring centers and help prioritize the most critical safety signals. Implementing these tools within Indian hospital systems could allow for real-time ADR monitoring, potentially identifying rare side effects in the local population much faster than traditional manual methods allow.
However, the adoption of such technology in India must be approached with a clear understanding of its limitations. Given that LLMs struggle with the Possible causality category, their use should be focused on augmenting the productivity of qualified pharmacologists rather than replacing them. Regulatory frameworks would need to evolve to ensure that AI-assisted reports meet the stringent standards of evidence required by the Central Drugs Standard Control Organisation (CDSCO). Furthermore, training local healthcare professionals on the nuances of AI prompting and output validation will be a critical step. By integrating these models into the existing safety infrastructure, India can enhance its position as a global leader in drug safety and public health surveillance.
The future of drug safety lies in the synergy between human expertise and automated intelligence. While the current study highlights the impressive capabilities of models like Gemini 2.5 and GPT-5.4, it also delineates a clear boundary for their current application. Moving forward, research must focus on evaluating these models using raw, unstructured narrative reports rather than pre-formatted data. Real-world medical notes are often messy, containing shorthand, grammatical errors, and contradictory information. Testing the robustness of LLMs in pharmacovigilance against this raw data will be the ultimate trial for their clinical utility and their readiness for widespread deployment in hospital settings.
As AI models continue to evolve, we can expect improvements in their ability to handle nuanced clinical reasoning and intermediate causality. The next generation of multimodal LLMs may even integrate laboratory results and imaging data to provide a more holistic view of an adverse event. Until then, these tools serve as highly efficient assistants that can help manage the sheer volume of safety data, allowing human experts to focus their attention on the most complex and clinically significant cases. By maintaining a cautious but optimistic approach, the medical community can harness the power of AI to create a safer, more responsive drug monitoring environment for patients worldwide.
The WHO-UMC causality assessment is a standardized method used globally in pharmacovigilance to determine the likelihood that a drug caused an adverse event. It classifies relationships into categories such as Certain, Probable, Possible, and Unlikely based on temporal relationships, clinical documentation, and the presence of alternative explanations. This framework ensures consistency in drug safety reporting across different countries and regulatory agencies, helping to identify and validate new safety signals effectively.
The Possible category is challenging because it involves cases where a drug-event relationship is plausible but could also be explained by the patient\'s underlying disease or other medications. Large Language Models often lack the deep clinical intuition and complex pathophysiological reasoning required to weigh these competing factors accurately. Consequently, while they excel at identifying clear-cut Certain or Unlikely cases, they often falter when faced with the nuanced ambiguity inherent in intermediate causality assessments.
Currently, LLMs cannot replace human experts but serve as powerful supportive tools. While models like Gemini 2.5 show substantial agreement with experts, their limited performance in complex, ambiguous cases means they are not suitable for independent decision-making. Their primary value lies in preliminary case triage and automating the processing of large datasets. Human oversight remains essential to validate AI outputs, especially in cases where clinical judgment and deep medical expertise are necessary to ensure patient safety.
Disclaimer: This content is for informational and educational purposes only and does not constitute medical advice or a substitute for professional clinical judgment. AI tools in medicine should be used as supportive technology under expert supervision. Refer to the latest local and national guidelines for clinical practice.
References
Ha YM et al. Large Language Models for World Health Organization-Uppsala Monitoring Centre Drug-Adverse Event Causality Assessment Using Food and Drug Administration Adverse Event Reporting System Cases: Comparative Performance Study. J Med Internet Res. 2026 Jul 08. doi: 10.2196/93237. PMID: 42418253.
Heckmann NS et al. Biomedical Large Language Models and Prompt Engineering for Causality Assessment of Individual Case Safety Reports in Pharmacovigilance. Pharm Res. 2026 May 23. doi: 10.1007/s11095-026-04121-4. PMID: 42174348.
Patil HG, Khairnar VS. Impact of AI on Pharmacovigilance: A Systematic Review. Int J Res Pharm Allied Sci. 2025 May 31. doi: 10.71431/IJRPAS.2025.4503.
"
Read summarized clinical updates, watch expert medical content, and earn CME certifications right from your smartphone.


A landmark study evaluates the performance of LLMs like Gemini 2.5 and GPT-5.4 in automating WHO-UMC drug-adverse event causality assessments. While showing substantial agreement with experts, AI remains a supportive tool rather than an independent decision-maker in complex pharmacovigilance workflows.
2 weeks back

Andhra Pradesh reported 10 new Covid-19 cases, taking the state tally to 49 while deaths remain at four. With 24 patients hospitalized and 16 under home isolation, the Health Department has intensified monitoring. Medical professionals should review regional distribution, diagnostic protocols, and management plans.
Today

An 11-year Swedish registry study of 618 uterine sarcoma patients found that minimally invasive surgery yielded survival comparable to open surgery in early stages. However, adjuvant chemotherapy conferred no survival benefit in localized or advanced disease, highlighting stage and histology as key outcomes.
3 days back

A cross-sectional study evaluates post-intensive care syndrome in cardiac patients 2-4 weeks post-ICU discharge, highlighting cognitive, psychological, and functional impairments and the need for structured multidisciplinary rehabilitation.
3 days back

Anterior cruciate ligament reconstruction failure lacks uniform definition. A narrative review proposes an integrative framework incorporating objective and subjective instability, persistent pain, restricted motion, graft rupture, and secondary meniscal injury to standardize clinical reporting.
3 days back

With World Obesity Atlas data warning that over 41 million Indian children are overweight or obese, ICMR and NIN have unveiled a 10-point policy roadmap. The initiative calls for mandatory front-of-pack labeling, HFSS taxes, strict marketing bans, and healthier school environments to curb non-communicable diseases.
Today