
Loading, please wait...

Loading, please wait...

Programmed death-ligand 1 (PD-L1) expression is an essential predictive biomarker that dictates first-line and maintenance immunotherapy regimens in advanced non-small cell lung cancer (NSCLC). Pathologists quantify this expression through the tumour proportion score (TPS), which calculates the percentage of viable neoplastic cells exhibiting membranous staining. However, traditional visual evaluation remains susceptible to substantial interobserver discordance, staining variability across antibody clones, and reader fatigue. To address these systemic hurdles, artificial intelligence (AI) models have emerged as decision-support instruments. Understanding the comparative landscape of automated PD-L1 TPS scoring is therefore critical for clinical laboratories and oncology teams seeking to standardize patient stratification.
Manual quantification of immunohistochemical slides presents significant analytical complexities for practicing pathologists. Specifically, readers must visually separate malignant epithelial cells from morphologically similar tumor-infiltrating lymphocytes, alveolar macrophages, and stromal fibroblasts. Because immune checkpoint inhibitors rely on stringent therapeutic cut-offs—most notably the 1% and 50% TPS thresholds—minor scoring errors can directly alter patient eligibility for monotherapy versus combination chemo-immunotherapy. Furthermore, preanalytical factors such as formalin fixation quality, biopsy tissue volume, and cytology cell block artifacts exacerbate interobserver variation. Consequently, even experienced pulmonary subspecialists may achieve only moderate concordance near borderline diagnostic thresholds. Automated image analysis platforms aim to overcome these limitations by executing rapid cell segmentation and membrane intensity classification across whole-slide digital images.
Current computational solutions for digital pathology fall into two primary categories: proprietary commercial software and non-commercial academic algorithms. A comprehensive review mapping seven commercial and 13 research-grade platforms revealed distinct reporting priorities between these development environments. Commercial developers consistently emphasized regulatory certifications, user-interface design, laboratory information system compatibility, and streamlined reader-assist modes. Conversely, non-commercial research groups provided deeper transparency regarding convolutional neural network architectures, open-source code repositories, and training supervision methods. However, both domains exhibited notable shortcomings in reporting granular performance at critical clinical decision cut-offs. Thus, clinicians must carefully differentiate between marketing claims and independently validated peer-reviewed evidence before clinical deployment.
To systematically appraise available algorithms, investigators formulated a comprehensive six-level evaluation structure. This framework examines data characteristics, data preparation standards, supervision methods, performance metrics, generalizability, and clinical deployment traceability. At the foundational levels, algorithms require rigorous training sets that feature multi-pathologist consensus panels as ground truth. Furthermore, the framework assesses whether tools employ supervised, weakly supervised, or self-supervised learning approaches. Importantly, analytical evaluation must advance beyond basic slide-level correlation coefficients to report Cohen's kappa values and categorical agreement precisely at the 1% and 50% decision boundaries. Ultimately, this structured appraisal enables laboratories to identify algorithm strengths and address potential diagnostic blind spots.
A critical finding from recent systematic reviews is that AI diagnostic performance is not automatically portable across different laboratory ecosystems. Discrepancies arise because deep-learning algorithms remain sensitive to variations in whole-slide scanner optics, digital compression formats, and distinct antibody clones such as 22C3, SP263, and 28-8. Moreover, differing patient demographics and histological sub-types alter baseline tissue architecture, which can impair algorithmic accuracy. Therefore, a platform that demonstrates excellent performance in a controlled vendor trial may experience degraded accuracy when introduced into a different hospital setting. Consequently, national and international pathology guidelines discourage universal algorithm rankings, emphasizing instead that each facility must conduct local verification using its own routine specimen workflows.
Safe clinical integration requires an assistive human-in-the-loop paradigm rather than fully autonomous scoring. In this collaborative workflow, the AI platform pre-populates region annotations, segments malignant clusters, and flags ambiguous borderline cases for closer human examination. As a result, pathologists can rapidly review flagged cells, correct misidentified histiocytes, and make final diagnostic determinations with greater speed and reproducibility. Furthermore, healthcare institutions must institute ongoing quality assurance programs to monitor software drift following scanner recalibrations or staining batch changes. By standardizing internal validation protocols and fostering transparent algorithmic reporting, modern pathology laboratories can safely harness AI to optimize immunotherapy selection in non-small cell lung cancer.
Artificial intelligence serves as a diagnostic assistance tool to quantify PD-L1 expression on tumor cell membranes in NSCLC. By automating the visual segmentation of malignant cells and excluding non-tumor stromal elements, AI helps pathologists reduce subjective interobserver variability. Consequently, these algorithms support more reproducible classification around critical treatment cut-offs like 1% and 50% TPS, ensuring appropriate patient selection for immunotherapy.
Universal rankings are unfeasible because algorithm accuracy varies significantly depending on specific immunohistochemical antibody clones, digital slide scanner models, and local tissue preparation protocols. Furthermore, published studies frequently use heterogeneous reference standards and disparate statistical metrics. Therefore, an algorithm validated on one specific scanner-assay combination may perform differently in another center, necessitating individualized local verification for each laboratory.
The human-in-the-loop model maintains the pathologist as the final diagnostic decision-maker while leveraging AI to accelerate cell counting and highlight borderline areas. Pathologists can readily review AI annotations, overrule misidentified macrophages, and adjust region boundaries. Consequently, this collaborative approach mitigates algorithmic errors, preserves clinical accountability, and improves both diagnostic precision and overall laboratory turnaround times.
Disclaimer: This content is for informational and educational purposes only. It does not constitute medical advice, diagnosis, or treatment recommendations. Clinical decisions should be guided by direct evaluation of individual patient circumstances, local regulatory approvals, and institutional protocols. Refer to the latest local and national guidelines for clinical practice.
References
1. Webers J et al. AI for tumour proportion scoring of programmed death-ligand 1 immunohistochemistry in non-small cell lung cancer: a review of commercial and non-commercial tools. Histopathology. 2026 Aug 23. doi: 10.1111/his.70268. PMID: 42634001.
2. Choi S et al. Artificial intelligence-powered programmed death ligand 1 analyser reduces interobserver variation in tumour proportion score for non-small cell lung cancer with better prediction of immunotherapy response. Eur J Cancer. 2022;170:17-26.
3. Lantuejoul S et al. Programmed death ligand 1 immunohistochemistry in non-small cell lung carcinoma. J Thorac Oncol. 2020;15(4):499-519.

Read summarized clinical updates, watch expert medical content, and earn CME certifications right from your smartphone.


Quantifying PD-L1 tumor proportion score (TPS) guides immunotherapy in NSCLC. This comprehensive review analyzes commercial and research AI tools, highlighting validation gaps, regulatory readiness, and practical steps for safe clinical integration.
Today

Preventive healthcare is rapidly emerging as the cornerstone of sustainable modern medicine in India. By identifying metabolic warnings early, integrating AI diagnostics, and scaling community screening through initiatives like Ayushman Bharat, clinicians can curb NCDs and protect long-term socioeconomic security.
Today

The Kerala government has launched an initiative to provide 24-hour specialty medical services across district and general hospitals. Covering medicine, surgery, orthopaedics, obstetrics, and paediatrics, the decentralised healthcare model aims to upgrade infrastructure and reduce tertiary hospital congestion.
Today

A contemporary UK study evaluates the link between dental exposure and oral flora infective endocarditis using transoesophageal echocardiography. We review key findings on valvular patterns, pathogen profiles, and antibiotic prophylaxis considerations.
Today

A systematic review confirms that free flap harvesting from paralyzed limbs is safe, reliable, and preserves vital functional tissue in healthy extremities without increasing donor site morbidity.
Today

The Obesity and Metabolic Surgery Society of India and the Endocrine Society of India have released joint clinical protocols. This landmark consensus moves beyond simplistic advice, recognizing obesity as a multifaceted chronic disease and recommending personalized, stage-based medical and surgical interventions.
Today