High-resolution computed tomography of the temporal bone remains the gold standard imaging modality for evaluating chronic middle ear diseases. Clinicians frequently rely on detailed radiological assessments to differentiate cholesteatoma from non-cholesteatomatous chronic otitis media. Because cholesteatoma exhibits aggressive osteolytic properties, accurate preoperative classification is crucial for guiding surgical planning and preventing intracranial complications. Recently, advanced artificial intelligence tools have emerged as potential diagnostic assistants across various medical disciplines. Specifically, multimodal large language models have demonstrated impressive capabilities in general medical vision tasks. However, their reliability in subtle anatomical regions, such as the middle ear cleft, requires rigorous validation. A landmark study published in Diagnostic and Interventional Radiology evaluated the performance of cutting-edge models in cholesteatoma HRCT diagnosis. The study assessed whether these advanced artificial intelligence systems could accurately process key temporal bone images under zero-shot conditions. Ultimately, the researchers aimed to determine whether modern multimodal models could function as independent second readers in complex otolaryngological imaging cases.
Background and Rationale for AI in Otology
Otitis media represents a significant public health burden globally, leading to substantial morbidity if left unmanaged. Chronic suppurative otitis media can manifest with or without cholesteatoma, requiring distinctly different therapeutic strategies. While non-cholesteatomatous chronic otitis media often responds to conservative management or tympanoplasty, cholesteatoma necessitates complete surgical eradication due to its progressive bone-eroding nature. High-resolution computed tomography provides exquisite anatomical detail of the temporal bone, including delicate ossicular structures, the facial nerve canal, and the tegmen tympani. Consequently, otolaryngologists rely heavily on radiologist expertise to identify subtle erosive changes, soft tissue masses, and scutum destruction.
In recent years, the rapid evolution of artificial intelligence has sparked widespread interest in automated diagnostic support. Multimodal large language models can process both textual prompts and visual inputs simultaneously. Therefore, researchers hypothesized that these systems might assist clinicians in interpreting specialized radiological images. Specifically, automating cholesteatoma HRCT diagnosis could theoretically streamline diagnostic workflows and assist non-specialist readers in remote settings. However, interpreting temporal bone computed tomography requires extraordinary spatial precision, as middle ear structures span only a few millimeters. Evaluating whether general-purpose multimodal models possess sufficient diagnostic accuracy for such intricate anatomy remains a fundamental question in contemporary radiology.
Evaluating Multimodal LLMs in Cholesteatoma HRCT Diagnosis
To rigorously test artificial intelligence capabilities, researchers conducted a retrospective study involving 101 surgical patients. The cohort comprised 48 patients with pathologically confirmed cholesteatoma and 53 patients with non-cholesteatomatous chronic otitis media. For each patient, an expert team selected six anonymized representative key images from temporal bone high-resolution computed tomography. These key images highlighted diagnostic criteria such as soft tissue attenuation, scutum erosion, and ossicular displacement. Subsequently, the researchers evaluated two leading multimodal large language models, GPT-5 and Gemini 2.5 Pro, via their official web interfaces.
The researchers used structured zero-shot prompts to query the models without prior training on temporal bone datasets. To assess reproducibility, the evaluation occurred across two distinct sessions separated by a one-week interval. Human expert performance served as the reference benchmark, with a board-certified head and neck radiologist achieving 96.0% accuracy and almost perfect agreement with surgical findings. In contrast, the multimodal language models demonstrated markedly lower overall performance. In the initial session, GPT-5 achieved an accuracy of 43.6%, whereas Gemini 2.5 Pro reached 49.5%. During the second session, accuracy remained low at 46.5% for GPT-5 and 47.5% for Gemini 2.5 Pro. These results revealed a significant performance gap between human experts and artificial intelligence models in cholesteatoma HRCT diagnosis.
Key Findings: High Sensitivity but Dismal Specificity
A deeper examination of diagnostic metrics revealed a concerning pattern of diagnostic bias across both artificial intelligence models. Specifically, both GPT-5 and Gemini 2.5 Pro exhibited exceptionally high sensitivity, ranging from 83.3% to 97.9% across test sessions. This high sensitivity indicates that the models successfully identified most true cases of cholesteatoma. However, this sensitivity came at the expense of extremely low specificity, which hovered between a striking 1.9% and 13.2%. Consequently, the models repeatedly misclassified benign chronic otitis media as cholesteatoma, yielding a high rate of false positive diagnoses.
Because of this severe imbalance, the balanced accuracy for both models ranged between 0.45 and 0.52, performing no better than random guessing. Furthermore, the positive predictive values were insufficient for clinical decision-making. High sensitivity combined with poor specificity means that if clinicians relied solely on these models, numerous patients with routine chronic otitis media would face unnecessary surgical interventions or unwarranted psychological distress. The models demonstrated a clear tendency to over-call pathology whenever soft tissue attenuation appeared in the tympanic cavity. Thus, achieving accurate cholesteatoma HRCT diagnosis requires diagnostic nuance that current vision-language architectures fail to deliver without specialized fine-tuning.
Inter-Model Discrepancies and Reproducibility Issues
Reliability and consistency are mandatory requisites for any artificial intelligence tool intended for clinical deployment. Unfortunately, the study uncovered notable limitations regarding short-interval reproducibility and inter-model concordance. Between-session agreement for GPT-5 yielded a Cohen kappa coefficient of 0.360, representing only fair reproducibility over a one-week period. Meanwhile, Gemini 2.5 Pro achieved a moderate reproducibility score with a kappa coefficient of 0.485. These variations occurred despite presenting identical key images and identical text prompts in both testing sessions.
Even more troubling was the near-total lack of agreement between the two commercial models. Inter-model agreement was categorized as slight, with kappa values of 0.035 in the first session and 0.086 in the second session. Although McNemar testing showed no statistically significant difference in overall accuracy between the models, their underlying decision-making pathways diverged significantly. One model frequently flagged non-erosive mucosal thickening as cholesteatoma, whereas the other model made inconsistent predictions across sessions. These discrepancies emphasize that general-purpose multimodal large language models lack internal stability when analyzing complex anatomical images. Therefore, relying on these tools without robust standardization introduces substantial diagnostic variability into otological practice.
Clinical Implications for Radiologists and Otolaryngologists
The clinical implications of this research are highly relevant for practicing otolaryngologists, radiologists, and healthcare technology developers. Currently, multimodal large language models cannot function as autonomous or independent second readers for temporal bone imaging. Their inability to distinguish subtle tissue density differences, localized scutum blunting, and ossicular erosion limits their clinical utility. Furthermore, their low specificity could increase unnecessary diagnostic workups and surgical consultations if used in primary care settings.
Nevertheless, these findings should not discourage further research into artificial intelligence applications in otology. Specialized deep learning algorithms, particularly dedicated convolutional neural networks trained specifically on temporal bone volumetric data, have previously shown promising results. General-purpose vision-language models may require extensive domain-specific fine-tuning, retrieval-augmented generation, or high-definition 3D input capabilities before achieving clinical readiness. Until then, verification by an experienced radiologist remains non-negotiable for chronic ear pathology. Clinicians must maintain critical oversight and avoid over-reliance on unvalidated artificial intelligence outputs in head and neck imaging.
Future Directions in AI-Assisted Otological Imaging
To overcome current diagnostic limitations, future artificial intelligence research must focus on architectural adaptation and specialized dataset curation. Temporal bone pathology rarely presents as isolated key frames; rather, radiologist readers assess continuous volumetric CT slices to trace the extent of disease. Providing multimodal models with full 3D datasets, rather than static 2D key images, could significantly enhance spatial contextual understanding. Furthermore, integrating clinical metadata, such as otoscopic findings and audiometric data, might help refine diagnostic specificity.
Additionally, researchers must establish robust validation protocols that test AI models against diverse, real-world clinical datasets. Developing hybrid systems that combine dedicated computer vision networks with advanced language models may bridge the gap between image processing and clinical reasoning. Until such domain-specific innovations are validated through prospective multi-center trials, human radiologist expertise remains the gold standard. Continuous monitoring, rigorous peer review, and strict adherence to established radiological criteria remain essential to safeguarding patient outcomes in otolaryngology.
Frequently Asked Questions
Can multimodal LLMs accurately replace radiologists in diagnosing cholesteatoma?
No, current multimodal large language models cannot replace radiologists in diagnosing cholesteatoma. In recent comparative studies, general multimodal models achieved accuracy rates below 50% on temporal bone CT scans. Although these models demonstrate high sensitivity, their extremely low specificity leads to frequent false positives. Therefore, human radiologist verification remains essential for accurate diagnosis and patient safety.
Why do general AI models exhibit low specificity on temporal bone CTs?
General AI models struggle with specificity because temporal bone CT interpretation requires fine-grained spatial resolution and subtle anatomical discrimination. Soft tissue masses in chronic otitis media often mimic cholesteatoma visually. General vision-language models lack domain-specific pretraining on 3D temporal bone anatomy, leading them to misinterpret benign mucosal inflammation or fluid as erosive cholesteatoma tissue.
How can AI performance in otological imaging be improved in the future?
Future AI performance can improve through domain-specific fine-tuning on large, pathologically validated temporal bone datasets. Additionally, transitioning from 2D key-image analysis to full 3D volumetric CT processing will provide crucial anatomical context. Combining specialized convolutional neural networks with language models, alongside clinical metadata like otoscopy and audiometry, will also significantly enhance diagnostic accuracy.
Disclaimer: This content is for informational and educational purposes only and does not constitute medical advice, diagnosis, or treatment. Always seek the advice of a qualified healthcare provider with any questions you may have regarding a medical condition. Refer to the latest local and national guidelines for clinical practice.
References
1. Yağcı B et al. Diagnostic performance and short-interval reproducibility of multimodal large language models in differentiating cholesteatoma from chronic otitis media using key-image temporal bone high-resolution computed tomography. Diagn Interv Radiol. 2026 Aug 11. doi: 10.4274/dir.2026.264082. PMID: 42577284.
2. Pathak MR, Gautam M, Acharya A, et al. An exploratory study of high resolution computed tomography of temporal bone in chronic otitis media. Nepalese Journal of Radiology. 2022;12(2):7-13.
3. Chavada PS, Khavdu PJ, Fefar AD, Mehta MR. Middle ear cholesteatoma: a study of correlation between HRCT temporal bone and intraoperative surgical findings. Int J Otorhinolaryngol Head Neck Surg. 2018;4(2):450-455.