The use of artificial intelligence for enhanced detection of urothelial bladder cancer among cystoscopy and urine cytology tests: A systematic review and meta-analysis
Matthew Feyissa, Alex Ng, Kimberley Chan, Buraq Ahmed, Peter Davies, Mohammad Alomari, Yiannis Philippou, Muheilan Muheilan, Jennifer Linehan, Clayton Lau, Niyati Lobo, Yuhong Yuan, Jeremy Teoh, Nikhil VasdevObjective:
Bladder cancer remains a significant burden on healthcare systems worldwide. The aim of this review is to evaluate the diagnostic performance of artificial intelligence against conventional first-line methods (cystoscopy and urine cytology) for bladder cancer.
Methods:
A PROSPERO-registered (CRD420261291622) systematic review and meta-analysis. Studies were included if they assessed artificial intelligence performance in definitive urothelial carcinoma detection via cystoscopy or urine cytology against a non-artificial intelligence human comparator. Bivariate random-effects meta-analysis was performed to assess diagnostic performance with the area under the summary receiver operating characteristic curve calculated from the hierarchical summary receiver operating characteristic curves.
Results:
Nine studies were included (six cytology, three cystoscopies; 8918 data points). Artificial intelligence demonstrated statistically significant greater sensitivity (0.927 vs 0.754), with a lower negative likelihood ratio (0.087 vs 0.254), suggesting stronger ‘rule-out’ performance. However, this came at the cost of higher false positives compared to conventional methods. Conventional methods (cystoscopy/cytology) demonstrated higher specificity (0.968 vs 0.841) and positive likelihood ratio (23.275 vs 5.849), reflecting stronger rule-in capability.
Conclusion:
Artificial intelligence algorithms potentially have enhanced screening performance for bladder cancer compared to first-line modalities. The utilisation of a hybrid model may improve outcomes and efficiencies. However, large-scale, prospective trials with standardised reporting and histological reference standards are required before artificial intelligence can safely and equitably be deployed.
Level of evidence:
2a