DOI: 10.12688/f1000research.175934.2 ISSN: 2046-1402

Leveraging Artificial Intelligence (AI) Based Algorithm for Accurate Estrogen Receptor (ER) and Progesterone Receptor (PR) Analysis in Breast Cancer Diagnostics: Potential to be a Crucial Aid in Routine Workflow

Kanthilatha Pai, Brij Mohan Kumar Singh, Chethana Babu Udupa, Madhavi Pai, Vani Verma, Sumit Jha, Kiran Aatre, Purnendu Mishra, Vishwapriya Mahadev Godkhindi, Swati Sharma, Gursevak Singh, Shubham Mathur, Akash Modi, Rajiv Kumar
Background Scoring of estrogen receptor (ER) and progesterone receptor (PR) expression in breast cancer is critical for identifying patients who would benefit with hormonal therapy. Since manual scoring of immunohistochemistry (IHC) is influenced by pathologist experience, fatigue, inter-observer variability, and subjectivity, artificial intelligence (AI)–based algorithms, trained on large datasets can aid to improve diagnostic accuracy. Methodology This study evaluated an AI-based algorithm for ER and PR IHC scoring in 297 ER and 293 PR cases of invasive breast carcinoma and compared the scores with that of pathologists (two senior and two junior) A pre-trained automated algorithm (Mimansa) identified region of interest and provided the scoreswhich was compared with the consensus score of pathologists-ground truth(GT).Concordance was evaluated using Cohen’s kappa and F1 score. Results For ER IHC, GT scores included 169 strong positive, 31 low positive, and 98 negative cases. Agreement with GT was 99% and 98% for senior pathologists, 97% for the AI algorithm, and 95% and 93% for junior pathologists. The algorithm correctly classified all strong positive cases but showed discordance in 16 low-score cases, with four false negatives and ten false positives. Notably, it identified two true positive cases missed by all pathologists. For PR IHC, agreement rates were 98% and 97% for senior pathologists, 92% for the algorithm, and 93% and 91% for junior pathologists. The algorithm achieved perfect accuracy in strong positive cases but produced 16 false negatives and eight false positives among low-score cases. Cohen’s kappa values were 0.91 (ER) and 0.84 (PR). Conclusion: The AI algorithm demonstrated high concordance with expert consensus, performing comparably to senior pathologists and outperforming junior pathologists in several metrics. It shows promise as a supportive second-reader tool, particularly in low-positive cases where diagnostic errors may significantly impact patient management.

More from our Archive