An LR-based risk stratification framework for assessing interlaboratory variability in HEp-2 IFA pattern interpretation
Chaochao Zhang, Yingxin Dai, Dan Cao, Rong Chen, Zhiyuan Gao, Yifan Gong, Anni Guo, Xiaowei Huang, Aiping Liu, Tiantian Liu, Ce Shi, Yi Sun, Wenjuan Wang, Hongkun Wu, Jing Yang, Yiwen Yao, Lei Yu, Ying Zhang, Haiyin Zheng, Kaimin Mao, Min Li, Bing ZhengAbstract
Objectives
Interlaboratory variability in HEp-2 indirect immunofluorescence assay (IFA) interpretation remains a major challenge. Conventional agreement metrics quantify disagreement but do not distinguish discrepancies by diagnostic relevance. We aimed to develop a likelihood ratio (LR)-based framework for clinically informed assessment of interpretation variability.
Methods
Pattern-specific positive LRs were derived from 33,690 routine HEp-2 IFA records and used to classify patterns into four risk tiers. A pattern discrepancy score (PDS) was developed to quantify discrepancies according to risk-tier displacement. Interlaboratory variability was assessed using a standardized 200-specimen panel tested by 18 laboratories, and diagnostic performance was explored in six laboratories.
Results
Pattern disagreement rates ranged from 17.5 to 50.0 %. PDS strongly correlated with disagreement rate (Spearman’s ρ=0.936, p<0.001), but laboratories with similar disagreement rates differed in cross-tier discrepancy distribution. Most discrepancies remained within the same risk tier; among discordant interpretations involving high-risk reference patterns, 20.4 % crossed into another tier. Pattern concordance varied markedly, while titer agreement was moderate to almost perfect (weighted κ, 0.569–0.821). Across six laboratories, diagnostic performance varied (Youden Index, 0.391–0.617), but was not significantly associated with interpretation consistency.
Conclusions
HEp-2 IFA interpretation variability differs in both frequency and potential diagnostic relevance. The LR-based framework and PDS complement conventional agreement measures by identifying discrepancies that alter diagnostic risk classification, potentially supporting risk-oriented proficiency assessment and targeted quality improvement.