Algorithmic Determinants of Performance Heterogeneity in Whole-Genome Sequencing-Based Prediction of Drug Resistance in Mycobacterium tuberculosis: A Systematic Review and Meta-Analysis
Baozhen Peng, Yang Zhou, Xiangchen Li, Huihui Liu, Bing Zhao, Ping Hou, Xichao Ou, Yanlin ZhaoWhole-genome sequencing (WGS) is an increasingly adopted platform for predicting drug resistance in Mycobacterium tuberculosis; however, diagnostic accuracy varies substantially across bioinformatic tools and analytical frameworks, generating considerable uncertainty for clinical laboratory implementation. We conducted a prospectively registered (PROSPERO: CRD420261342739), PRISMA-DTA-compliant systematic review and meta-analysis of diagnostic accuracy studies. PubMed (MEDLINE), Embase, Web of Science, and Cochrane CENTRAL were searched from 1 January 2000 through 28 January 2026. Primary overall sensitivity and specificity were estimated using a tool-level bivariate random-effects model. Exploratory subgroup analyses and meta-regression examined the association between algorithm category and diagnostic-performance heterogeneity. Twenty-eight drug-level evaluations from seven tools (rifampicin, isoniazid, ethambutol, and pyrazinamide for each tool) were compiled from the extracted 2 × 2 data. For the primary tool-level composite analysis, pooled sensitivity was 0.930 (95% CI: 0.907–0.948) and pooled specificity was 0.962 (95% CI: 0.929–0.981). In secondary drug-specific analyses, sensitivity was highest for rifampicin (0.960, 95% CI: 0.934–0.976) and isoniazid (0.933, 95% CI: 0.906–0.953), and lowest for pyrazinamide (0.860, 95% CI: 0.800–0.904). Exploratory tool-level comparisons produced pooled sensitivity estimates of 0.920 for rule-based tools, 0.899 for machine learning tools, and 0.951 for hybrid tools. These comparisons involved only seven tool-level analytic units and cannot disentangle algorithm type from individual tool identity, training data, mutation catalogue version, or validation population. WGS-based bioinformatic tools provide highly specific and generally sensitive predictions of Mycobacterium tuberculosis resistance for first-line drugs across diverse clinical settings. Exploratory differences between tool categories should not be interpreted as causal effects of algorithmic architecture. Future studies should use prospective head-to-head evaluations on shared, geographically diverse isolate collections, alongside continued improvement of resistance catalogues and external validation.