Real-Time AI-Based Assessment of Glottic Lesions During Flexible Laryngoscopy: Comparison with Expert Evaluation and Histopathology
Johannes Kränzlein, Konstantinos Anagnostopoulos, Christoph Arens, Nikolaos DavarisBackground: Endoscopic risk assessment of glottic lesions is examiner-dependent. This prospective single-center diagnostic evaluation study assessed the real-time analytical performance of the CE-MDR Class IIb-certified Zeno AI system during structured flexible laryngoscopy. Methods: A total of 109 patients underwent protocol-based flexible laryngoscopy. Post-epiglottic visualization of the glottic level and additional acquisition characteristics were documented systematically. Three experienced examiners independently classified each case as morphologically unremarkable, benign, or malignant/suspicious; majority vote served as the expert reference. Zeno AI classified findings as unremarkable, benign, malignant, or uncertain. The system was evaluated in shadow mode, and AI outputs were included only when the same classification remained stable for at least three seconds. The primary analysis assessed agreement between Zeno AI and the expert majority using Cohen’s κ. Interobserver agreement among the examiners was assessed using Fleiss’ κ. Diagnostic performance against histopathology was evaluated secondarily in operated patients using a primary intent-to-diagnose analysis and a secondary definitive-output analysis. Results: Expert majority classified 64/109 findings as unremarkable, 34/109 as benign, and 11/109 as malignant/suspicious. AI classified 65/109 as unremarkable, 26/109 as benign, 12/109 as malignant, and 6/109 as uncertain. Excluding uncertain outputs, AI and the expert majority agreed in 99/103 cases (96.1%; κ = 0.927). Histopathology was available in 36 patients, with 27 findings negative for malignancy and 9 positive for malignancy. In the primary intent-to-diagnose analysis, treating uncertain outputs as test-negative, sensitivity was 88.9%, specificity 92.6%, PPV 80.0%, and NPV 96.2%. In the secondary analysis restricted to definitive AI outputs, sensitivity was 100%, specificity 91.7%, PPV 80.0%, and NPV 100%. No finding positive for malignancy was classified as benign; one high-grade dysplasia was classified as uncertain. Conclusions: Zeno AI showed high agreement with expert assessment during structured flexible laryngoscopy. However, its potential clinical benefit was not evaluated in this shadow-mode study. Uncertain outputs should be regarded as non-definitive and require expert assessment, while histopathology remains essential for definitive diagnosis.