DOI: 10.25259/cytojournal_192_2025 ISSN: 1742-6413

Automating fine needle aspiration and non-gynecologic quality control using a large language model: A real-world evaluation of diagnostic accuracy and efficiency

Sanskriti Naik, Vikas Nishadham, Christopher Jackson

Objectives:

Cytopathology laboratories are required to track diagnostic category distributions as part of quality control (QC) under Clinical Laboratory Improvement Amendments regulations. This process is typically manual, time-consuming, and prone to inconsistency. We evaluated whether a large language model (LLM) could accurately and efficiently categorize cytopathology reports to support regulatory QC processes.

Material and Methods:

A total of 472 fine needle aspiration and non-gynecologic (NG) cytology reports from 1–31 January 2024 were processed using DeepSeek-R1:32B, a locally deployed open-source LLM. GYN cytology reports were excluded because their diagnostic categories are already captured as discrete data in our laboratory information system, eliminating the need for natural language processing. The model extracted the final diagnosis, diagnostic category, and confidence score from each report. A board-certified cytopathologist independently reviewed all cases to serve as the ground truth. Cases were categorized as correct or incorrect. Time required for manual versus artificial intelligence (AI)-assisted review was also measured.

Results:

The model achieved 98.5% categorization agreement (95% confidence interval: 97.0–99.3%; Cohen’s Kappa = 0.973) with expert review (465/472 cases were classified as correct), with only 1.5% (7/472) raw misclassifications. Most misclassifications were due to semantic ambiguity, missing context, or institutional interpretation conventions. The pipeline implemented a human-in-the-loop mechanism by flagging low-confidence or non-standard outputs for manual review, enhancing safety. The AI-assisted workflow reduced total review time by 98.5%, from 279 to 4.15 min.

Conclusion:

This study demonstrates that LLMs can accurately and safely support cytopathology QC tasks, significantly reducing workload while maintaining reporting consistency. With further development and validation, LLM-powered tools may become integral to pathology workflows across specialties.

More from our Archive