DOI: 10.1055/a-2920-4203 ISSN: 2377-0813

Artificial vs. Human Intelligence: Evaluating CPT Coding Accuracy in Extremity Microsurgical Reconstruction

Jonathan Childs, Juan Enrique Berner Gomez, Oriana Haran, Manuel Ricardo Ortiz Llorens

Background Current Procedural Terminology (CPT) coding in microsurgical reconstruction requires precise translation of complex procedures into billable codes. Manual coding requires training, experience and it is a time-consuming task, in which mistakes are possible. Artificial intelligence (AI) tools may offer an efficient alternative, but their accuracy in this specialized context remains unclear. Methods We evaluated three leading AI large language models (LLMs): ChatGPT, Gemini 2.5 Flash, and DeepSeek-V3. These were tasked to retrieve CPT codes from 30 microsurgical operative notes, retrieved from our prospective extremity reconstruction database. Using expert-assigned codes as the gold standard, we compared AI model performance via precision, recall, and Jaccard index. Results All AI models demonstrated poor agreement with human coders. ChatGPT performed best but achieved low precision (0.215), recall (0.230), and Jaccard index (0.148). Gemini and DeepSeek showed further reduced accuracy. Notably, models consistently overcoded, erroneously including add-on codes, which human coders deemed unnecessary by institutional standards. Conclusions Current LLMs are unreliable for independent CPT coding in microsurgical reconstruction, risking billing errors and compliance violations. While AI may aid as an adjunct tool, improvements in training data, payer-policy integration, and validation are needed before clinical implementation.

More from our Archive