Lakehouse‐Driven Surgical Databases: Scoping Review and Implementation of a Free Flap Database
Louis‐Xavier Barrette, Sanjana M. Velu, Mark Troftgruben, Brendan Dooley, D. Gregory Farwell, Robert M. BrodyABSTRACT
Objective
To review the landscape of cloud‐native analytics platforms and large language models (LLMs) applied to surgical outcomes research; to present the development of an automated head and neck free flap outcomes database as a reference implementation for using lakehouse architectures in the development of clinical databases.
Data Sources
PubMed and Embase from January 2018 through June 2026.
Methods
A scoping review of PubMed and Embase was conducted using five searches combining data architecture (e.g., “data lakehouse,” “cloud‐native”), clinical text extraction (“natural language processing,” “LLM”) surgical outcomes (“postoperative,” “surgical outcome”), and domain‐specific databases (“free flap,” “microvascular reconstruction,” “automated”). Studies were reviewed if they described database development incorporating automated outcomes classification.
Results
Of 311 studies identified, five studies included elements of incremental processing, pattern‐matching‐gated LLM classification, or fault‐tolerant parallel processing for prospective database maintenance. No implementation examples combined all architectural characteristics. We present the development of a production database for microvascular surgery that automatically identifies free flap cases, collects and analyzes operative notes, validates cohort membership through LLM reasoning, and classifies outcomes across multiple dimensions (flap type, indication, flap failure, hardware complications, bony complications, etc.), all operating incrementally as new cases enter the electronic health record and existing cases are updated.
Conclusion
We describe what appears to be the first data lakehouse‐based automated surgical database applied to otolaryngology, detail architectural decisions required for production, and identify critical unresolved challenges including LLM validation standards, data access barriers, and the translation of clinical domain expertise into computational logic.
Level of Evidence
N/A.