The Gemini Eye in Microsurgery: Video-Based Capillary Refill and Chromatic Assessment for Free Flap Monitoring
Ebru Aşiret, Burak Yaşar, Büşra Taş Efe, Süleyman Ege Tozan, Hasan Murat Ergani, Ramazan Erkin ÜnlüBackground/Objectives: Postoperative free flap monitoring relies on clinical assessment of colour, turgor, and capillary refill time (CRT). The high frequency of required assessments renders this process labour-intensive and inherently subjective, with sensitivity dependent on observer experience and fatigue. This study evaluated the feasibility of a large multimodal AI (Artificial Intelligence) model (Gemini 3 Flash, Google AI Studio) applied without any task-specific training or fine-tuning, with each video analysed in an independent session, to establish whether meaningful diagnostic agreement is achievable before any domain-specific training is introduced, with the ultimate goal of supporting the development of a machine-based secondary safety net that augments, rather than replaces, the primary clinical assessment of the responsible surgeon in postoperative free flap surveillance. Methods: One hundred and forty-three postoperative video recordings from 143 different patients who underwent fasciocutaneous free flap reconstruction for extraoral defects were analysed. The cohort was deliberately enriched for pathological cases to ensure adequate representation of each vascular compromise category. Assessment was based on intrapatient comparison between flap and adjacent native tissue, without task-specific training or prior clinical information. AI outputs were compared against consensus assessments of three senior plastic surgeons (two associate professors, one full professor) for four parameters: flap colour, turgor, CRT interpretation, and clinical diagnosis. Agreement was evaluated using Cohen’s kappa (κ). Receiver operating characteristic (ROC) analysis assessed the discriminative performance of the AI confidence score. Results: Statistically significant agreement was demonstrated across all four parameters (p < 0.001 for all): flap colour (κ = 0.613, 95% CI: 0.489–0.737, 80.4%), turgor (κ = 0.676, 95% CI: 0.558–0.794, 83.9%; assessed from visual surrogates rather than direct palpation), CRT interpretation (κ = 0.487, 95% CI: 0.360–0.614, 74.1%), and clinical diagnosis (κ = 0.524, 95% CI: 0.399–0.649, 74.1%). Normal flaps were correctly identified in 81.0% of cases and venous compromise in 69.4%. ROC analysis identified a confidence score percentile cutoff of 89 as the optimal threshold for discriminating compromised from normal flaps, with scores below 89 indicating a compromised (pathological) result (AUC = 0.667, 95% CI: 0.579–0.755; sensitivity 77.0%, specificity 52.4%; overall accuracy 62.9%). Conclusions: A structured-prompted multimodal LLM (Large Language Model) demonstrated statistically significant agreement with the consensus judgement of senior plastic surgeons in postoperative free flap assessment without task-specific training, relying solely on intrapatient visual comparison. These findings constitute a proof-of-concept for AI-based video analysis as a potential adjunctive approach in free flap surveillance, warranting prospective validation with independent clinical outcome data before any claim of clinical utility, with possible broader applicability across clinical domains pending such validation. Key limitations include the absence of independent clinical outcome or confirmation of diagnostic categories, deliberately enriched sampling, a small arterial insufficiency subgroup, and the current absence of a formal data processing agreement for the cloud-based AI platform used.