Enhancing global longitudinal strain accuracy: impact of a web based training program versus fully automated measurement using artificial intelligence
J Sen, B Thampinathan, M Signorile, M Sooriyakanthan, C Yu, T Marwick, P ThavendiranathanAbstract
Background
Global longitudinal strain (GLS) detects early cardiac dysfunction during cancer therapy but is under-utilized. Although AI can aid strain measurements, a skilled human must still measure and verify GLS.
Purpose
To evaluate whether a structured, web-based educational program can improve clinicians’ accuracy, consistency, and diagnostic performance in GLS measurement compared with fully automated AI-driven strain analysis.
Methods
A web-based educational program was delivered in three phases: Learning Phase (strain education, analysis of 10 transthoracic echocardiograms (TTEs) from five breast-cancer patients, and personalized feedback), Consolidation-Application Phase within 4 weeks (20 TTEs: 10 original + 10 new from five additional patients); Retention Phase: 1-year follow-up in a subgroup (20 TTEs). GLS-CTRCD was defined as >15% relative reduction from baseline (present in 3/5 patients during Learning Phase, 2/5 during Application Phase). Expert reads and cardiac MRI served as reference standards. Fully automated AI-driven GLS from three vendors was also compared with the reference. Agreement, diagnostic accuracy, and inter-observer variability were assessed with mixed-effects models, Cohen’s κ, and intraclass correlation coefficients (ICC), respectively.
Results
Among 74 participants from 17 countries, mean absolute GLS error versus reference decreased from Learning Phase (1.42±0.63%) to Consolidation Phase (1.17±0.46%; Δ = -0.25% [95%CI: -0.39,-0.11]) and to Application Phase (1.03 ±0.30%; Δ = -0.39% [-0.53.-0.26]) (p<0.001 for both), with retention of gains at 1 year. Using AI, the mean absolute error was 0.18±1.99%, which was lower on average but with greater variability. κ for identifying GLS-CTRCD rose from Learning (0.71 [95%CI: 0.64, 0.78]) to Consolidation (0.94 [0.90,0.97]) and Application (0.97 [0.95,1.0]) (p<0.001), remaining high at Retention (0.98 [0.94,1.0]). Participants who completed training outperformed AI (κ = 0.80 [0.59,1.0]). Inter-observer ICC increased from Learning (0.62 [95%CI: 0.58,0.65]) to Consolidation (0.72 [0.69,0.75]), Application (0.81 [0.79,0.83]), and remained stable at Retention (0.77 [0.72,0.81]). The ICC of the AI vendors was 0.76 [0.66,0.84], comparable to trained participants.
Conclusions
A web-based educational program markedly improved GLS measurement accuracy, CTRCD detection, and inter-observer reliability, with benefits persisting for at least one year and surpassing current AI performance. Such training may facilitate reliable clinical use of GLS.