How Often Do Large Language Models Agree with Each Other—And with the Truth? A Consensus- and Complexity-Stratified Analysis of Data Extraction for Neuroimaging AI
Nafiye Sanlier, Umid Sulaimanov, Ariorad Moniri, Behman Demir, Gular Ismayilova, Melih Yucel Sanlier, Ugur Erginoglu, Ahmed Rasim Bayramoglu, Maryam Sabah Al-Jebur, Simon Gashaw Ammanuel, Erkin Otles, Abdullah Keles, Ufuk Erginoglu, Mustafa K. BaskayaBackground: The reliable integration of large language models (LLMs) into neuroimaging data extraction workflows remains unresolved. Prior benchmarking shows that exact-match accuracy underestimates LLM extraction performance, but whether inter-model consensus and variable complexity can guide automation remains unclear. We evaluated whether inter-model consensus can serve as a confidence signal for human–artificial intelligence (AI) extraction and can guide complexity-stratified workflow triage. Methods: Four frontier LLMs were queried via OpenRouter with an identical zero-shot structured prompt to extract 22 predefined variables from 91 peer-reviewed neuroimaging AI articles, yielding 2002 article–variable items per model. Variables were stratified a priori into low- (n = 7), medium- (n = 8), and high-complexity (n = 7). Performance was compared with an expert reference using exact-match and semantic-equivalence accuracy. Item-level consensus and five triage strategies characterized the efficiency–accuracy trade-off. Results: Semantic-equivalence accuracy converged to 80.5–83.4% across models despite approximately ten percentage-point exact-match differences. Unanimous 4/4 consensus occurred in 45.6% (910/1994) of items, with exact-match accuracy of 85.8%, rising to 95.3% after semantic normalization; however, 14.2% still failed to match the reference. Reliability was complexity-dependent: 96.6% for low-complexity variables, 73.2% for medium-complexity variables, and 38.1% for high-complexity variables. A hybrid strategy auto-accepting 4/4 items and routing 3/4 items to rapid verification reduced estimated review effort by approximately 59%. Conclusions: Inter-model consensus is useful, but it is incomplete and depends on variable complexity. We show that LLM-assisted extraction in neuroimaging AI is a complexity-stratified workflow design problem: low-complexity neuroimaging variables may be selectively automated, while medium-complexity variables require rapid verification, and high-complexity methodological variables should remain human-led.