Accelerating metadata annotation in collaborative research centers: A hybrid AI workflow for biomedical entities
Manuel Watter, Felix Engel, Aref Kalantari, Claudia Giuliani, Karin Schuller, Claus-Werner Franzke, Markus Sperandio, Harald Binder, Klaus KaierAbstract
Background
Collaborative Research Centers rely on FAIR-compliant, richly structured metadata, yet manual annotation is a major bottleneck. We implemented a search-augmented large language model (LLM) workflow within a local research data management system to pre-annotate biomedical entities, using human-in-the-loop verification to ensure data quality.
Methods
The pipeline uses Gemini 3 Pro in a two-step prompting strategy: (1) identify dataset deposits and stable identifiers in articles converted to Markdown; (2) extract structured fields from curated repository landing pages rendered via a headless browser. To handle a highly hierarchical metadata schema, we flattened the schema for prompting and remapped outputs to strict JSON with granular provenance tags. Authors received pre-filled metadata and could accept, edit, or delete entries (TP, FP, FN mapping). Performance metrics (precision, recall, F1) were estimated as proportions and synthesized via random-effects meta-analysis.
Results
Among 51 screened articles, the LLM identified a repository deposit in 31; authors responded for 17 (55% of articles with an identified deposit; 33% of all screened articles), yielding 39 verified datasets. False positives were rare (mean 0.23) and false negatives low (mean 1.46). Precision was consistently high at 99.65% (95% CI 98.42%–100.00%). Recall showed more variability at 93.75% (95% CI 89.79%–96.96%), and the F1 score reached 96.17% (95% CI 93.78%–98.11%).
Conclusions
The hybrid workflow achieved high author-verified precision with moderate recall, shifting effort from drafting to reviewing while maintaining schema compliance. As these metrics rely on author acceptance and response rates were modest, broader, multi-site validation is needed.