DOI: 10.1515/jib-2025-0058 ISSN: 1613-4516

The DSMZ Digital Diversity Annotation Hub: a pipeline for database expansion via text mining and human curation

Emanuel Quadros, Lorenz C. Reimer, Julia Koblitz

Abstract

Maintaining scientific databases that depend on continuous curation of research literature often requires labor-intensive, slow, and error-prone annotation processes. To address these challenges, we present a pipeline that integrates text mining with expert supervision to support database expansion. Using the BRENDA enzyme database as a case study, we compiled a relation extraction dataset by aligning document-level annotations with literature references through distant supervision. We then developed a neural model that performs entity recognition and relation classification, enabling the extraction of enzyme-strain associations from full-text articles. To close the loop between machine learning and expert curation, we designed a web-based interface that allows annotators to review and refine predicted relations. While preliminary, our initial experiments show the potential of combining weak supervision and human-in-the-loop validation to accelerate the integration of literature-derived information into knowledge bases.

More from our Archive