The DSMZ Digital Diversity Annotation Hub: a pipeline for database expansion via text mining and human curation
Emanuel Quadros, Lorenz C. Reimer, Julia KoblitzAbstract
Maintaining scientific databases that depend on continuous curation of research literature often requires labor-intensive, slow, and error-prone annotation processes. To address these challenges, we present a pipeline that integrates text mining with expert supervision to support database expansion. Using the BRENDA enzyme database as a case study, we compiled a relation extraction dataset by aligning document-level annotations with literature references through distant supervision. We then developed a neural model that performs entity recognition and relation classification, enabling the extraction of enzyme-strain associations from full-text articles. To close the loop between machine learning and expert curation, we designed a web-based interface that allows annotators to review and refine predicted relations. While preliminary, our initial experiments show the potential of combining weak supervision and human-in-the-loop validation to accelerate the integration of literature-derived information into knowledge bases.