Information modeling of scientific articles for semantic publishing: theory, models and applications
Mengjuan Weng, Xiaoguang Wang, Ningyuan Song, Cuiyu Qin, Hua JianPurpose
The digital proliferation of scientific articles since the 1980s has made strategic reading an essential skill for researchers. To support automatic filtering, linking and analysis of scientific literature, fine-grained scientific content, extensive semantic links and machine-readable formats are required. This review aims to analyze and compare existing models of scientific articles for semantic publishing and to provide future directions for their development.
Design/methodology/approach
The review followed the PRISMA 2020 guidelines and systematically searched two major databases (WoS and Scopus). A total of 159 articles were screened and synthesized, and key studies were selectively cited to support the narrative review. We synthesized 34 existing models, the theories they employ, and the applications of these models, resulting in 3 theoretical perspectives, 4 types of modeling content and 4 types of applications.
Findings
The review categorizes model content into four orthogonal aspects based on the granularity of textual units: bibliographic records, textual structure, discourse structure and entity types and relationships. The review reveals three main theoretical perspectives on modeling: the scientific paper as an instance of a text model, the scientific paper as a genre of scientific discourse and the scientific paper as an argumentation of scientific claims. Different types of information models are found to serve distinct purposes.
Originality/value
This review provides direction for the automated processing and analysis of scientific papers and reference for the modeling of other types of scientific documents.