Ontological Learning Through Topic Modeling: Identifying Conceptual Candidates for Enriching a Leadership Domain Ontology
Carlos Mauricio Zuluaga-Ramírez, Manuela Gómez-Suta, Manuela Idárraga-Grajales, Carolina Gómez-Carrillo, Julio Cesar Chavarro-Porras, Sandra Estrada-Mejía, José Soto-MejíaOntologies enable the construction of formal and reusable semantic representations of domain knowledge; however, keeping them updated in response to the dynamic evolution of concepts remains a challenge, particularly in the leadership domain, where the diversity of approaches, theories, models, and styles generates conceptual ambiguity and information overlap. To address this limitation, this study proposes an ontology learning methodology based on topic modeling (TM) to create an adaptive component capable of identifying candidate concepts for enriching a previously developed leadership ontology. The methodology is structured into four stages: (1) corpus preprocessing; (2) manual labeling of the corpus evaluated through inter-annotator agreement using the Kappa index; (3) vocabulary construction by comparing five statistical weighting schemes (Inverse Document Frequency (IDF), Entropy, and three variants of Ochoa’s proposal) against a gold standard; and (4) sentence classification and concept construction using four probabilistic topic models (Latent Semantic Analysis (LSA), Non-Negative Matrix Factorization (NMF), Latent Dirichlet Allocation (LDA), and Biterm Topic Model (BTM)), taking as input both a corpus fully associated with the initial ontology and one partially associated. The IDF weighting scheme achieved the best balance between precision and recall (F-measure: 67.03%), and a Support Vector Machine (SVM) classifier reached an accuracy of 98.19% in sentence classification. The LDA and NMF models obtained the highest scores for thematic coherence. However, the expert evaluation revealed that the thematic coherence metric alone is insufficient for evaluating the resulting structures, as it favored the LDA model, while the NMF model generated more conceptually diverse structures. The proposed methodology enables the semi-automatic transformation of unstructured text into structured candidate concepts for subsequent ontology enrichment.