DOI: 10.1177/29498732261469335 ISSN: 2949-8732

NSORN: Designing a Benchmark Dataset for Neurosymbolic Ontology Reasoning With Noise

Julie Loesch, Gunjan Singh, Raghava Mutharaju, Remzi Celebi

In the field of neurosymbolic computing, systematic evaluation of ontology reasoning systems under noisy conditions remains underexplored. In particular, there is a need for benchmark frameworks that enable reproducible assessment of how neurosymbolic reasoners behave when ontological data are corrupted. Thus, this work aims to develop a mechanism for introducing noise into an ontology, particularly focusing on the Assertional Box, and evaluate the performance of existing neurosymbolic reasoners on commonly used ontologies under varying levels of noise. We developed NSORN (Neurosymbolic Ontology Reasoning with Noise), a first-step benchmark framework that consists of three techniques to introduce noise into ontologies: logical, statistical, and random noise. Logical noise uses logical violations of disjoint axioms and domain/range constraints. While random noise corrupts existing triples by replacing either the subject or object of a triple with a random entity, statistical noise is introduced using graph neural networks to add noisy facts with low-probability scores. We evaluated the performance of existing neurosymbolic reasoners by introducing noise to Pizza , Family , and OWL2Bench ontologies under these noise types with various levels. The resulting benchmarks were tested on two state-of-the-art neurosymbolic reasoners, Box2EL and OWL2Vec* , as well as a purely neural method, R-GCN . We focus on reasoning tasks, such as class membership and object property assertions, to test how these reasoners handle noise. In our experimental setup, R-GCN consistently outperforms Box2EL and OWL2Vec* in robustness to noisy ontological data—even under the most destructive form, logical noise—maintaining stable membership and object property assertion performance while embedding-based models degrade sharply.

More from our Archive