DOI: 10.1145/3837113 ISSN: 2836-6573
From Global Search to Local Lookup: Scalable Knowledge Base Completion via Differentiable Guarded Logic
Yizheng ZhaoState-of-the-art knowledge base completion models, ranging from geometric embeddings to GNNs, achieve impressive predictive performance but share a critical limitation: they act as black boxes that often ignore data consistency, producing predictions that violate rigorous ontological constraints. While neuro-symbolic frameworks attempt to bridge this gap by integrating logical rules, they face a prohibitive scalability barrier. Existing approaches rely on unguarded universal quantification that demands global search, causing combinatorial explosions and execution failures on large-scale knowledge bases. In this paper, we propose
GUARDNET,
a differentiable reasoning framework designed to break this expressivity-scalability deadlock. Our core contribution is leveraging the Guarded Fragment (GF) of first-order logic to fundamentally restructure the computational graph of reasoning. We demonstrate that the GF's syntactic ''guard'' acts as a topological constraint that transforms intractable global quantification into efficient neighborhood-restricted lookups that strictly align logical evaluation with the underlying graph structure. For sparse knowledge bases, this paradigm shift eliminates the quadratic grounding overhead that plagues traditional neuro-symbolic systems, reducing complexity to linear in the number of edges without sacrificing the semantic expressiveness required for complex reasoning. This construction is the differentiable counterpart of the guarded chase in database theory: the guard atom acts as an index that confines grounding to existing tuples, inheriting the PTIME data complexity of guarded tuple-generating dependencies. Experiments on benchmarks with up to 377K concepts show that
GUARDNET
is the first neuro-symbolic framework to scale to such large knowledge bases, succeeding where neuro-symbolic and probabilistic baselines fail to converge within 72 hours, while significantly outperforming state-of-the-art geometric embedding and GNN models in both accuracy & consistency. We release our code at