DOI: 10.3390/ai7080321 ISSN: 2673-2688

Benchmarking Normative AI Assistants Under Inconsistent Evidence with Paraconsistent Trace Semantics

Maksim V. Ulizko, Aleksandr V. Chernikov, Ivan V. Tomilov, Natalia F. Gusarova, Aleksandra S. Vatian

Normative AI assistants are increasingly used in domains governed by duties, permissions, prohibitions, exceptions, priorities, and institutional policies. Existing retrieval-augmented generation (RAG) and legal AI benchmarks evaluate answer accuracy, retrieval quality, citation grounding, natural-language inference, clause extraction, or general legal reasoning ability. These dimensions are necessary but insufficient when supplied evidence is incomplete, mutually inconsistent, or defeasible. The objective of this study is to introduce ParaTraceBench, a paraconsistent trace-based benchmarking framework for post-retrieval normative reasoning over fixed evidence packages. Each scenario contains a query, evidence fragments, extracted facts, defeasible rules, typed attack edges, priority relations, an expected conclusion status, and a gold diagnostic trace. The formalism uses evidence-grounded arguments, a single edge-based attack representation, explicit attack-licensing rules, acyclic priority bases with a transitive closure, grounded argument labeling, trace-normal-form alignment, and deterministic scoring. The operational NER metric is explicitly interpreted as inconsistency-conditioned unsupported-conclusion avoidance rather than proof of logical non-explosion. We evaluated the framework using 140 scenarios, external validation on 567 anonymized Russian-language cases from Russian Federation and EAEU-related materials, reasoning-oriented baseline adaptations, five-run prompt-fairness and stability controls, and a deterministic component-dependency audit. On the full external set, the trace-based configuration reached 85.7% answer-status accuracy, 85.5% contradiction-localization accuracy, 94.2% operational NER, 84.1% priority-handling accuracy, and 83.7% belief-revision accuracy. These results indicate that contradiction-aware trace evaluation provides diagnostic information beyond final-answer accuracy under the evaluated fixed-evidence conditions, while not establishing causal architectural superiority, logical non-triviality, or end-to-end RAG performance.

More from our Archive