Exploiting Natural Information Redundancy for Reliability Control in Electronic Document Management: A Formal Framework and an Error-Annotated Corpus
Abdinabi Mukhamadiyev, Isroil Jumanov, Khusan Karshiev, Rustam Rakhimov, Akmaljon Abdumalikov, Erkin HafizovElectronic document management systems rely on code-redundancy methods that protect the transport and link layers, yet most errors in administrative documents arise at the presentation and application layers, where a document can be syntactically valid but semantically, logically or arithmetically wrong. This paper proposes a framework that uses the natural redundancy already present in document information—statistical, logical, semantic, structural and technological—as a single resource for reliability assessment. Reliability is distributed across a five-level hierarchy of document elements and aggregated by validity-weighted estimation with a dispersion-sensitive correction. We define proximity functions over attribute elements, derive a probabilistic matching model whose outputs form a proper probability distribution, and show that the running time is linear in the number of documents and attributes under bounded error enumeration. We also release and characterise a benchmark corpus of 4300 administrative documents with programmatically generated errors of six types. The corpus shows no evidence that error risk varies across document types, departments or time, but two pairs of error types co-occur more often than chance, which supports assessing error classes jointly. The recorded metadata carry no predictive information about errors. The framework is specified and analysed; its detection performance is left to future work.