DOI: 10.3390/bdcc10080269 ISSN: 2504-2289

From Unstructured Reports to Exploratory Causal Modeling: A Modality-Aware AI Pipeline for Infrastructure Delay Analysis

Florence Gundidza, Masato Kikuchi, Tadachika Ozono

Infrastructure project reports contain rich narrative evidence on delay causes, yet transforming such unstructured text into reliable causal knowledge remains challenging because reports mix confirmed events with hypothetical, conditional, or localized statements. This study proposes an eight-stage computational pipeline that converts infrastructure project evaluation reports into a Bayesian-network model for exploratory structure learning and probabilistic dependency modeling. The central methodological contribution is a modality-aware extraction layer that distinguishes confirmed, project-wide delay evidence from conditional, hypothetical, or component-level statements before causal analysis. The pipeline was evaluated on 55 road infrastructure project reports financed by the Asian Development Bank, the African Development Bank, and JICA, from which delay events across 15 cause categories were extracted and stratified by epistemic modality and scope. Ablation analysis shows that the principal dependency structure recovered by the Bayesian network is not recoverable without modality-aware filtering, indicating that evidence-quality stratification materially shapes downstream causal-structure exploration. Among the recovered dependencies, a financial-to-project-management pathway was the most consistent signal: its undirected skeleton edge was the only relationship recovered by all four causal-discovery algorithms tested (with the orientation determined only by the score-based search), its association was nominally positive—though weak and not uniformly discernible—across nine extraction models spanning three commercial vendors and open-weight families, and it is consistent with prior delay-factor literature. Its model-based scenario contrast (ΔP=+0.638, 95% CI [0.470,0.764]) is reported as hypothesis-generating rather than as a validated policy effect: under structure-learning uncertainty, the interval extends to zero, and the effect magnitude and the specific learned edge depend on the extraction model and the small effective sample. These findings suggest that incorporating modality awareness into narrative-evidence extraction improves the reliability of exploratory causal-structure analysis from infrastructure project reports.

More from our Archive