DOI: 10.1371/journal.pone.0353137 ISSN: 1932-6203

A mechanism-guided simulation framework for synthetic tax-risk data: Constraint validation and reproducible benchmarking

Xiaojing Fan

Synthetic benchmarks can support tax-risk method development when administrative records are inaccessible, but a simulator can also predetermine the patterns later reported as detection findings. We present the Tax Data Generation Model (TDGM), a mechanism-guided simulator that enforces specified income-statement and balance-sheet identities, generates a three-year panel for 10,000 firms, and injects four prespecified risk strategies. This study uses no administrative audit data; its purpose is to test whether the implementation reproduces its design assumptions and to provide a reproducible benchmark. All predictive evaluations used five repeated 70/30 group splits by firm, so no firm appeared in both training and test data. Observed marginal effects were heterogeneous rather than uniformly small: for all risky observations, Cohen’s

d
was −1.29 for profit margin, −1.20 for tax burden, 1.05 for the transfer-pricing indicator, and 0.15 for the income-gap ratio. In a controlled cross-sectional comparison, raw CTGAN failed both row-level accounting identities, whereas deterministic projection restored 100% validity but worsened dependence fidelity. Under the revised classifier specification, a one-hidden-layer neural network with conventional L2 regularization achieved mean ROC-AUC 0.9942 and PR-AUC 0.9710; XGBoost achieved 0.9714 and 0.8975, respectively. Removing the income-gap ratio changed XGBoost ROC-AUC only from 0.9714 to 0.9702, while restricting the model to raw Level 1 variables reduced it to 0.8941. At 1% prevalence, XGBoost ROC-AUC remained 0.9720 but PR-AUC fell to 0.6703, showing why prevalence-sensitive metrics are required. Together, the three contributions are a constraint-by-construction simulator, a reusable audit protocol for mechanism-guided generators, and explicit measurement of the accounting-repair/statistical-fidelity trade-off. Because selected flows are changed after stock generation, derived ratios can carry design-induced information. These are internal properties of a synthetic data-generating process, not estimates of real tax-evasion behavior or deployable audit performance.