A Stacking‐Based Hybrid Machine Learning Framework for Detecting Cost‐of‐Sales Manipulation in Data‐Constrained Financial Statements
Aysel Topsir, Melih Agraz, Selcuk CebiABSTRACT
Detecting financial statement fraud in emerging markets is difficult due to scarce labeled data and severe class imbalance. This study develops a stacking‐based hybrid machine learning framework for detecting cost‐of‐sales manipulation in firms listed in Türkiye. Rather than proposing a general fraud/non‐fraud detection model, the study focuses on fraud‐type‐specific detection for cost‐of‐sales manipulation under data‐constrained conditions. Because publicly labeled fraud cases are limited, a Türkiye‐specific semi‐synthetic dataset is constructed from KAP financial disclosures, in which unmodified financial statements are combined with manipulated versions generated through rule‐based, audit‐driven manipulation scenarios and domain expert knowledge. The proposed hybrid architecture combines diverse base learners through stacking to capture both linear and nonlinear structures in financial ratios, with the aim of improving generalization and reducing misclassification. Empirical results show that the stacking hybrid model yields higher point estimates than the strongest baseline model across most evaluation metrics and markedly reduces false‐positive classifications. Paired bootstrap confidence intervals support the improvement in fraud‐class precision, whereas the differences in the remaining metrics fall within sampling uncertainty. The resulting error distribution is therefore more favorable for audit planning and decision support.