DOI: 10.3390/make8100310 ISSN: 2504-4990

A Comprehensive Survey of Missing Data-Handling Methods from a Machine Learning Perspective

Jingxuan Li, Weiping Ding, Xingquan Zhu, Guoqing Chao

With the advent of the big data era, data becomes increasingly important, especially for many machine learning tasks where data quality is vital. However, missing data is inevitable in many real-world applications, making its handling a challenging problem. Although some reviews exist, they either treat it as a statistical problem or focus only on traditional machine learning methods. This survey provides a novel taxonomy of missing data-handling methods. According to feature or label missing, we classify them into feature-missing handling methods and label-missing handling methods. For feature-missing handling methods, we further categorize them into four classes: traditional, deep learning, soft computing, and large language model missing data-handling methods. Each class is further split to introduce representative methods. For missing-label handling methods, we focus on label completion and distinguish active learning, which selects unavailable labels for acquisition from an oracle, from semi-supervised learning, which estimates unavailable labels from labeled and unlabeled data. For each category, we list their advantages, disadvantages, and applications. To facilitate engineers and researchers, some commonly used software is introduced. To promote further development, we point out challenges requiring deeper investigation.