DOI: 10.3390/electronics15194488 ISSN: 2079-9292

A Study on Methods for Evaluating the Quality of Datasets for Intelligent Algorithm Evaluation

Lu Bai, Yuqing Gu, Wei Zhang, Zilong Liu

The credibility of artificial intelligence (AI) algorithm testing and evaluation directly determines the assessment of their intelligence level and application reliability, which heavily depends on the quality of input datasets. Currently, datasets for AI algorithm testing generally suffer from unclear sources, arbitrary annotations, and heterogeneous formats, while lacking quantitative evaluation specifications that integrate metrological characteristics with algorithm requirements, leading to poor comparability and low credibility of evaluation results. To address these gaps, this study proposes a dataset quality evaluation method oriented to AI algorithm testing by integrating metrological traceability and measurement uncertainty control concepts. Firstly, following the principles of scientificity, quantifiability, and traceability, a hierarchical technical framework and a full-process technical roadmap were constructed. Secondly, a five-dimensional quantitative evaluation index system was established, covering repeatability, completeness, accuracy, consistency, and traceability, with an accompanying uncertainty evaluation, where weighted scoring and measurement uncertainty evaluation methods were introduced to realize quantitative classification of dataset quality. Verification experiments on a temperature and humidity meter image dataset show that the proposed method can effectively assess and grade dataset quality levels. The evaluated dataset obtained a score of 0.9765 ± 0.010 (k = 2), reaching the “excellent” grade, demonstrating the framework’s ability to recognize high-quality datasets. This proof-of-concept demonstrates the framework’s feasibility for assessing high-quality datasets. Formal Standard Reference Data certification is beyond the scope of this study and requires further validation. This research provides a feasible path for the standardized construction of candidate reference data in the field of digital metrology and holds significant value for promoting the standardized application of AI technologies in critical domains such as measurement and control.