DOI: 10.3390/electronics15194439 ISSN: 2079-9292

Sparse k-Means Clustering with Lasso-Based Feature Selection

Miin-Shen Yang, Shazia Parveen

Clustering high-dimensional data is challenging when irrelevant or weakly informative features obscure the underlying cluster structure. To address this issue, we propose two sparse k-means clustering algorithms, S-KM1 and S-KM2, that perform cluster estimation and feature selection by sparse feature weighting simultaneously. Clustering has been applied to a wide range of disciplines, including image segmentation, social network analysis, medical imaging, market segmentation, anomaly detection, etc. In clustering algorithms, the most popular and widely used method is k-means. However, the k-means algorithm always treats feature (attribute) components in datasets equally. In practice, different features may contribute unequally to the cluster structure and should therefore not necessarily be weighed equally. This has been demonstrated by Yang and Benjamin (2023) with two types of sparse possibilistic c-means (PCM) methods called SPCM1 and SPCM2, using the Lasso concept. Motivated by SPCM1 and SPCM2, we propose two sparse k-means clustering algorithms, S-KM1 and S-KM2, in this paper. The proposed methods S-KM1 and S-KM2 consider the k-means clustering subject to different feature weight constraints based on the Lasso and L1 penalty so that they can shrink the irrelevant features towards zero and achieve sparsity in features. The proposed S-KM1 and S-KM2 are compared with some of the most popular sparsity-clustering techniques on several numerical and real-life datasets using different clustering performance measures. Comparisons and experimental findings demonstrate the validity, efficacy, and effectiveness of the proposed S-KM1 and S-KM2 clustering algorithms, and they also surpass the existing advanced algorithms in terms of efficiency and usefulness.