A Semantic Clustering Framework for Discovering Latent Offense Patterns: A Case Study of Thai Police Records
Krittakom Srijiranon, Tanatorn Tanantong, Nattanon Keeratiwattapong, Nawarerk Chalarak, Usanut SangtongdeeCrime offense descriptions are often recorded as unstructured text, making large-scale analysis and categorization difficult. This study proposes a semantic clustering framework for Thai crime offense descriptions using sentence embeddings, dimensionality reduction, and unsupervised clustering. Two datasets were obtained from Thonglor Metropolitan Police Station and Mueang Nonthaburi Police Station, Thailand. After preprocessing, the datasets contained 962 and 902 unique offense descriptions, respectively. Each description was transformed into a 768-dimensional embedding using SimCSE-PhayaThaiBERT. The embeddings were represented in Principal Component Analysis (PCA) Space and Uniform Manifold Approximation and Projection (UMAP) Space and clustered using K-Means, DBSCAN, HDBSCAN, and OPTICS. The results showed that UMAP Space generally provided more useful clustering results than PCA Space. Although DBSCAN achieved the highest internal clustering scores, it classified most records as noise. In contrast, HDBSCAN provided a more balanced result by maintaining strong clustering quality while retaining more records for interpretation. Qualitative analysis showed that the discovered clusters corresponded to meaningful offense categories. The proposed framework can support exploratory analysis of Thai crime records without requiring manually labeled data.