Can Publicly Available Information Predict the Popularity of Library Materials? A Machine Learning-Based Approach Using Open Loan Data from Public Libraries in South Korea
Jong Hwan Suh, Minseok Kim, Kyuhwan KongRecommendation systems in public libraries rely on loan data skewed toward past popularity, making it difficult for unborrowed and newly published titles to reach users. Hence, we propose and evaluate a machine learning-based approach predicting library material popularity using open loan data from South Korean public libraries. Three feature sets were constructed: word2vec-based title embeddings (F1), borrower demographic features (gender and age group; F2), and topic features from the Korean Decimal Classification (KDC) main class (F3). Seven machine learning models were evaluated using title-level grouped cross-validation, and XGBoost was selected as the best-performing model. Using this model, the effect and marginal contribution of the feature sets were examined via pairwise t-tests on title-level grouped cross-validation repeated 30 times. Consequently, the full feature set F outperformed all two-feature-set combinations, and F3 emerged as a key feature set. The feature sets were consistently ranked F3 > F2 > F1 in both predictive performance and model fitness. A cold-start evaluation confirmed near-identical performance for entirely unseen titles. Thus, library material popularity can be predicted using only publicly available information, suggesting the feasibility of a privacy-preserving approach to informing library material recommendations, relevant to library use, digital inclusion, and social justice research.