A Machine Learning Approach to Latent Structure Learning for Zero-Inflated Patent Keyword Count Data
Sunghae JunPatent document–keyword count data are typically high-dimensional, sparse, and dominated by zero entries, which makes it difficult to simultaneously reconstruct keyword frequencies and identify meaningful technological structures. This study proposes a machine learning approach to latent structure learning for zero-inflated patent keyword count data. The proposed zero-gated latent factor model (ZG-LFM) combines nonnegative matrix factorization (NMF) with keyword-specific logistic occurrence models. NMF is used to extract interpretable document–factor and factor–keyword representations, while the occurrence gate estimates the probability that each keyword appears in a given patent document. The method was evaluated in an initial domain-specific case study using a document–keyword matrix constructed from 9434 quantum computing patent documents and 175 keywords, of which 87.60% of the entries were zero. Predictive performance was assessed using root mean squared error, mean absolute error, and the area under the receiver operating characteristic curve across different numbers of latent factors. The experimental results showed that NMF provided more accurate keyword count reconstruction, whereas the proposed model consistently achieved better discrimination between zero and nonzero keyword entries. These findings indicate that latent count reconstruction and keyword occurrence modeling provide complementary information for analyzing sparse patent data. The learned latent factors further revealed coherent quantum computing subdomains, including hybrid quantum–classical execution, quantum machine learning, quantum state measurement and error analysis, quantum cryptography, superconducting chips, quantum circuits, optical control, qubit devices, and optimization algorithms. The proposed framework therefore provides interpretable latent technology structures while improving the identification of keyword occurrence patterns in zero-inflated patent data. These findings demonstrate the feasibility of the framework within the analyzed quantum computing corpus; its generalizability across other technological domains remains to be evaluated.