DOI: 10.1021/acs.jpclett.6c02089 ISSN: 1948-7185

Zipf-like Statistical Regularities in Molecular Sequence Representations for Chemical Language Models

Anyu Liu, Chao Fang, Yuntao Li, Zongguo Wang, Tao Qi, Guoping Hu

Abstract

Large language model (LLM)-based approaches increasingly use molecular strings such as SMILES and SELFIES for molecular generation and property prediction. However, the statistical properties of molecular token distributions have not been systematically characterized. Here, we analyzed rank–frequency distributions of tokens derived by byte-pair encoding (BPE) across large molecular databases. BPE-derived tokens showed approximate Zipf-like rank–frequency scaling across the examined representations and chemical spaces, with fitted slopes moderately steeper than the canonical value of −1, paralleling a statistical pattern widely observed in natural language. Moreover, when BERT models were pretrained using BPE vocabularies with different rank–frequency slopes, the closeness of these slopes to the ideal Zipf value of −1 strongly correlated with performance on molecular property prediction tasks (Pearson r = 0.91, p < 0.001) and remained associated after adjustment for vocabulary size (partial r = 0.86, p < 0.001). Together, these findings show that molecular BPE vocabularies exhibit an approximate Zipf-like rank–frequency regularity and that slope closeness provides an empirical diagnostic for comparing vocabulary sizes within the examined SMILES/SELFIES BPE framework.

More from our Archive