Uncovering lexical semantic variation in Mandarin varieties based on word embeddings: the case of Hong Kong written Chinese
Hongzhi Xu, Yuanbing Zhao, Xueyi Wen, Jingxia LinAbstract
Identifying lexical semantic differences across closely related language varieties remains methodologically challenging, as traditional approaches often rely on intuition, manual inspection, or researcher-constructed word lists with limited coverage. This study proposes a computational framework that uses word embedding models to automatically detect lexical semantic variation in large comparable corpora. Using Hong Kong written Chinese as the primary case and Mainland China written Chinese as a point of comparison, we show that the approach not only recovers well-documented differences reported in previous research, but also identifies less salient cases that are easily missed by manual methods. The results demonstrate that the proposed framework provides broader empirical coverage, reduces reliance on subjective discovery, and scales effectively to large datasets. More broadly, the study highlights the value of embedding-based methods for investigating lexical semantic variation across language varieties.