DPBERT: A domain-specific pre-trained model for Chinese data policy texts and its interpretability
Zhang Tao, Zhang Ce, Wang Hangong, Ma Haiqun, Jiang LeiChinese data policy texts are characterised by dense terminology, standardised expressions and complex domain-specific semantics, which pose challenges for general-purpose language models. This study develops Data-Policy BERT (DPBERT), a domain-specific pre-trained model tailored to Chinese data policy texts. Using Chinese-bidirectional encoder representations from transformer-whole-word masking as the backbone, we conduct continued pre-training on 65,498 Chinese data-related policy documents with two masking strategies: masked language modelling and whole-word masking. The resulting models are evaluated on policy text classification, named entity recognition and an interpretability analysis based on gradient-based saliency. Experimental results show that DPBERT-whole-word masking outperforms the baseline models on both downstream tasks, indicating that whole-word masking is better suited to capturing Chinese policy terms and composite concepts. We further construct an interpretability evaluation framework using gradient-based saliency and a feature salient value metric to examine how models attend to core policy tokens. DPBERT-whole-word masking exhibits more concentrated attention on high-saliency policy features. DPBERT and Large Language Model Meta AI are compared under the same Chinese data, fine-tuning procedure and evaluation metrics; the results indicate that DPBERT achieves better performance on policy text classification and named entity recognition within this study’s task scope. This work provides a reference for semantic modelling, entity recognition and interpretable analysis of Chinese data policy texts.