This study investigates the applicability of domainspecific pretraining for classifying Chinese online petitions, with the objective of enhancing governmental operational efficiency and facilitating citizen engagement. Utilizing data from China’s “Message Board for Leaders” platform, the research examines how domain-specific text corpora enhance the semantic understanding and classification performance of a Chinese BERT language model for petitions, called CnPBERT. The approach taken consists in further pretraining the Chinese BERT model on a large, domain-relevant dataset followed by fine-tuning for petition categorization with varying training depths. Experimental results indicate that CnPBERT outperforms the standard BERT-base-Chinese model and demonstrates superior computational efficiency, particularly when only adding a classification head to the encoder stack, rendering it highly suitable for resourceconstrained government applications. The study’s findings have implications for improving the processing and response times to citizen petitions, potentially leading to more effective governance and increased public engagement. Furthermore, the success of this domain-specific approach suggests its potential applicability to other areas of e-governance and public administration.
Zhang et al. (Wed,) studied this question.