Abstract A wealth of linguistic data has been annotated by corpus linguists, and this extant annotated data can be used to automatically replicate and apply the linguist’s annotation scheme by means of machine learning models. This paper accompanies the release of documented code notebooks, which allow corpus linguists to use manually categorized examples or ‘training data’ as input for a predictive language model. By means of a case study of Early Modern English - ing forms, we describe how the predictive language model MacBERTh can be used to accurately replicate the manual data classification scheme employed in previous corpus linguistic studies. Additionally, we discuss how manual error analysis and post-correction may help improve the model’s output. By openly releasing the data and code used in this paper, we hope to stimulate the use of machine learning models such as MacBERTh in corpus linguistics.
Fonteyn et al. (2025) studied this question.