PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 20, 2025International Journal of Corpus Linguistics2 citationsOpen Access

Using machine learning to automate data annotation in corpus linguistics

View Full Paper
LFLauren FonteynEMEnrique ManjavacasJRJaleesa De Regt

Key Points

  • The machine learning model accurately replicates the manual annotation scheme of corpus linguistics studies.
  • Using training data, the model successfully applies a predictive language approach to linguistic data.
  • Manual error analysis improves model accuracy, indicating potential for better data classification outcomes.
  • Open release of code and data encourages wider adoption of machine learning in corpus linguistic research.

Abstract

Abstract A wealth of linguistic data has been annotated by corpus linguists, and this extant annotated data can be used to automatically replicate and apply the linguist’s annotation scheme by means of machine learning models. This paper accompanies the release of documented code notebooks, which allow corpus linguists to use manually categorized examples or ‘training data’ as input for a predictive language model. By means of a case study of Early Modern English - ing forms, we describe how the predictive language model MacBERTh can be used to accurately replicate the manual data classification scheme employed in previous corpus linguistic studies. Additionally, we discuss how manual error analysis and post-correction may help improve the model’s output. By openly releasing the data and code used in this paper, we hope to stimulate the use of machine learning models such as MacBERTh in corpus linguistics.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Fonteyn et al. (2025) studied this question.

synapsesocial.com/papers/68d469c831b076d99fa66742https://doi.org/10.1075/ijcl.22088.fon
Ask AI
Helpful
Bookmark
Share
View Full Paper