Abstract Recurrent multiword units (RMUs) are central to language processing, yet their systematic identification and networked organization remain understudied. This study combines corpus-based and network-analytic methods to examine how RMUs contribute to the emergence of constructional schemas. Drawing on a 185-million-word corpus of Taiwan Mandarin, we pursue two aims. First, we propose a quantitative method for identifying cohesive RMUs based on word predictability in context. Second, we model RMUs as a network in which nodes represent RMUs and edges encode structural and semantic similarity, estimated with a state-of-the-art large language model. A comparison with a random sequence network confirms the non-random structure of the RMU network. Analysis of its topology reveals exemplar-based semantic groupings that support higher-level generalizations. These findings highlight RMUs as key building blocks in linguistic categorization, where subgroupings emerge through sequential lexical associations that underlie the formation of grammatical patterns and hierarchical structure.
Alvin Cheng-Hsien Chen (2026) studied this question.