Machine learning analysis reveals linguistic features predict perceived reading difficulty in second language learners, highlighting word age-of-acquisition as the primary driver.
Understanding how linguistic features shape text difficulty is fundamental to assessing second language (L2) reading. This study investigates how lexical, syntactic, and discourse-level features jointly influence L2 learners’ judgments of English text difficulty. Eight hundred eighty Chinese university students completed a comparative judgment task, evaluating the relative difficulty of 4,722 English texts from the CommonLit Ease of Readability Corpus. A random forest regression model trained on 19 linguistic features predicted learners’ judgment scores, explaining 49.2% of the variance in the training set (3,305 texts) and 50% of the variance in a held-out test set (1,417 texts). The model was further validated on an independent dataset. Word age-of-acquisition emerged as the strongest predictor of perceived difficulty. These findings underscore the multidimensional nature of L2 text difficulty and highlight the value of incorporating cognitively-grounded linguistic features when modeling L2 readers’ perceptions. The results offer practical implications for L2 materials development and readability assessment.
No takes yet. Share an insight, caveat, or question.
Jiang et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: