This study advances research on second language (L2) Chinese development by identifying linguistically grounded metrics capable of reliably capturing writing complexity, a central dimension of the widely adopted CAF framework (Complexity, Accuracy, Fluency) for evaluating learner development. While previous work has largely examined the correlation between traditional complexity (lexical, syntactic, and topic–comment–based) features (TCFs) and complexity using statistical inference, the present study adopts a predictive perspective and develops a new set of Dependency Construction–based Complexity Features (DCFs). Built upon well-established insights from Construction and Dependency Grammar, DCFs provide a linguistically informed framework consisting of 52 features that quantify both the diversity and the sophistication of learners’ syntactic constructions. Using a corpus of 6,975 graded compositions spanning three complexity levels, we trained machine learning models (logistic regression and random forest) to predict complexity. Comparative experiments show that DCFs consistently yield stronger predictive performance than traditional TCF measures, with random forest achieving an accuracy of 91.61%. These results demonstrate that DCFs offer a substantially more effective representation of L2 Chinese writing complexity, underscoring their value for automated assessment and for empirical research on L2 Chinese development.
Yin et al. (Thu,) studied this question.