Abstract This study investigates how Large Language Models (LLMs) encode second language (L2) writing proficiency distinctions compared to human learners, focusing on the structural alignment between synthetic outputs and human developmental patterns. We analyzed CEFR-graded Write therefore, CALF captures core structural proficiency (syntax, lexis, accuracy, fluency) but not discourse-pragmatic qualities or human variability. The analyses provide level-calibrated evidence that LLM texts show clearer distinctions than learners’, under identical prompts, whose development is gradual and overlapping. This positions LLMs as both tool (exemplars, rubric calibration) and challenge (assessment validity and fairness), while offering SLA researchers insight into how proficiency constructs are encoded in human versus model-based writing.
Carlo et al. (2026) studied this question.