We reverse the poverty-of-the-stimulus argument and ask whether both Large Language Models (LLMs) and native speakers learn the robust pattern of NP-internal agreement in Icelandic as the evidence in the input is abundant but not poor. Given recent claims in the literature that LLMs “are better than theoretical linguists at theoretical linguistics, at least in the domain of verb argument structure” (Ambridge & Blything 2024) and that “language models refute Chomsky’s approach to language” (Piantadosi 2024), it is important to investigate LLM performance on a variety of languages, including ones that are represented in the training data of said LLMs, but to a much lesser degree than English. In this paper, we look at NP-internal agreement in Icelandic with a task consisting of grammaticality judgments and translations from English to Icelandic. We compare the results for eleven different LLMs and 188 human speakers. Our findings show that the pattern of NP-internal agreement in native speakers is robust but that models fail to generalize it from clear evidence in the training data, making errors humans generally do not produce. We argue that this shows that claims about LLMs’ linguistic performance need to be evaluated in a variety of (lower-resourced) languages, including those with richer morphology.
Ármannsson et al. (Fri,) studied this question.