By positing a relationship between naturalistic reading times and information-theoretic surprisal, surprisal theory This paper re-evaluates a claim due to Goodkind and Bicknell ( By extending Goodkind and Bicknell's analysis to modern neural architectures, we show that the proposed relation does not always hold for Long Short-Term Memory networks, Transformers, and pre-trained models. We introduce an alternate measure of language modeling performance called predictability norm correlation based on Cloze probabilities measured from human subjects. Our new metric yields a more robust relationship between language model quality and psycholinguistic modeling performance that allows for comparison between models with different training configurations.
No takes yet. Share an insight, caveat, or question.
Hao et al. (2020) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: