PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 27, 2026ACM Transactions on Intelligent Systems and Technology0 citations

When Do Large Language Models (LLMs) Struggle to Count Letters?

View Full Paper
TFTairan FuRFRaquel FerrandoJCJavier Conde

Key Points

  • This research aims to investigate when large language models fail to accurately count letters in words.
  • Conducted an experimental analysis on a representative group of large language models.
  • Evaluated the relationship between model errors and word frequency, token complexity, and counting operation intricacy.
  • Analyzed a large dataset of words to assess patterns in counting errors.
  • Models recognize letters but fail to count them accurately.
  • Frequency of words and tokens does not significantly affect counting errors.
  • Stronger counting error correlations observed with words containing letters that appear more than once.

Abstract

Large Language Models (LLMs) have achieved unprecedented performance on many complex tasks, being able, for example, to answer questions on almost any topic. However, they struggle with other simple tasks, such as counting the occurrences of letters in a word, as illustrated by the inability of many LLMs to count the number of ”r” letters in ”strawberry”. Several works have studied this problem and linked it to the tokenization used by LLMs, to the intrinsic limitations of the attention mechanism, or to the lack of character-level training data. In this paper, we conduct an experimental study to evaluate when LLMs fail to count the letters. In more detail, we study the relations between the LLM errors when counting letters including 1) the frequency of the word and its components in the training dataset and 2) the complexity of the counting operation. We present a comprehensive analysis of the errors of LLMs when counting letter occurrences by evaluating a representative group of models over a large number of words. The results show a number of consistent trends in the models evaluated: 1) models are capable of recognizing the letters but not counting them; 2) the frequency of the word and tokens in the word does not have a significant impact on the LLM errors; 3) there is a positive correlation of letter frequency with errors, more frequent letters tend to have more counting errors, 4) the errors show a strong correlation with the number of letters or tokens in a word and 5) the strongest correlation occurs with the number of letters with counts larger than one, with most models being unable to correctly count words in which letters appear more than twice. These results suggest that the problems of LLMs to count letters are not related to the frequency of words or tokens in the training data but to the complexity of the counting operation. However, further studies are needed to build a better understanding of the limitations of LLMs to count the letters in a word.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Fu et al. (2026) studied this question.

synapsesocial.com/papers/6a1689ce0c924ddd1bd587b3https://doi.org/10.1145/3818606
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond2024 · 490 citations
  2. 2Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books2015 · 2,074 citations
  3. 3CUTE: Measuring LLMs’ Understanding of Their Tokens2024 · 5 citations
  4. 4HellaSwag: Can a Machine Really Finish Your Sentence?2019 · 784 citations
  5. 5Untitled7,295 citations