Many texts are broadly disseminated on online social media platforms each day. Topic modeling is a Natural Language Processing and Unsupervised Learning technique used in this scenario. It identifies the topics in a collection of texts — that is, the most relevant groups of words in the context of all the texts are analyzed. The characteristics of the texts are a relevant factor for identifying topics. Unlike traditional sources, texts published on social media are usually short, because even when a character limit per publication (e.g., Twitter/X) is not imposed, users tend to be objective in the texts they write. This work evaluates the performance of four categories of topic modeling algorithms: traditional, Dirichlet Multinomial Mixture (DMM)-based, self-aggregation-based, and global co-occurrence-based. Real texts generated by social media users were used for the evaluation. Model performance was evaluated using quality metrics accepted in the literature. Finally, the results were analyzed such that the performance of each algorithm was weighted down, and to clarify whether there would be any detriment to the results of topic modeling using traditional algorithms on short texts.
Santos et al. (Fri,) studied this question.