Randomized trial explores urban discourse assessment in Edinburgh and London, highlighting insights for social sciences.
The availability of social media data and tools to (semi-)automatically analyse them has sparked research in various fields to assess human behaviour, opinions and attitudes on a global scale. Besides challenges with data preprocessing to ensure the quality of the database, improving the accuracy of automated classifications, such as topic modelling and sentiment analysis, is frequently discussed in natural language processing (NLP) and beyond. In discourse studies, the interpretability of automated classifications is important for integrating such tools into more fine-grained textual analyses. For instance, the distinction into positive and negative polarity rendered by sentiment analysis constitutes a starting point for assessing human perceptions and opinions, but the mere polarity labels do not necessarily allow for interpreting the meaning of sentiments across topics. To explore this, we use Twitter data from Edinburgh and London posted during the COVID-19 pandemic and assess emerging topics and their evaluation. Focussing on methodological aspects, we integrate corpus-linguistic methods to validate and complement NLP techniques to study urban discourses. This study thus contributes to enhancing the toolkit for corpus-assisted discourse analysis, and, at the same time, underscores the value of linguistic insights for social sciences research.
No takes yet. Share an insight, caveat, or question.
Lehnen et al. (2026) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: