Topic modeling, in particular the Latent Dirichlet Allocation (LDA) model, has recently emerged as an important tool for understanding large datasets, in particular, user-generated datasets in social studies of the Web. In this work, we investigate the instability of LDA inference, propose a new metric of similarity between topics and a criterion of vocabulary reduction. We show the limitations of the LDA approach for the purposes of qualitative analysis in social science and sketch some ways for improvement.
No takes yet. Share an insight, caveat, or question.
Koltcov et al. (2014) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: