The use of automated coding procedures to scale up content analysis has risen over the last years. Using a mixed-method approach, we examine researchers’ justifications to scale up content analysis and assess the methodological adjustments—or lack thereof—when employing supervised machine learning for 38 large-scale content analyses. Almost all of the included studies displayed deficiencies in study design, primarily related to the uncritical use of frequentist statistics on datasets containing the entire statistical population, or employing supervised machine learning without methodological adjustments to account for misclassifications. Our findings question the need for large datasets and automated coding in the first place.
Linde et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: