Item rescoring, or the post hoc collapsing of Likert-type response categories, is an important yet underexplored psychometric practice to address category malfunctioning issues such as disordered thresholds or underused response options. These issues could lead to inaccurate interpretation of scale scores and flawed conclusions in research and clinical settings. Despite its frequent use, limited empirical guidance exists on whether rescoring is psychometrically justifiable and how it should be implemented. This dissertation comprises two studies designed to address these gaps. Study 1 was a scoping review examining how rescoring has been applied in existent literature, with a focus on depression measures. By synthesizing findings from 13 studies, this review offered four key recommendations: (1) researchers should aim to prevent response scale malfunctioning in the first place; (2) rescoring should be applied thoughtfully, when malfunctioning issues arise; (3) uniform rescoring may serve as a starting point for practicality; and (4) rescoring decisions should be clearly justified in reporting. This review also identified critical gaps, including the need for clearer criteria to define low responses, greater clarity on the use of individual rescoring, and more systematic evaluations of the psychometric consequences of rescoring. Building on one of these gaps, Study 2 systematically evaluated the impact of four uniform rescoring methods (tail combination, middle combination, presence scoring, and persistence scoring) on the Patient Health Questionnaire-9’s psychometric properties, including internal consistency reliability, test-retest reliability, factor structure, and convergent and discriminant validity. Results suggested that rescoring generally had minimal impact on psychometric quality, although middle combination rescoring corrected threshold disordering most effectively. Extending the investigation longitudinally, Study 2 also found that, while threshold disordering and low responses persisted across time points, the specific items affected varied. Nonetheless, the rescoring methods that resolved these issues at baseline remained consistently effective at subsequent waves, suggesting that the rescoring needs associated with each type of data issue are not necessarily different over time. Together, the two studies provide valuable insights into the implementation of rescoring in both cross-sectional and longitudinal contexts and offer converging evidence that rescoring is a psychometrically defensible practice that does not compromise measures’ psychometric quality.
Xuyan Tang (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: