When constructed response items are used on more than one occasion, a natural concern is whether the scoring is consistent (e.g. not more lenient or strict) across the occasions. It is common to conduct trend scoring in which a set of occasion A responses are re-scored at occasion B. The responses are usually selected according to some rescore design, such as being balanced (with an equal number from each score category), proportional to the distribution of occasion A scores, or a mixed version of these two designs. Recent work has demonstrated that treating the two-way table as if it arose from multinomial sampling is incorrect, and can yield seriously biased estimates of whether the scores are lower/higher at occasion B. The present study builds on these results by incorporating ordinal measures of change. It contrasts the usual trend analysis with an alternative analysis that explicitly conditions on the rescore design and finds only the latter is effective. Omnibus measures based on combining the individual t-tests/d-statistics are examined. Measures were somewhat conservative in Type I error control and had good power to detect drift. Omnibus measures based on t-tests had marginally higher power, having higher correct detection rates than those based on the d-statistic in 1-8% of the cases. The difference between the best versions (E weighted, which is based on t-tests, v. D weighted, which is based on d-statistics) was only 1.8%. Keywords: constructed response
Donoghue et al. (Mon,) studied this question.