Twelve observers rated the stuttering severity of 400 different samples during four sessions to see if the experience of listening to and rating many samples affected parameters of observer agreement. Indices of agreement showed no change within or over sessions; overall level of rating and patterns of errors for observers remained moderately stable within and over sessions. Simply rating many samples should not be construed as training. The data from this study can be a baseline for comparing conditions that may increase observer agreement.
No takes yet. Share an insight, caveat, or question.
Martin A. Young (1969) studied this question.