How can the contributions of raters and tasks to error variance be estimated? Which source of error variance is usually greater? Are interrater coefficients adequate estimates of reliability? What other facets contribute to unreliability in performance assessments?
No takes yet. Share an insight, caveat, or question.
Brennan et al. (1995) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: