Reliability as measured by the extent of agreement is often a problem for complex global judgments. Empirically, the use of multiple raters improved reliability consistent with predictions from the Spearman-Brown formula. Implications for the reliability of clinical diagnosis are suggested.
No takes yet. Share an insight, caveat, or question.
James D. Roff (1981) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: