PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
November 1, 1969Psychological Bulletin1,535 citations

Large sample standard errors of kappa and weighted kappa.

View Full Paper
JFJoseph L. FleissJCJacob CohenBEB. S. Everitt

Key Points

  • To identify mathematical inconsistencies in established standard error formulas for Cohen's kappa and weighted kappa and formulate valid large-sample variance expressions.
  • Analyzed prior standard error derivations by Cohen and Everitt based on fixed marginal totals and binomial distributions.
  • Modeled the cross-classification of N subjects assigned independently into k nominal categories by two raters.
  • Demonstrated that traditional standard error formulas for kappa and weighted kappa are mathematically invalid due to conflicting assumptions of fixed marginal totals and binomial cell variation.
  • Showed that simplified binomial approximations previously suggested for null-parameter variances are similarly flawed and unsuitable for routine application.

Abstract

The statistics kappa (Cohen, 1960) and weighted kappa (Cohen, 1968) were introduced to provide coefficients of agreement between two raters for nominal scales. Kappa is appropriate when all disagreements may be considered equally serious, and weighted kappa is appropriate when the relative seriousness of the different possible disagreements can be specified. The papers describing these two statistics also present expressions for their standard errors. These expressions are incorrect, having been derived from the contradictory assumptions of fixed marginal totals and binomial variation of cell frequencies. Everitt (1968) derived the exact variances of weighted and unweighted kappa when the parameters are zero by assuming a generalized hypergeometric distribution. He found these expressions to be far too complicated for routine use, and offered, as alternatives, expressions derived by assuming binomial distributions. These alternative expressions are incorrect, essentially for the same reason as above. Assume that N subjects are distributed into k* cells by each of them being assigned to one of k categories by one rater and, independently, to one of the same k categories by a second

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Fleiss et al. (1969) studied this question.

synapsesocial.com/papers/69d5760e734fad6b67f4c922https://doi.org/10.1037/h0028106
Ask AI
Helpful
Bookmark
Share
View Full Paper