Any orthopaedist can tell the difference between an apple and an orange. Most people can tell the difference between an orange and a tangerine, but when comparing tangelos and mandarins, things become a bit more difficult. Alas, the latter appears to be the case with fracture classifications, their purpose, and their continued usefulness in orthopaedic trauma surgery. Clinical publications dealing with fracture management generally attempt to answer the question: What is the best expected prognosis for a given fracture pattern depending on what treatment is used? Authors advocate a treatment method based on a specific fracture pattern, hailing its benefits. But when that treatment fails, these same authors are unsure whether this was due to the fracture pattern or the treatment method. As a result, newly improved, and generally more complex, classifications are created. Therein lies the rub. As Swiontkowski et al. have correctly pointed out in this issue, while the variation of fracture patterns for an individual bone sits on a continuum, the patterns we devise are nothing more than arbitrary compartments. The end result is uncertainty. In an effort to determine when observers agree, statisticians have developed a measure of inter- and intraobserver reliability, known as the kappa coefficient (5). Interobserver reliability implies that two different observers looking at the same item will see the same thing. Intraobserver reliability implies that the same observer, looking at the same item at different times, will see the same thing. One would expect that this would be straightforward, but unfortunately it is not true. To highlight this point, in a study regarding the Garden classification of femoral neck fractures, the best correlation between observers Frandsen et al. (2) could obtain was 0.47. In evaluating the Neer classification of proximal humerus fractures, interobserver reliability ranged from 0.40 to 0.53 in one study, 0.35 to 0.45 in another study, and 0.48 to 0.52 in a third (4,6,7). In one of the three articles dealing with pilon fractures in this issue, Swiontkowski et al. were able to achieve an average reliability of 0.57 when evaluating fracture types (AO/OTA 43A, B, and C). This dropped to 0.43 for fracture group and 0.41 for fracture subgroup. Similarly, Martin et al., in evaluating AO/ASIF 43 type fractures had reliability as high as 0.60 for fracture types, but this dropped as low as 0.38 when fracture groups were evaluated. The message is clear. To make a fracture classification work, one must keep it simple (2-4,6-8). For the readership to believe that a treatment modality works, they must know that the fracture this treatment is based upon is a reproducible pattern that all orthopaedists can reliably recognize. The editors of the Journal of Orthopaedic Trauma were aware of these difficulties when they adopted the AO/OTA fracture classification (1) as the standard for this publication. Our clinical studies will have real meaning only by adhering to this existing scheme and continually improving it. It is the editorial policy of this journal to require all manuscripts submitted to conform to these standards. What the authors of the articles in this issue have shown us is that drawing conclusions regarding fracture management based on patterns more stratified than the broad fracture types of the AO/OTA classification (A, B, and C) is, at this point in time, meaningless. Any data regarding treatment based on fracture groups or subgroups should first be validated with an interobserver reliability of greater than 0.55. Furthermore, any new classification that is to be presented for publication should be tested for inter- and intraobserver reliability before it is put into clinical practice. The fourth article in this issue, by Chan et al., attempts to do just that. The authors of all four articles in this series should be commended for questioning the very foundation of our treatment assumptions: namely, that apples are not oranges, unless they happen to look like tangelos,... or was it mandarins? Roy W. Sanders, M.D. Editor-in-Chief Tampa, Florida
No takes yet. Share an insight, caveat, or question.
Roy Sanders (1997) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: