This study tested how reliably five experienced shoulder surgeons could classify the same 95 proximal humerus fractures using both the Neer and AO/ASIF systems. Each surgeon classified all fractures twice, 8 weeks apart, to measure both inter- and intraobserver agreement. The study also asked whether a three-view extended trauma series improved classification reliability over standard two-view imaging.
When you use the Neer or AO/ASIF classification to guide an operative decision, recognize that the classification itself carries real uncertainty. Two experienced shoulder surgeons will disagree on the fracture pattern in roughly half of cases — and the same surgeon will reclassify the same fracture differently one time in three.
This has direct clinical implications. Treatment algorithms that rely on distinguishing a two-part from a three-part fracture, or a Neer Group III from Group IV, are built on a substrate of poor reproducibility. This paper is part of why we obtain CT for complex proximal humerus fractures before committing to a classification-driven operative plan.
For multicenter research, this study established that comparing outcomes of "similarly classified" fractures across institutions is not scientifically valid. Each center's classification may systematically differ from another's.
The key pearl: the disagreement in the literature about AVN rates and treatment outcomes for displaced proximal humerus fractures may be largely a classification artifact, not a true difference in biology or surgical results.
This study tested how reliably five experienced shoulder surgeons could classify the same 95 proximal humerus fractures using both the Neer and AO/ASIF systems. Each surgeon classified all fractures twice, 8 weeks apart, to measure both inter- and intraobserver agreement. The study also asked whether a three-view extended trauma series improved classification reliability over standard two-view imaging.
When you use the Neer or AO/ASIF classification to guide an operative decision, recognize that the classification itself carries real uncertainty. Two experienced shoulder surgeons will disagree on the fracture pattern in roughly half of cases — and the same surgeon will reclassify the same fracture differently one time in three.
This has direct clinical implications. Treatment algorithms that rely on distinguishing a two-part from a three-part fracture, or a Neer Group III from Group IV, are built on a substrate of poor reproducibility. This paper is part of why we obtain CT for complex proximal humerus fractures before committing to a classification-driven operative plan.
For multicenter research, this study established that comparing outcomes of "similarly classified" fractures across institutions is not scientifically valid. Each center's classification may systematically differ from another's.
The key pearl: the disagreement in the literature about AVN rates and treatment outcomes for displaced proximal humerus fractures may be largely a classification artifact, not a true difference in biology or surgical results.