This study tested how reliably four observers (2 residents, 2 shoulder fellows) could classify the same 20 proximal humeral fractures using the Neer system. Observers reviewed plain radiographs and CT scans on two separate occasions, and kappa statistics measured both self-consistency and agreement between observers. The central question: does adding CT improve Neer classification reliability, and is disagreement caused by displacement thresholds or by fragment identification?
The Neer classification is ubiquitous — every proximal humerus fracture gets labeled 2-part, 3-part, or 4-part. But this paper reveals that the label is unreliable even among experts. With plain radiographs, two fellowship-trained shoulder surgeons will disagree on the Neer classification roughly one-third of the time. Ordering a CT scan will not fix this.
The practical takeaway: when you use Neer classification to guide a treatment decision, recognize that the classification itself carries substantial uncertainty. The good news is that treatment decisions (operative vs. Non-operative) are more reproducible than the underlying classification, suggesting experienced surgeons integrate information beyond the fracture label.
This paper is foundational context for understanding why newer classification efforts (such as AO/OTA) and 3D CT reconstructions have been pursued, and why some advocate for intraoperative classification as the only reliable standard.
This study tested how reliably four observers (2 residents, 2 shoulder fellows) could classify the same 20 proximal humeral fractures using the Neer system. Observers reviewed plain radiographs and CT scans on two separate occasions, and kappa statistics measured both self-consistency and agreement between observers. The central question: does adding CT improve Neer classification reliability, and is disagreement caused by displacement thresholds or by fragment identification?
The Neer classification is ubiquitous — every proximal humerus fracture gets labeled 2-part, 3-part, or 4-part. But this paper reveals that the label is unreliable even among experts. With plain radiographs, two fellowship-trained shoulder surgeons will disagree on the Neer classification roughly one-third of the time. Ordering a CT scan will not fix this.
The practical takeaway: when you use Neer classification to guide a treatment decision, recognize that the classification itself carries substantial uncertainty. The good news is that treatment decisions (operative vs. Non-operative) are more reproducible than the underlying classification, suggesting experienced surgeons integrate information beyond the fracture label.
This paper is foundational context for understanding why newer classification efforts (such as AO/OTA) and 3D CT reconstructions have been pursued, and why some advocate for intraoperative classification as the only reliable standard.