This 2009 JBJS symposium paper by Schemitsch and colleagues audits the quality of orthopedic evidence. It asks whether RCTs, observational studies, and expert opinion each adequately support clinical decision-making. The answer: orthopedic literature is dominated by low-level evidence, and even its RCTs are frequently underpowered and poorly reported.
When you read an orthopedic RCT that reports "no significant difference," ask whether it was powered to detect one. With 91% of fracture care RCTs underpowered, a negative result almost never constitutes proof of equivalence.
When a multi-outcome study reports a statistically significant finding, count how many outcomes were tested. Testing 16 outcomes gives a 55% chance of a spurious p < 0.05 by chance alone — the more outcomes reported, the less each individual positive result should be trusted without correction.
When a new implant or technique generates exciting single-center retrospective data, treat it as hypothesis-generating only. The unreamed femoral nail became standard practice on retrospective data, then failed in a powered RCT with a 4.5-fold higher nonunion rate. The pattern repeats.
Level of evidence is not the same as study quality. Level-I and Level-II RCTs in this review did not differ significantly in Cochrane quality scores (15.2 vs 11.7 points, p = 0.08). A Level-I label warrants scrutiny, not automatic trust.
This 2009 JBJS symposium paper by Schemitsch and colleagues audits the quality of orthopedic evidence. It asks whether RCTs, observational studies, and expert opinion each adequately support clinical decision-making. The answer: orthopedic literature is dominated by low-level evidence, and even its RCTs are frequently underpowered and poorly reported.
When you read an orthopedic RCT that reports "no significant difference," ask whether it was powered to detect one. With 91% of fracture care RCTs underpowered, a negative result almost never constitutes proof of equivalence.
When a multi-outcome study reports a statistically significant finding, count how many outcomes were tested. Testing 16 outcomes gives a 55% chance of a spurious p < 0.05 by chance alone — the more outcomes reported, the less each individual positive result should be trusted without correction.
When a new implant or technique generates exciting single-center retrospective data, treat it as hypothesis-generating only. The unreamed femoral nail became standard practice on retrospective data, then failed in a powered RCT with a 4.5-fold higher nonunion rate. The pattern repeats.
Level of evidence is not the same as study quality. Level-I and Level-II RCTs in this review did not differ significantly in Cochrane quality scores (15.2 vs 11.7 points, p = 0.08). A Level-I label warrants scrutiny, not automatic trust.