Kellgren and Lawrence graded 510 radiographs from 85 adults (ages 55-64) across 11 joint groups to create a standardized 5-grade scale for radiographic OA severity. The study simultaneously quantified how much reader identity alone distorts OA prevalence estimates in population research.
The question 'does this patient have OA?' has a specific radiographic answer tracing directly to this 1957 paper: Grade 2 means definite OA is present. Every major OA cohort since — Framingham, the Osteoarthritis Initiative, ACL and meniscus outcome trials. Uses this same threshold to define disease presence and enroll patients.
When reviewing a knee or hip film, Grade 1 (doubtful) should not drive replacement or surgical decisions. The paper documents that Grade 1 reproducibility is poor even between trained observers.
The ±31% inter-observer error has direct implications for imaging reporting: if two surgeons grade the same films independently, their prevalence estimates can differ by nearly a third. This is why multicenter imaging studies require centralized readers or consensus panels.
For the boards: the wrist is the one joint where Kellgren-Lawrence grading breaks down. Rheumatoid overlap confounds OA grading there, and this paper provides the original data to prove it.
Kellgren and Lawrence graded 510 radiographs from 85 adults (ages 55-64) across 11 joint groups to create a standardized 5-grade scale for radiographic OA severity. The study simultaneously quantified how much reader identity alone distorts OA prevalence estimates in population research.
The question 'does this patient have OA?' has a specific radiographic answer tracing directly to this 1957 paper: Grade 2 means definite OA is present. Every major OA cohort since — Framingham, the Osteoarthritis Initiative, ACL and meniscus outcome trials. Uses this same threshold to define disease presence and enroll patients.
When reviewing a knee or hip film, Grade 1 (doubtful) should not drive replacement or surgical decisions. The paper documents that Grade 1 reproducibility is poor even between trained observers.
The ±31% inter-observer error has direct implications for imaging reporting: if two surgeons grade the same films independently, their prevalence estimates can differ by nearly a third. This is why multicenter imaging studies require centralized readers or consensus panels.
For the boards: the wrist is the one joint where Kellgren-Lawrence grading breaks down. Rheumatoid overlap confounds OA grading there, and this paper provides the original data to prove it.