Residency · Residency · Orthopedic Surgery
Evidence-Based Orthopedics: Reading and Appraising the Literature
Introduction
Evidence-based medicine (EBM) is the conscientious, explicit, and judicious use of current best evidence in making decisions about the care of individual patients. In orthopedic surgery, where surgical techniques and implant technologies evolve rapidly, the ability to critically appraise the literature is essential for delivering optimal patient care. This lecture provides a structured framework for reading, interpreting, and applying orthopedic research to clinical practice.
Hierarchy of Evidence
Levels of Evidence in Orthopedic Research
| Level | Study Design | Example |
|---|---|---|
| I | High-quality RCT; systematic review of Level I studies | DRAFFT2 trial; Cochrane reviews |
| II | Lesser-quality RCT; prospective comparative study | RCT with >20% loss to follow-up |
| III | Case-control; retrospective comparative study | Registry-based comparison |
| IV | Case series (no comparison group) | Single-center surgical outcomes report |
| V | Expert opinion; case reports; narrative reviews | Technique descriptions; consensus statements |
The levels of evidence form a hierarchy that guides the weight given to different study designs. Level I represents high-quality randomized controlled trials (RCTs) or systematic reviews of Level I studies. Level II includes lesser-quality RCTs and prospective comparative studies. Level III encompasses case-control studies and retrospective comparative studies. Level IV consists of case series without a comparison group. Level V includes expert opinion, case reports, and narrative reviews.
The Role of Systematic Reviews and Meta-Analyses
Cochrane Reviews represent the gold standard for systematic evidence synthesis. Meta-analyses pool data from multiple studies to increase statistical power, but heterogeneity between studies (measured by the I-squared statistic) must be assessed before pooling results. Forest plots visually display individual study effects and the pooled estimate.
Critical Appraisal of Clinical Studies
Evaluating Randomized Controlled Trials
When appraising an RCT, one must assess the randomization method and concealment of allocation, determine whether blinding was feasible (double-blind designs are rare in surgical trials), calculate the number needed to treat (NNT) to understand clinical significance, evaluate whether intention-to-treat versus per-protocol analysis was used, and review loss to follow-up rates, since greater than 20% introduces significant bias.
Understanding Bias in Orthopedic Research
Selection bias arises from non-random allocation of patients to treatment groups. Performance bias results from differences in care beyond the intervention being studied. Detection bias occurs with non-blinded outcome assessment, particularly relevant for subjective outcomes such as pain scores. Attrition bias develops from differential loss to follow-up between groups. Publication bias reflects the tendency for positive results to be published more readily than negative findings.
Assessing Observational Studies
Cohort studies follow patients over time but lack randomization, while case-control studies compare patients with an outcome to those without. Confounding variables must be controlled through multivariate analysis or propensity score matching. The Newcastle-Ottawa Scale is commonly used to assess quality of non-randomized studies.
Statistical Concepts for the Orthopedic Surgeon
Measures of Treatment Effect
Relative risk (RR) is the ratio of event rates between groups and is used in cohort studies and RCTs. Odds ratio (OR) is the ratio of odds of an event, used primarily in case-control studies. Hazard ratio (HR) is used in survival analysis to compare time-to-event data. Confidence intervals are essential for interpretation: a 95% CI that crosses 1.0 for RR or OR indicates no statistically significant difference.
P-Values and Clinical Significance
A p-value less than 0.05 is the conventional threshold for statistical significance, but statistical significance does not equal clinical significance. The minimal clinically important difference (MCID) defines the smallest change in an outcome measure that patients perceive as meaningful. Common orthopedic MCID values include 10 points for the DASH score, 2 cm for the VAS pain scale, and 18 points for the Harris Hip Score.
Power Analysis and Sample Size
Type I error (alpha) is the probability of a false positive, typically set at 0.05. Type II error (beta) is the probability of a false negative, typically set at 0.20. Power (1 minus beta) represents the ability to detect a true difference, with 80% power being the standard threshold. Underpowered studies in orthopedics are common and may miss true treatment effects.
Applying Evidence to Orthopedic Practice
Patient-Reported Outcome Measures (PROMs)
PROMs capture the patient perspective on treatment effectiveness. Common validated instruments include the WOMAC, KOOS, DASH, and SF-36. Ceiling and floor effects must be considered when selecting outcome tools, and responsiveness to change is a critical psychometric property.
Shared Decision-Making
Integrating the best available evidence with patient values and preferences is the foundation of shared decision-making. Decision aids can improve patient understanding of risks and benefits, while the surgeon's clinical expertise remains an essential component of EBM.
Common Pitfalls in Orthopedic Literature
Surrogate endpoints (such as radiographic alignment) may not correlate with patient-centered outcomes. Industry-sponsored studies may overestimate treatment effects. Short follow-up periods may miss late complications such as implant failure or adjacent segment disease. Subgroup analyses without pre-specification increase the risk of false-positive findings. Lack of standardized reporting remains a problem, though the CONSORT statement for RCTs and STROBE for observational studies provide essential checklists.
Key Clinical Pearls
One should always assess the level of evidence before changing clinical practice based on a single study. Statistical significance (p < 0.05) without exceeding the MCID should not drive treatment decisions. Underpowered studies reporting "no difference" may simply lack the sample size to detect a real effect, making it essential to check the power analysis. Validated appraisal tools such as the CONSORT checklist for RCTs and the Newcastle-Ottawa Scale for observational studies should be used to systematically evaluate study quality.
References
- Bhandari M, Tornetta P. Evidence-Based Orthopedics. Clinical Orthopaedics and Related Research. 2003;413:19-27.
- Slobogean GP, Leung R, Englesbe MJ. Evidence-based practice in surgery: the hierarchy of evidence. Journal of Bone and Joint Surgery. 2015;97(18):e65.
- Wright JG, Swiontkowski MF, Heckman JD. Introducing levels of evidence to the journal. Journal of Bone and Joint Surgery (American). 2003;85(1):1-3.
- Poolman RW, Struijs PA, Krips R, et al. Reporting of outcomes in orthopaedic randomized trials: does blinding of outcome assessors matter? Journal of Bone and Joint Surgery (American). 2007;89(3):550-558.