Residency · Residency · Plastic Surgery
Evidence-Based Practice and Critical Appraisal in Plastic Surgery
Introduction
Evidence-based medicine (EBM) integrates the best available research evidence with clinical expertise and patient values to guide clinical decision-making. Defined by Sackett (1996) as "the conscientious, explicit, and judicious use of current best evidence in making decisions about the care of individual patients". Plastic surgery has historically relied heavily on expert opinion and case series (lower levels of evidence); there is an increasing drive toward higher-level evidence. Critical appraisal skills allow surgeons to evaluate the validity, significance, and applicability of published research to their own practice.
Understanding study design, statistical analysis, bias, and outcome measurement is essential for every plastic surgery resident and practitioner.
Levels of Evidence
Oxford Centre for Evidence-Based Medicine Hierarchy
| Level | Study Type | Examples in Plastic Surgery |
|---|---|---|
| I | Systematic reviews of RCTs; individual RCTs | MSLT-II (melanoma SLNB), iBRA study |
| II | Cohort studies; outcomes research | NSQIP analyses, DIEP vs. implant outcomes |
| III | Case-control studies | Risk factor analyses for complications |
| IV | Case series | Single-surgeon technique reports |
| V | Expert opinion; bench research | Consensus guidelines, cadaveric studies |
Level I: systematic reviews of randomized controlled trials (RCTs); individual RCTs with narrow confidence intervals. Level II: systematic reviews of cohort studies; individual cohort studies; outcomes research; ecological studies. Level III: systematic reviews of case-control studies; individual case-control studies. Level IV: case series; poor-quality cohort or case-control studies. Level V: expert opinion; bench research; physiological studies.
Current State of Plastic Surgery Literature
The majority of plastic surgery publications remain Level IV (case series) and Level V (expert opinion). Only 3-5% of plastic surgery literature represents Level I evidence. Barriers to RCTs in plastic surgery include ethical concerns (withholding potentially beneficial procedures), difficulty with blinding, surgical technique variability, and small patient populations for rare conditions. Pragmatic trials and well-designed observational studies can provide valuable evidence when RCTs are not feasible.
Study Design
Experimental Studies
Randomized controlled trial (RCT): gold standard for evaluating interventions; participants randomly assigned to intervention or control groups. Parallel-group: most common; separate groups receive different treatments simultaneously. Crossover: participants serve as their own controls; receive both treatments in sequence; reduces intersubject variability. Cluster randomization: groups (clinics, surgeons) rather than individuals are randomized; used when individual randomization is impractical.
Non-inferiority trial: designed to show a new treatment is not worse than the standard by more than a predefined margin; useful for establishing alternatives with fewer side effects.
Observational Studies
Cohort study: follows groups with and without an exposure over time; can be prospective or retrospective; assesses relative risk. Case-control study: compares patients with an outcome (cases) to those without (controls) and looks back for exposure differences; assesses odds ratio; efficient for rare outcomes. Cross-sectional study: assesses exposure and outcome at a single time point; cannot determine causality; useful for prevalence estimation. Case series/case report: descriptive; no control group; useful for documenting rare conditions, novel techniques, and generating hypotheses.
Systematic Reviews and Meta-Analyses
Systematic review: comprehensive, reproducible search and appraisal of all relevant studies on a specific question; follows PRISMA guidelines. Meta-analysis: statistical pooling of data from multiple studies to increase power and precision; generates a forest plot with summary effect estimate. Require assessment of heterogeneity (I² statistic: >50% indicates significant heterogeneity) and publication bias (funnel plot asymmetry). Quality depends on the quality of included studies: "garbage in, garbage out".
<image>Pyramid diagram showing the hierarchy of evidence from bottom to top: expert opinion and case reports, case series, case-control studies, cohort studies, randomized controlled trials, and systematic reviews/meta-analyses at the apex, with labels indicating the corresponding Oxford levels of evidence (V through I) and the proportion of plastic surgery literature at each level</image>
Critical Appraisal Framework
Validity (Internal Validity)
Selection bias: were groups comparable at baseline? Was randomization adequate? Was allocation concealment used? Performance bias: were participants and clinicians blinded to the intervention? (Difficult in surgery but possible for outcome assessors). Attrition bias: was follow-up complete? Was there differential dropout between groups? Intention-to-treat (ITT) analysis preserves randomization. Detection bias: were outcome assessors blinded? Were outcomes measured consistently across groups?
Reporting bias: were all prespecified outcomes reported? Check trial registration (ClinicalTrials.gov) for selective outcome reporting.
Significance
Statistical significance: p-value <0.05 by convention; indicates the probability of observing the result (or more extreme) if the null hypothesis is true. Clinical significance: the magnitude of the effect and its relevance to patient care; a statistically significant difference may not be clinically meaningful. Confidence intervals: provide a range of plausible values for the true effect; narrow CIs indicate greater precision; if the CI crosses the null value (0 for differences, 1 for ratios), the result is not statistically significant. Effect size measures: mean difference, risk ratio (RR), odds ratio (OR), number needed to treat (NNT).
Applicability (External Validity)
Can the results be applied to my patient population? Consider differences in demographics, comorbidities, surgical technique, and healthcare setting. Were the outcomes measured relevant to my patients? Patient-reported outcomes (PROs) are increasingly valued over surgeon-assessed outcomes. Is the intervention feasible in my practice setting? Consider resources, expertise, and cost.
Common Statistical Errors
Type I error (alpha): false positive; concluding an effect exists when it does not; typically set at 0.05. Type II error (beta): false negative; failing to detect a true effect; related to statistical power (1-beta); adequately powered studies require sufficient sample size. Multiple comparisons: testing multiple outcomes inflates Type I error; Bonferroni correction or similar adjustments should be applied. Confounding: a third variable associated with both exposure and outcome creates a spurious association; control through randomization, matching, or multivariate analysis.
Correlation vs. causation: observational studies demonstrate associations but cannot prove causality; Hill's criteria provide a framework for causal inference.
Outcome Measurement in Plastic Surgery
Patient-Reported Outcome Measures (PROMs)
PROMs capture the patient's perspective on their health, function, and satisfaction; increasingly recognized as essential endpoints. Validated instruments specific to plastic surgery: BREAST-Q: domains for satisfaction with breasts, psychosocial well-being, physical well-being, and satisfaction with care; validated for augmentation, reduction, reconstruction, and mastopexy. FACE-Q: facial aesthetic and reconstructive outcomes.
HAND-Q: hand surgery outcomes. BODY-Q: body contouring outcomes. CLEFT-Q: cleft lip and palate outcomes. PROMs must be validated (content validity, construct validity, responsiveness) and reliable (test-retest, internal consistency).
Clinician-Reported Outcomes
Surgeon-assessed outcomes (e.g., complication rates, flap survival, aesthetic grading) remain important but are subject to observer bias. Standardized photography with blinded panel assessment improves objectivity. Complication grading systems: Clavien-Dindo classification adapted for plastic surgery provides a standardized framework.
Composite Endpoints
Combining multiple outcomes into a single endpoint increases statistical power but can be misleading if components have different clinical importance. Each component should be reported individually alongside the composite.
<image>Illustration showing the three pillars of evidence-based medicine as a Venn diagram: best available research evidence (systematic reviews, RCTs, cohort studies), clinical expertise (surgical skill, pattern recognition, experience), and patient values and preferences (goals, concerns, quality of life priorities), with evidence-based clinical decisions at the intersection of all three</image>
Landmark Trials in Plastic Surgery
Selected High-Impact Studies
MSLT-I and MSLT-II: sentinel lymph node biopsy in melanoma; MSLT-II showed no survival benefit for completion lymph node dissection after positive SLNB. iBRA study: largest prospective study of immediate breast reconstruction outcomes in the UK; identified risk factors for complications. POISE (Perioperative Ischemic Evaluation): demonstrated increased stroke risk with perioperative beta-blocker initiation (applicable to all surgical specialties). Ilizarov principles: demonstrated the biology of distraction osteogenesis; transformed limb reconstruction. Anderson and Parrish (1983): selective photothermolysis; foundational principle for all laser therapy.
Practice-Changing Evidence
Progressive tension sutures (Pollock): reduced seroma rates in abdominoplasty from 15% to <1%. Nipple-sparing mastectomy: oncologic safety data supporting preservation of the nipple-areolar complex in selected breast cancer patients. Negative pressure wound therapy (VAC): evidence supporting its use in open wounds, skin grafts, and surgical incisions. Enhanced Recovery After Surgery (ERAS) protocols: multimodal perioperative pathways reducing length of stay and complications in breast reconstruction and body contouring.
Implementing Evidence-Based Practice
Barriers in Plastic Surgery
Surgical variability: each surgeon's technique differs; standardizing the "intervention" is challenging. Learning curve effects: outcomes improve with surgeon experience; early results may not reflect the true efficacy of a technique. Small sample sizes: many plastic surgery conditions are rare, limiting statistical power. Industry influence: device and implant manufacturers fund research; potential for bias in study design and reporting.
Publication bias: positive results are more likely to be published; negative or inconclusive studies are underrepresented.
Strategies for Improvement
Multicenter registries: pooling data across institutions increases sample sizes and generalizability (e.g., TOPS database, BRA-Day studies). Standardized outcome reporting: using validated PROMs and agreed-upon complication definitions facilitates comparison across studies. Journal club: regular critical appraisal of current literature develops appraisal skills and updates clinical practice. Clinical practice guidelines: synthesize evidence into actionable recommendations (e.g., ASPS Evidence-Based Clinical Practice Guidelines).
Shared decision-making tools: decision aids that present evidence-based information to patients facilitate informed consent and align treatment with patient values.
<image>Table comparing the critical appraisal checklist items for different study designs: RCTs (randomization, blinding, ITT analysis, follow-up), cohort studies (exposure measurement, confounding control, follow-up duration), case-control studies (case definition, control selection, recall bias), and systematic reviews (search strategy, quality assessment, heterogeneity, publication bias)</image>
Key Clinical Pearls
Evidence-based practice integrates research evidence with clinical expertise and patient values; it does not mean blindly following RCTs or ignoring clinical judgment. Most plastic surgery evidence is Level IV (case series); there is an urgent need for well-designed prospective studies and RCTs, but well-conducted observational studies also provide valuable evidence. When appraising a study, systematically assess validity (bias), significance (statistical and clinical), and applicability (to your patient population). Patient-reported outcome measures (BREAST-Q, FACE-Q) capture what matters most to patients and should be included as primary or co-primary endpoints in clinical research. Understanding common statistical pitfalls (Type I/II errors, confounding, multiple comparisons) prevents misinterpretation of results and guides appropriate clinical application.
References
- Sackett DL, Rosenberg WMC, Gray JAM, Haynes RB, Richardson WS. Evidence based medicine: what it is and what it isn't. BMJ. 1996;312(7023):71-72.
- Pusic AL, Klassen AF, Scott AM, Klok JA, Cordeiro PG, Cano SJ. Development of a new patient-reported outcome measure for breast surgery: the BREAST-Q. Plast Reconstr Surg. 2009;124(2):345-353.
- Burns PB, Rohrich RJ, Chung KC. The levels of evidence and their role in evidence-based medicine. Plast Reconstr Surg. 2011;128(1):305-310.
- Defined by Oxford Centre for Evidence-Based Medicine. OCEBM Levels of Evidence Working Group. Oxford Centre for Evidence-Based Medicine; 2011. Available at: https://www.cebm.ox.ac.uk/resources/levels-of-evidence


