Premed · Premed · Statistics Biostatistics

Lecture 27: Interpreting Medical Literature Critically

Statistics / Biostatistics


Learning Objectives

By the end of this lecture, students will be able to:

  1. Apply a structured framework for critically appraising a clinical study
  2. Evaluate the validity of study design, analysis, and conclusions
  3. Identify common statistical errors and misleading presentations in published research
  4. Interpret forest plots and meta-analyses
  5. Distinguish between association and causation using the Bradford Hill criteria

Lecture Content

I. Why Critical Appraisal Matters

Medical knowledge is constantly evolving, and clinicians must evaluate new evidence as it emerges. Not all published studies are well-designed or correctly analyzed. Evidence-based medicine requires the ability to assess whether findings are valid (internal validity), determine whether findings apply to a specific patient (external validity or generalizability), and weigh the benefits and harms of interventions. Critical appraisal is a core competency for all healthcare professionals.

II. A Framework for Critical Appraisal

Three fundamental questions guide the appraisal of any study. First, are the results valid (internal validity)? Second, what are the results in terms of magnitude and precision? Third, will the results help in caring for patients (applicability)?

Structured tools exist to facilitate this process. The CONSORT checklist is used for RCTs, the STROBE checklist for observational studies, the PRISMA checklist for systematic reviews, and the CASP (Critical Appraisal Skills Programme) worksheets provide general guidance.

III. Evaluating Internal Validity

Evaluating internal validity requires asking several questions. Was the study design appropriate for the question? Therapy questions call for RCTs, prognosis questions for cohort studies, diagnosis questions for cross-sectional studies with gold standard comparison, and etiology or harm questions for cohort or case-control designs.

For RCTs, was randomization performed properly? Was allocation concealed? Were groups similar at baseline? Was blinding adequate, and who was blinded (participants, clinicians, outcome assessors)? Was follow-up complete -- what was the dropout rate, was it similar between groups, and was ITT analysis performed? Were outcomes measured validly and reliably? Were potential confounders identified and controlled?

IV. Evaluating the Results

The results should be evaluated along several dimensions. What is the effect size? Relative measures include RR, OR, and HR, while absolute measures include ARR and NNT. For continuous outcomes, the mean difference or standardized mean difference is relevant. How precise is the estimate? The width of the confidence interval answers this -- narrow CIs indicate precise estimates, wide CIs indicate uncertainty.

Is the result statistically significant, based on the p-value relative to alpha? But statistical significance is not the same as clinical significance. The key question is whether the effect size is large enough to change practice, considered in light of the minimal clinically important difference (MCID).

<image>An annotated excerpt from a hypothetical clinical trial results table. The table shows primary and secondary outcomes with columns for Treatment group (events/n), Control group (events/n), Effect measure (RR or mean difference), 95% CI, and p-value. Annotations around the table highlight key appraisal points: "Is the CI narrow or wide?" "Does the CI exclude 1.0 (or 0)?" "Is the effect size clinically meaningful?" "Was the primary outcome pre-specified?" Color-coded highlights distinguish significant from non-significant results.</image>

V. Common Statistical Errors in the Literature

Several statistical errors frequently appear in published research. Multiple comparisons without correction inflates the false positive rate when many outcomes or subgroups are tested. Subgroup analyses presented as primary findings are particularly problematic because post-hoc subgroup results are hypothesis-generating, not confirmatory. Confusing correlation with causation is common in observational studies, which can only suggest associations. P-hacking involves selectively reporting analyses that produce p < 0.05. Underpowered studies may produce negative results that simply reflect insufficient sample size. Reporting relative risk without absolute risk can exaggerate the perceived benefit. Ignoring intention-to-treat and analyzing only completers introduces attrition bias. Survival bias involves analyzing only those who survived to a certain point. The ecological fallacy draws individual-level conclusions from group-level data.

VI. Systematic Reviews and Meta-Analysis

A systematic review is a structured, reproducible approach to identifying, evaluating, and synthesizing all relevant studies on a topic. It follows a pre-specified protocol (registered on PROSPERO), uses a comprehensive search strategy, and includes quality assessment of all included studies.

A meta-analysis is the statistical pooling of results from multiple studies to provide a more precise estimate of the effect. A fixed-effect model assumes one true effect size and weights studies by their precision. A random-effects model allows for between-study variability (heterogeneity) and is more conservative. The standard graphical display is the forest plot, where each study is shown as a square (with size proportional to its weight) and a CI line, and a diamond at the bottom represents the pooled estimate.

Heterogeneity -- variability in results across studies -- is assessed with the I^2 statistic, which represents the percentage of total variability due to between-study differences. An I^2 of 0% indicates no heterogeneity, 25-50% indicates low to moderate heterogeneity, and I^2 above 75% indicates high heterogeneity, where the pooled estimate may not be meaningful. Cochran's Q test formally tests whether observed differences exceed chance expectation.

Publication bias -- the tendency for studies with positive or significant results to be published more often -- is assessed with a funnel plot (a scatterplot of effect size versus study precision, where asymmetry suggests bias) and Egger's test (a formal statistical test for funnel plot asymmetry).

<image>A two-panel meta-analysis figure. Panel A: A forest plot with 8 studies. Each row shows the study author and year, a square with 95% CI line, and the numerical OR with CI. The pooled estimate (diamond) is at the bottom. Squares vary in size (weight). A vertical line at OR = 1.0 marks no effect. The I-squared statistic (e.g., 35%) is annotated. Panel B: A funnel plot for the same studies. The x-axis shows the log(OR) and the y-axis shows the standard error (inverted). Points are roughly symmetric around the pooled estimate, suggesting no major publication bias. The expected triangular "funnel" outline is drawn.</image>

VII. Association vs. Causation: Bradford Hill Criteria

Nine criteria help evaluate whether an observed association may be causal. Strength: strong associations are more likely causal, though weak ones are not excluded. Consistency: the association is observed repeatedly across different studies, populations, and settings. Specificity: the exposure leads to a specific outcome, though this criterion has limited applicability. Temporality: the exposure precedes the outcome -- this is the only essential criterion. Biological gradient (dose-response): greater exposure leads to greater risk. Plausibility: a biologically plausible mechanism exists. Coherence: the association is consistent with known biology and epidemiology. Experiment: interventional evidence supports the association (for example, removing the exposure reduces disease). Analogy: similar exposures cause similar effects.

These are guidelines, not rigid rules. No single criterion (except temporality) is necessary or sufficient for establishing causation.

VIII. Applying Evidence to Clinical Practice

When applying evidence to clinical practice, several considerations come into play. The patient's specific characteristics matter: were similar patients included in the study, and are there reasons the results might not apply (age, comorbidities, genetics)? The clinical setting matters: resource availability and healthcare system differences can affect applicability. Patient values and preferences should be incorporated through shared decision-making. The totality of evidence, not just a single study, should be considered. The GRADE framework (Grading of Recommendations Assessment, Development and Evaluation) provides a systematic approach to rating the quality of evidence and the strength of recommendations.

<image>A GRADE evidence quality summary table for a clinical intervention. Rows represent different outcomes (mortality, hospitalization, adverse events). Columns include: number of studies, study design, risk of bias assessment, inconsistency, indirectness, imprecision, publication bias, overall quality (high/moderate/low/very low), and the effect estimate with CI. Color-coded quality ratings (green for high, yellow for moderate, orange for low, red for very low) help visualize the evidence strength for each outcome.</image>


Lecture 27: Interpreting Medical Literature Critically — figure 1
Lecture 27: Interpreting Medical Literature Critically — figure 2
Lecture 27: Interpreting Medical Literature Critically — figure 3

Read this lecture as Markdown