Residency · Residency · Cardiothoracic Surgery

Understanding CT Surgery Clinical Trials and Evidence Hierarchy

Introduction

Evidence-based practice is the foundation of modern cardiothoracic surgery. The hierarchy of evidence guides clinical decision-making by ranking study designs based on their susceptibility to bias. Understanding clinical trial design, interpretation, and application is essential for the CT surgeon to critically appraise the literature, participate in research, and deliver optimal patient care grounded in the best available evidence.

Hierarchy of Evidence

Levels of Evidence

Level I evidence includes systematic reviews and meta-analyses of randomized controlled trials (RCTs) and well-designed large RCTs. Level II encompasses smaller RCTs and prospective cohort studies with concurrent controls. Level III includes retrospective cohort studies and case-control studies. Level IV comprises case series and cross-sectional studies. Level V consists of expert opinion, case reports, and bench research.

Grades of Recommendation

Grade A (Strong) recommendations are supported by Level I evidence where benefits clearly outweigh risks. Grade B (Moderate) recommendations are supported by Level II-III evidence where benefits likely outweigh risks. Grade C (Weak) recommendations are supported by Level IV-V evidence where the balance of benefits and risks is uncertain. Practice guidelines from STS, AATS, EACTS, and AHA/ACC use these frameworks to classify recommendations.

Randomized Controlled Trials in CT Surgery

Landmark RCTs

TrialComparisonKey FindingClinical Impact
SYNTAXPCI (DES) vs. CABG for LM/3VDCABG superior for complex disease (high SYNTAX score)SYNTAX score guides revascularization decisions
ARTSingle vs. bilateral IMA for CABGNo significant survival difference at 10 yearsBIMA use remains debated
PARTNER 1-3TAVR vs. SAVR across risk categoriesTAVR non-inferior or superior across risk spectrumTAVR expanded to low-risk patients
STICHCABG + medical vs. medical alone (ischemic CM)Long-term survival benefit with CABG (10-year)Supports revascularization in ischemic HF
ACOSOG Z0030LN sampling vs. complete dissection (NSCLC)Sampling non-inferior for early-stageReduced morbidity without staging compromise
CALGB 140503Segmentectomy vs. lobectomy (≤2 cm NSCLC)Segmentectomy non-inferiorLung-sparing approach for small peripheral tumors
CTSN (Mitral)Repair vs. replacement for ischemic MRHigher MR recurrence with repair; no mortality differencePatient selection critical for repair

The SYNTAX Trial compared PCI with drug-eluting stents versus CABG for left main and three-vessel coronary artery disease and established CABG superiority for complex disease with high SYNTAX scores. The ART Trial compared single versus bilateral IMA grafting for CABG, and its 10-year follow-up showed no significant difference in survival, though the results remain controversial. The PARTNER Trials established transcatheter aortic valve replacement (TAVR) as an alternative to surgical AVR across risk categories. ACOSOG Z0030 demonstrated that mediastinal lymph node sampling is non-inferior to complete dissection for early-stage NSCLC staging. CALGB 140503 showed that sublobar resection (segmentectomy) is non-inferior to lobectomy for small (2 cm or less) peripheral NSCLC.

Challenges of RCTs in Surgery

Double-blinding is nearly impossible in surgical trials, and sham surgery raises ethical concerns. Surgeon skill and technique vary, and training effects and learning curves introduce variability that makes standardization difficult. Crossover occurs when patients randomized to one arm move to the other, such as medical therapy patients requiring surgery. Recruitment is hampered when patients and surgeons resist randomization because one treatment is perceived as superior. Relatively uncommon conditions in CT surgery make large RCTs logistically difficult due to sample size requirements. Surgical outcomes may take years or decades to manifest, and long follow-up periods lead to attrition that complicates analysis.

Observational Studies in CT Surgery

Registry-Based Research

The STS National Database is the largest CT surgery database, containing more than 7 million records and enabling risk-adjusted outcomes analysis, benchmarking, and observational research. The TVT Registry tracks all commercial TAVR procedures in the United States. The ELSO Registry is an international registry for extracorporeal membrane oxygenation outcomes. Registry-based research offers strengths including large sample sizes, real-world generalizability, and long-term follow-up, but is limited by selection bias, confounding variables, missing data, and the inability to establish causation.

Propensity Score Methods

Propensity score matching is a statistical technique that creates comparable groups from observational data by matching patients based on the probability of receiving the treatment. Inverse probability of treatment weighting (IPTW) is an alternative approach that weights observations to create a pseudo-randomized population. These methods can only adjust for measured confounders, and unmeasured confounders remain a source of bias. Propensity score methods are increasingly used in CT surgery literature to address selection bias in the absence of RCTs.

Meta-Analysis

Meta-analysis provides a systematic quantitative synthesis of results from multiple studies addressing the same question. A fixed-effects model assumes a single true effect size across studies, while a random-effects model accounts for heterogeneity between studies and is more conservative. Heterogeneity is assessed using the I-squared statistic, and publication bias is evaluated with funnel plots and Egger's test. Network meta-analysis allows indirect comparison of treatments that have not been directly compared in head-to-head trials.

Critical Appraisal of Surgical Literature

Key Questions When Evaluating a Study

When evaluating a study, one should consider whether the study design was appropriate for the clinical question, whether adequate randomization and allocation concealment were present for RCTs, and whether the groups were similar at baseline with confounders addressed. Follow-up should be adequate with acceptable loss to follow-up (less than 20%). Outcomes should be assessed using intention-to-treat analysis. Results should be evaluated for clinical significance, not just statistical significance, and the generalizability to one's own patient population should be considered.

Common Biases in Surgical Research

Selection bias involves systematic differences between comparison groups. Performance bias arises from differences in care provided beyond the intervention being studied. Detection bias involves differences in outcome assessment between groups. Attrition bias stems from systematic differences in withdrawals between groups. Publication bias occurs because positive results are more likely to be published than negative results. Survivorship bias results from analyzing only patients who survived to a certain time point.

Applying Evidence to Practice

From Evidence to Guidelines

Clinical practice guidelines synthesize the best available evidence into actionable recommendations. STS/AATS guidelines cover CABG, valve surgery, lung cancer, and other CT surgery domains. Guidelines must be interpreted in the context of individual patient factors, values, and preferences. Shared decision-making integrates evidence, clinical expertise, and patient preferences.

Quality of Evidence vs. Strength of Recommendation

High-quality evidence may still generate a weak recommendation if benefits and harms are closely balanced. Conversely, low-quality evidence may generate a strong recommendation when the intervention is clearly beneficial and risk is low. The GRADE framework (Grading of Recommendations Assessment, Development and Evaluation) is widely adopted for guideline development.

Key Clinical Pearls

The SYNTAX trial remains the cornerstone for guiding revascularization decisions between PCI and CABG, with the SYNTAX score objectively quantifying coronary disease complexity. Surgical RCTs are inherently challenging due to the inability to blind, surgeon variability, and recruitment difficulties, making well-designed observational studies with rigorous statistical methods valuable complements. Propensity score matching can reduce but never eliminate selection bias, as unmeasured confounders remain a fundamental limitation of observational data. When reading a study, one should always distinguish between statistical significance (p-value) and clinical significance (effect size and clinical relevance). The best clinical decisions integrate the highest-quality evidence with individual patient anatomy, comorbidities, and informed preferences.

References

  1. Defined by Oxford Centre for Evidence-Based Medicine. Levels of Evidence Working Group. Oxford Centre for Evidence-Based Medicine. 2011.
  2. Mohr FW, Morice MC, Kappetein AP, et al. Coronary artery bypass graft surgery versus percutaneous coronary intervention in patients with three-vessel disease and left main coronary disease: 5-year follow-up of the randomised, clinical SYNTAX trial. Lancet. 2013;381(9867):629-638.
  3. Mack MJ, Leon MB, Thourani VH, et al. Transcatheter aortic-valve replacement with a balloon-expandable valve in low-risk patients (PARTNER 3). New England Journal of Medicine. 2019;380(18):1695-1705.
  4. Altman DG, Bland JM. How to obtain the P value from a confidence interval. BMJ. 2011;343:d2304.

Read this lecture as Markdown