Medical Research · Supplementary · from Medical Research
Case 3: Evaluating Systematic Review Quality
Patient Presentation
Demographics: 45-year-old female primary care physician
Chief Complaint: "A systematic review says we should be screening all adults for depression annually, but I'm not sure the evidence supports this."
History of Present Illness: A 45-year-old primary care physician attends a departmental journal club where a recently published systematic review and meta-analysis is presented. The review concludes that universal annual depression screening in primary care reduces depression severity at 6 months (standardized mean difference -0.32, 95% CI -0.48 to -0.16, p<0.001) and recommends implementation in all primary care settings.
The physician is skeptical because she has noticed that her practice already uses the PHQ-2 as a screening tool, and she has observed that many patients who screen positive do not follow through with recommended treatment. She wants to critically evaluate the systematic review before changing her practice's screening interval from the current "opportunistic" approach to a mandated annual protocol.
Upon detailed review, she identifies several methodological concerns: the search strategy was limited to two databases (PubMed and PsycINFO) and excluded non-English publications, the included studies had heterogeneous screening tools (PHQ-9, BDI-II, CES-D, HADS), only 4 of 12 included studies were RCTs (the remainder were observational), the I² statistic was 74% indicating substantial heterogeneity, no funnel plot or Egger's test for publication bias was reported, and the GRADE certainty of evidence was not assessed.
Past Medical History:
- Not applicable (physician is the evaluator, not the patient)
Medications:
- Not applicable
Social History:
- Not applicable
Family History:
- Not applicable
Physical Examination
- Not applicable (this is a research methodology case)
Workup and Results
Systematic Review Quality Assessment (AMSTAR-2 Checklist):
| AMSTAR-2 Domain | Assessment | Rating |
|---|---|---|
| PICO research question | Clearly defined | Adequate |
| Protocol registered a priori | No registration found (PROSPERO not mentioned) | Critical flaw |
| Literature search strategy | Only 2 databases; no grey literature; English only | Critical flaw |
| Study selection in duplicate | Yes, two reviewers with kappa reported | Adequate |
| Data extraction in duplicate | Yes | Adequate |
| Excluded studies listed with reasons | Only broad categories provided | Minor flaw |
| Included study descriptions | Adequate PICO descriptions | Adequate |
| Risk of bias assessment | Cochrane RoB for RCTs, but no tool for observational studies | Critical flaw |
| Meta-analysis methods appropriate | Random effects model used, but mixing RCTs and observational | Critical flaw |
| Publication bias assessed | Not reported | Critical flaw |
| Heterogeneity explored | I²=74% noted but not adequately explored via subgroup/sensitivity analysis | Critical flaw |
| Conflicts of interest | Lead author consultancy with screening tool manufacturer disclosed | Moderate concern |
| GRADE assessment | Not performed | Critical flaw |
Meta-Analysis Statistical Summary:
| Parameter | Value | Interpretation |
|---|---|---|
| Pooled SMD | -0.32 | Small-to-moderate effect |
| 95% CI | -0.48 to -0.16 | Statistically significant |
| I² heterogeneity | 74% | Substantial heterogeneity |
| Number of studies | 12 | - |
| RCTs included | 4 of 12 | Majority observational |
| Total participants | 14,832 | Adequate sample |
| Prediction interval | -0.71 to +0.07 | Crosses null -- true effect uncertain |
Clinical Image
Flowchart diagram illustrating the AMSTAR-2 critical appraisal tool for evaluating systematic review quality, highlighting the distinction between critical and non-critical domains and their impact on overall confidence ratings. Source: Educational illustration.
Diagnosis
Critically Low-Quality Systematic Review with Multiple Methodological Flaws; Insufficient Evidence to Change Screening Practice
Key Diagnostic Criteria:
- Seven critical flaws identified on AMSTAR-2 assessment (overall rating: critically low confidence)
- No protocol pre-registration
- Inadequate search strategy (only 2 databases, English-only)
- Mixing of RCTs and observational studies in meta-analysis without appropriate sensitivity analysis
- Substantial unexplored heterogeneity (I²=74%)
- No publication bias assessment
- No GRADE evidence certainty rating
- Prediction interval crosses the null, suggesting the true effect may include no benefit
Treatment Plan
- Practice decision: Do not change current opportunistic depression screening protocol based on this systematic review alone
- Search for higher-quality systematic reviews on the same topic (e.g., USPSTF systematic evidence review)
- Evaluate the USPSTF Grade B recommendation for depression screening, which is based on a more rigorous evidence synthesis
- If adopting annual screening, ensure adequate follow-up infrastructure (counseling access, psychiatry referral, care coordination)
- Conduct local quality improvement analysis: What percentage of PHQ-2-positive patients in the practice currently receive appropriate follow-up?
- Present journal club critique findings to the department with AMSTAR-2 assessment results
- Advocate for evidence-based guidelines to inform practice changes rather than individual published reviews
Key Learning Points
- AMSTAR-2 is the validated tool for critical appraisal of systematic reviews; it distinguishes critical from non-critical methodological flaws
- A statistically significant meta-analysis result does not guarantee high-quality evidence; the underlying methodology determines confidence in the findings
- The prediction interval, unlike the confidence interval, accounts for between-study heterogeneity and provides a range of plausible treatment effects in future settings
- Mixing RCTs and observational studies in meta-analysis without sensitivity analysis inflates confidence in findings that may be driven by confounding
- Publication bias assessment (funnel plots, Egger's test) is essential because small negative studies are less likely to be published, inflating pooled effect estimates