Residency · Residency · Pathology

Test Utilization and Evidence-Based Laboratory Medicine

Introduction

Inappropriate laboratory test utilization, both overutilization and underutilization, contributes to unnecessary healthcare costs, patient harm from false positives or missed diagnoses, and diagnostic confusion. Evidence-based laboratory medicine (EBLM) applies principles of evidence appraisal to laboratory test selection, interpretation, and clinical decision-making.

The Scope of the Problem

Overutilization

An estimated 20-50% of laboratory tests are unnecessary or inappropriate. Drivers include defensive medicine, routine order sets, standing orders, lack of knowledge about test performance, and "shotgun" testing. Consequences include unnecessary follow-up testing, invasive procedures for false positives, diagnostic confusion, increased costs, and patient discomfort from blood loss and phlebotomy pain. Laboratory costs represent approximately 3-5% of hospital budgets but influence approximately 60-70% of clinical decisions.

Underutilization

Failure to order appropriate tests leads to missed or delayed diagnoses. Examples include failure to order troponin in atypical chest pain, missing HbA1c screening, and inadequate coagulation testing in bleeding patients. Underutilization is less studied than overutilization but equally impactful on patient outcomes. Diagnostic stewardship addresses both overuse and underuse.

Principles of Evidence-Based Laboratory Medicine

Evaluating Diagnostic Test Performance

Sensitivity is the probability of a positive test in patients with the disease (true positive rate). Specificity is the probability of a negative test in patients without the disease (true negative rate). Positive predictive value (PPV) is the probability of disease given a positive test and depends on prevalence. Negative predictive value (NPV) is the probability of no disease given a negative test. Likelihood ratios (LR+ equals sensitivity divided by 1 minus specificity; LR- equals 1 minus sensitivity divided by specificity) are independent of prevalence.

Pre-Test and Post-Test Probability

Bayesian reasoning dictates that the clinical value of a test depends on the pre-test probability of disease. A test with excellent sensitivity and specificity may have poor PPV in low-prevalence populations, producing excessive false positives. Conversely, a test with moderate sensitivity can be highly useful when pre-test probability is high. The Fagan nomogram is a graphical tool for calculating post-test probability from pre-test probability and likelihood ratio. Pathologists should understand and communicate these concepts when consulted about test selection.

Study Design for Diagnostic Tests

STARD (Standards for Reporting Diagnostic Accuracy) provides guidelines for transparent reporting of diagnostic accuracy studies. The reference standard must be clearly defined and applied independently of the test under evaluation. Spectrum bias arises if the study population differs from the clinical population where the test will be used, meaning performance may not generalize. Verification bias occurs if the reference standard is only applied to test-positive patients, resulting in overestimation of sensitivity.

Test Utilization Strategies

Utilization Management Approaches

Hard stops prevent ordering of certain tests without meeting predefined criteria, such as no repeat BMP within 24 hours. Reflex and reflexive testing uses algorithm-driven sequential testing, for example TSH with reflex to free T4 only if TSH is abnormal. Minimum reorder intervals are system-enforced time limits between repeat orders for the same test. Order set optimization involves reviewing and modifying order sets to remove unnecessary tests, with involvement of clinical champions. Choosing Wisely, an ABIM Foundation initiative, has specialty societies publishing lists of overused tests, including pathology-specific recommendations.

Clinical Decision Support (CDS)

Best practice alerts (BPAs) are pop-up notifications at the time of ordering that are effective but prone to alert fatigue. Interruptive alerts are more effective than non-interruptive ones but are more likely to cause fatigue. Peer comparison reports show individual ordering patterns relative to peers as a "nudge" approach. Educational outreach involves active engagement with clinical teams about appropriate test utilization.

Formulary and Test Menu Management

The laboratory test formulary is analogous to a pharmacy formulary. Tests are added or removed based on clinical evidence, institutional needs, and cost-effectiveness. Send-out test management involves reviewing high-cost reference laboratory testing for appropriateness. Periodic review of low-volume tests considers potential removal or consolidation.

Specific Examples of Utilization Targets

High-Impact Targets

Ordering a complete metabolic panel (CMP) versus basic metabolic panel (BMP) should be considered carefully, as liver function tests are often unnecessary in patients without hepatic indication. Daily labs on stable inpatients involve excessive daily chemistry and CBC orders, and evidence supports less frequent monitoring. Type and screen in low-risk surgical patients may be unnecessary when transfusion is rare in many elective procedures, avoiding routine crossmatch. Coagulation cascade testing with PT/INR and aPTT as routine preoperative screening in patients without bleeding history has low yield. ESR and CRP are often ordered together, but CRP is generally more useful and ESR adds limited incremental value in most situations.

Molecular and Genetic Test Utilization

This rapidly growing area presents increasing utilization concerns due to expensive tests with complex interpretation. Syndromic panels (broad molecular panels) may detect clinically irrelevant organisms, highlighting the need for diagnostic stewardship. Cancer genomic panels require ensuring appropriate tumor type, adequate specimen, and clinical indication before ordering. Pharmacogenomic testing should be ordered for specific gene-drug pairs with strong clinical evidence, avoiding untargeted panels.

Measuring Utilization Outcomes

Key Metrics

Test volume per patient-day should be tracked over time and compared with benchmarks. Cost per patient encounter captures laboratory-related costs. Repeat order rates measure the frequency of identical tests within defined time windows. Diagnostic yield is the proportion of abnormal results for tests targeted for utilization review. Clinical outcome measures ensure that reduced testing does not adversely affect patient care.

Utilization Committees

A multidisciplinary committee should include pathologists, laboratorians, clinicians, informaticists, and administrators. The committee conducts regular review of utilization data, new test proposals, and send-out expenditures, and has the authority to implement utilization management interventions. Pathologist leadership is critical for credibility and clinical integration.

Communicating Test Results

Interpretive Reporting

Adding interpretive comments to complex test results guides clinical action. Reference ranges should include clinical context, and unexpected or discordant results should be flagged. Algorithmic interpretation applies cascade logic to multi-test panels, such as anemia workup or thyroid function testing. Pathologist consultation should be available for complex interpretive questions.

Reflective Testing Algorithms

For iron studies, ferritin is ordered first, with serum iron, TIBC, and transferrin saturation added if results are equivocal. For thyroid function, TSH serves as the initial screen, with reflex to free T4 and/or free T3 only when TSH is abnormal. Hemoglobin A1c reflexes to estimated average glucose and interpretation. These algorithms reduce unnecessary testing while ensuring clinically relevant information is provided.

Clinical Pearls

Approximately 20-50% of laboratory tests ordered are estimated to be unnecessary, contributing to patient harm through false-positive cascades, diagnostic confusion, and unnecessary costs. The positive predictive value of any test is heavily influenced by disease prevalence; ordering tests in low-probability populations increases false-positive rates and should be discouraged. Reflex and reflexive testing algorithms (such as the TSH-first thyroid cascade) are effective utilization strategies that reduce unnecessary testing while maintaining diagnostic quality. Pathologists should lead institutional test utilization committees and serve as consultants for evidence-based test selection and interpretation.

References

  1. Zhi M, et al. The landscape of inappropriate laboratory testing: a 15-year meta-analysis. PLoS One. 2013;8(11):e78962.
  2. Naugler C, et al. Reducing overutilization of laboratory testing. CMAJ. 2017;189(18):E654.
  3. Price CP, et al. Evidence-based laboratory medicine: principles, practice, and outcomes. Clin Chem Lab Med. 2012;50(3):403-412.
  4. Choosing Wisely. American Society for Clinical Pathology recommendations. Available at: https://www.choosingwisely.org/societies/american-society-for-clinical-pathology/.

Read this lecture as Markdown