Premed · Premed · Statistics Biostatistics

Lecture 1: Introduction to Statistics in the Health Sciences

Statistics / Biostatistics


Learning Objectives

By the end of this lecture, students will be able to:

  1. Define statistics and biostatistics and explain their role in the health sciences
  2. Distinguish between descriptive and inferential statistics
  3. Identify key terms: population, sample, parameter, statistic
  4. Explain why statistical reasoning is essential for evidence-based medicine
  5. Describe the general workflow of a biostatistical investigation

Lecture Content

I. What Is Statistics?

Statistics is the science of collecting, organizing, analyzing, interpreting, and presenting data. When these methods are applied specifically to biological, medical, and public health problems, the field is known as biostatistics. The core purpose of statistics is to enable informed decisions under uncertainty. Because clinical medicine is inherently uncertain, statistics provides the essential framework for managing that uncertainty and bridging the gap between raw data and meaningful clinical conclusions.

<image>A flowchart showing the statistical pipeline in health research. Panel A: "Research Question" flows into "Study Design," then "Data Collection," "Data Analysis," and finally "Interpretation and Decision." Panel B: At each stage, common pitfalls are annotated (e.g., bias at collection, confounding at analysis, overinterpretation at the final stage).</image>

II. Descriptive vs. Inferential Statistics

Statistics is broadly divided into two branches. Descriptive statistics encompasses the methods used to summarize and organize data from a sample or population, including tools like means, medians, proportions, tables, and graphs. Its goal is to provide a clear snapshot of the data at hand. Inferential statistics, on the other hand, uses methods to draw conclusions about a population based on sample data. Examples include hypothesis tests, confidence intervals, and regression models, all of which aim to generalize findings beyond the observed data.

Both branches are essential in clinical research. Descriptive statistics characterize the study sample, giving readers a concrete picture of who was studied and what was measured, while inferential statistics test hypotheses and estimate population effects, allowing researchers to make broader claims about how a treatment or exposure behaves in the population at large.

III. Key Terminology

A population is the entire group of individuals or observations of interest -- for example, all adults with Type 2 diabetes in Canada. A sample is a subset of the population selected for study, such as 500 adults with Type 2 diabetes recruited from three clinics. A parameter is a numerical summary of a population and is usually unknown; an example would be the true mean HbA1c of all Canadian diabetic adults. A statistic is a numerical summary computed from a sample, such as the sample mean HbA1c from the 500 recruited participants. A variable is any characteristic that can be measured or categorized, including blood pressure, sex, or treatment group, while an observation or data point is a single measurement or record in a dataset.

<image>A Venn-style diagram showing the relationship between population and sample. A large circle labeled "Population (N)" contains a smaller highlighted circle labeled "Sample (n)." Arrows from the sample point outward labeled "Inference" and arrows from the population inward labeled "Sampling." Key parameters (mu, sigma) are listed beside the population circle, and corresponding statistics (x-bar, s) beside the sample circle.</image>

IV. Why Statistics Matters in Medicine

Evidence-based medicine (EBM) requires the critical appraisal of research findings, and clinicians must be able to evaluate whether study results are valid and applicable. Statistical literacy helps clinicians interpret journal articles and clinical guidelines, understand drug efficacy data such as number needed to treat and relative risk reduction, communicate risk to patients in an understandable way, and avoid common pitfalls like confusing correlation with causation.

Real-world applications are everywhere. Clinicians use statistics when evaluating whether a new drug is superior to placebo, determining if a screening test is worth implementing, or assessing whether an observed adverse event rate is higher than expected. Without a solid grasp of statistics, clinicians risk misinterpreting evidence and making decisions that may not serve their patients well.

V. The Biostatistical Investigation Workflow

A biostatistical investigation follows a structured series of steps. The process begins with formulating a research question that is specific, measurable, and clinically relevant, such as "Does Drug A reduce 30-day mortality compared to Drug B in post-MI patients?" Next comes study design, which involves choosing an appropriate design (RCT, cohort, case-control, or cross-sectional) and defining inclusion and exclusion criteria, outcome measures, and sample size.

The third step is data collection, where the researcher ensures valid and reliable measurement instruments and minimizes bias through randomization, blinding, and standardized protocols. Data analysis follows, requiring the selection of appropriate statistical methods based on data type and study design, typically beginning with descriptive analysis before moving to inferential methods. Finally, the investigator interprets and reports results, considering clinical significance alongside statistical significance, reporting effect sizes, confidence intervals, and p-values, and acknowledging limitations of the study.

<image>A timeline diagram showing a typical clinical study from conception to publication. Key milestones include: protocol development, ethics approval, patient enrollment, data collection, interim analysis, final analysis, manuscript preparation, and peer review. Below the timeline, icons represent the statistical input required at each stage (e.g., sample size calculation during protocol, descriptive tables during analysis, forest plots during reporting).</image>

VI. Scales of Measurement (Preview)

Data can be measured on four scales, each with different properties. A nominal scale consists of categories with no inherent order, such as blood type (A, B, AB, O). An ordinal scale has categories with a meaningful order but unequal intervals, such as a pain scale (mild, moderate, severe). An interval scale is a numeric scale with equal intervals but no true zero, such as temperature in Celsius. A ratio scale is a numeric scale with equal intervals and a true zero, such as weight in kilograms or blood glucose.

The scale of measurement is important because it determines which statistical methods are appropriate for a given variable. This concept will be explored in depth in Lecture 2.

VII. Historical Context

Biostatistics has deep historical roots in the work of pioneering scientists. John Graunt produced early mortality tables in London in 1662. Florence Nightingale used statistical graphics in 1858 to advocate for sanitary reforms, demonstrating the power of data visualization. Ronald Fisher developed ANOVA, randomization, and experimental design principles in the 1920s, laying the groundwork for modern statistics. Austin Bradford Hill conducted the first modern randomized controlled trial in 1948, testing streptomycin for tuberculosis.

The evolution of biostatistics has been driven by practical clinical and public health needs, and modern biostatistics increasingly relies on computing power and large datasets to tackle increasingly complex questions.

<image>A historical timeline with portraits or icons of key figures in biostatistics. Panel A: John Graunt with a mortality table. Panel B: Florence Nightingale with her polar area diagram (coxcomb chart). Panel C: Ronald Fisher with an ANOVA table. Panel D: Bradford Hill with a diagram of an RCT design. Each panel includes a brief caption of their contribution.</image>


Lecture 1: Introduction to Statistics in the Health Sciences — figure 1
Lecture 1: Introduction to Statistics in the Health Sciences — figure 2
Lecture 1: Introduction to Statistics in the Health Sciences — figure 3
Lecture 1: Introduction to Statistics in the Health Sciences — figure 4

Read this lecture as Markdown