Premed · Premed · Genetics

Lecture 20: Quantitative Genetics and Complex Traits

Genetics


Learning Objectives

By the end of this lecture, students will be able to:

  1. Distinguish between qualitative (Mendelian) and quantitative (complex) traits
  2. Explain the polygenic model of inheritance and how multiple loci produce continuous phenotypic variation
  3. Define heritability (broad-sense and narrow-sense) and interpret heritability estimates
  4. Describe the methods used to study complex traits: twin studies, adoption studies, GWAS
  5. Explain the concepts of QTL mapping and genome-wide association studies (GWAS)
  6. Discuss the genetic architecture of common human diseases (e.g., diabetes, heart disease, schizophrenia)

Lecture Content

I. Qualitative vs. Quantitative Traits

Qualitative traits fall into discrete phenotypic categories, such as Mendel's round versus wrinkled seeds or ABO blood type. They are typically controlled by one or a few genes with large effects and produce clear Mendelian ratios in crosses. Quantitative traits, by contrast, exhibit continuous phenotypic variation, such as height, blood pressure, skin color, and disease susceptibility. They are measured on a continuous scale, are often normally distributed in a population, and are controlled by many genes (polygenic) plus environmental factors. These are also called complex traits or multifactorial traits.

The multiple-factor hypothesis, developed by Nilsson-Ehle in 1909 and East in 1916, explains how multiple genes, each contributing a small additive effect, can produce continuous variation. As the number of contributing loci increases, the phenotypic distribution approaches a normal (Gaussian) curve. Wheat kernel color illustrates this nicely: 1 gene gives 3 classes, 2 genes give 5 classes, and n genes give (2n + 1) classes. The threshold model extends this concept to traits that appear qualitative (affected versus unaffected) but have an underlying continuous distribution of liability. Liability is polygenic plus environmental, and disease manifests only when liability exceeds a threshold. Examples include cleft palate, type 2 diabetes, and schizophrenia.

<image>Panel A: Comparison of qualitative vs. quantitative trait distributions — left panel shows a bar graph with discrete phenotypic classes (e.g., flower color: red, pink, white) with clear Mendelian ratios; right panel shows a smooth bell curve of continuous variation (e.g., human height) with the population mean marked. Panel B: Progressive demonstration of the polygenic model — a series of histograms showing phenotypic distributions for 1 gene (3 classes), 2 genes (5 classes), 3 genes (7 classes), and many genes (smooth normal curve), illustrating how increasing loci produce continuous variation. Panel C: Threshold model diagram — a normal distribution curve representing underlying liability, with a vertical threshold line; individuals to the right of the threshold are affected, those to the left are unaffected; shift of the liability distribution in relatives of affected individuals is shown as a second overlapping curve with higher proportion exceeding the threshold.</image>

II. Partitioning Phenotypic Variance

Total phenotypic variance (Vp) can be decomposed into genetic variance (Vg), environmental variance (Ve), and gene-environment interaction variance (Vgxe). Genetic variance is further partitioned into additive genetic variance (Va), which reflects the additive effects of individual alleles and is the component that responds to selection; dominance variance (Vd), arising from interactions between alleles at the same locus; and epistatic variance (Vi), arising from interactions between alleles at different loci. Thus Vg = Va + Vd + Vi.

Broad-sense heritability (H^2) equals Vg/Vp and represents the proportion of total phenotypic variance due to all genetic effects. Narrow-sense heritability (h^2) equals Va/Vp and captures only the additive genetic component. Narrow-sense heritability is more useful because it predicts the response to selection and ranges from 0 (no genetic contribution) to 1 (all variation is genetic).

Several important caveats apply to heritability estimates. Heritability is a population-level statistic, not a measure for an individual. It depends on the specific population and environment studied, so changing either can change h^2. High heritability does not mean the trait is uninfluenced by environment. And heritability does not explain differences between groups.

III. Methods for Estimating Heritability in Humans

Twin studies compare concordance or correlation between monozygotic (MZ) twins, who share approximately 100% of their DNA, and dizygotic (DZ) twins, who share approximately 50%. If MZ concordance is much greater than DZ concordance, a strong genetic component is indicated. Falconer's formula estimates h^2 as approximately 2(r_MZ - r_DZ), where r is the correlation coefficient. This approach assumes equal environments for MZ and DZ twins, and its limitations include shared prenatal environments, epigenetic differences between MZ twins, and gene-environment correlation. Adoption studies compare adopted individuals to their biological versus adoptive families: resemblance to biological parents indicates a genetic contribution, while resemblance to adoptive parents indicates an environmental one. Family studies measure correlations between relatives of known genetic relatedness; for example, the slope of parent-offspring regression approximates h^2/2 for one parent or h^2 for the mid-parent value.

Approximate heritability estimates for common traits include: height h^2 approximately 0.80, BMI 0.40-0.70, blood pressure 0.30-0.50, type 2 diabetes liability 0.25-0.50, schizophrenia liability approximately 0.80, and intelligence (IQ scores) 0.50-0.80.

<image>Panel A: Twin study design diagram — side-by-side comparison of MZ twins (from one zygote, genetically identical) and DZ twins (from two zygotes, ~50% shared DNA), with arrows showing how concordance rates are compared to estimate heritability. Panel B: Bar chart of heritability estimates for various human traits (height, BMI, blood pressure, schizophrenia, type 2 diabetes, IQ), with the genetic component (Vg) and environmental component (Ve) stacked in different colors. Panel C: Parent-offspring regression plot — mid-parent height on the x-axis and offspring height on the y-axis, with data points and a regression line whose slope estimates narrow-sense heritability; annotations showing that a slope of 0.80 indicates h^2 ~ 0.80 for height.</image>

IV. QTL Mapping and Linkage Analysis for Complex Traits

A Quantitative Trait Locus (QTL) is a region of the genome that contributes to variation in a quantitative trait. QTL mapping in model organisms involves crossing two inbred strains that differ in the trait of interest, genotyping F2 or backcross progeny at many markers across the genome, and using statistical methods (LOD scores, interval mapping) to identify markers associated with trait variation. A LOD score greater than 3 is typically considered significant for linkage.

Linkage analysis in human families uses large pedigrees with multiple affected individuals and tests for co-segregation of genetic markers with the trait. While successful for Mendelian disorders, linkage analysis has limited power for complex traits because individual genes have small effects. QTL intervals are typically large (10-30 cM), containing hundreds of genes, and statistical power is low for genes of small effect. These limitations drove the development of association-based approaches.

V. Genome-Wide Association Studies (GWAS)

GWAS is a hypothesis-free approach that tests hundreds of thousands to millions of SNPs across the genome for association with a trait. It is grounded in the common disease-common variant (CDCV) hypothesis, which proposes that common diseases are influenced by common genetic variants (minor allele frequency greater than 1-5%), each with a small effect.

The typical study design compares allele or genotype frequencies between affected individuals and controls (case-control design) or tests associations between SNP genotype and trait values in a population sample (quantitative trait design). Genotyping is performed on SNP arrays with 500,000 to 5 million SNPs, and imputation uses linkage disequilibrium to infer genotypes at untyped positions. Statistical considerations are rigorous: the genome-wide significance threshold is p < 5 x 10^-8 to account for multiple testing, results are displayed as Manhattan plots (-log10 p-value against chromosomal position), QQ plots assess inflation, and population stratification is corrected using genomic control or principal components.

GWAS have identified thousands of loci for hundreds of traits and diseases, but most individual effect sizes are very small (odds ratios of 1.05-1.30). A persistent challenge is missing heritability: GWAS-identified loci typically explain only 10-30% of the estimated heritability. Possible explanations include rare variants, gene-gene interactions, gene-environment interactions, epigenetics, and structural variants. Polygenic risk scores (PRS) sum the effects of risk alleles across many loci, weighted by their effect sizes, to stratify individuals by genetic risk. Clinical utility is growing but still limited, with important concerns about portability across ancestry groups.

<image>Panel A: GWAS study design schematic — large cohort of cases and controls genotyped on SNP arrays, statistical testing of each SNP, results plotted as a Manhattan plot with chromosomes on the x-axis and -log10(p-value) on the y-axis, with significant peaks above the genome-wide significance line (p = 5 x 10^-8) labeled. Panel B: The concept of missing heritability — a pie chart showing the proportion of heritability for a trait (e.g., height) explained by GWAS-identified common variants versus unexplained heritability, with labels for possible sources of missing heritability (rare variants, epistasis, gene-environment interactions, structural variants). Panel C: Polygenic risk score illustration — a bell curve showing PRS distribution in a population, with the tails highlighted to show individuals at high genetic risk (right tail) and low genetic risk (left tail), and a comparison of disease prevalence in different PRS percentile groups.</image>

VI. Gene-Environment Interactions and Epigenetics in Complex Traits

Gene-environment interaction (GxE) occurs when the effect of a genotype on phenotype depends on the environment, and vice versa. PKU provides a clear example: the PAH genotype causes disease only with phenylalanine exposure from the diet. The FTO obesity risk allele shows an effect size modified by physical activity level. GxE can explain why heritability estimates vary across different environments.

Gene-environment correlation (rGE) describes the situation in which individuals with certain genotypes are more likely to be exposed to certain environments. Passive rGE occurs when parents provide both genes and a matching environment (for example, books in the home and genes for reading ability). Active rGE occurs when individuals seek environments matching their genotype. Evocative rGE occurs when an individual's genotype-influenced behavior elicits specific environmental responses.

Epigenetic contributions through DNA methylation and histone modifications can mediate environmental effects on gene expression and may contribute to missing heritability if epigenetic states are heritable but not captured by SNP arrays. Ultimately, the phenotype for any complex trait is a product of genes, environment, and their interactions; rarely is a complex trait purely genetic or purely environmental.

VII. Clinical Relevance of Quantitative Genetics

Risk prediction for common diseases is inherently probabilistic rather than deterministic. Genetic counseling for complex traits differs fundamentally from counseling for Mendelian disorders: simple recurrence risks cannot be given, and empirical risk data must be used instead. Family history remains the most practical tool for assessing genetic predisposition. Pharmacogenomic variation is often quantitative, with drug metabolism rates falling on a continuum. Understanding the genetic architecture of disease also informs drug target identification, as GWAS loci enriched for drug targets have been shown to have higher success rates in clinical trials.


Lecture 20: Quantitative Genetics and Complex Traits — figure 1
Lecture 20: Quantitative Genetics and Complex Traits — figure 2
Lecture 20: Quantitative Genetics and Complex Traits — figure 3

Read this lecture as Markdown