# Genome-Wide Association Studies: Methodology and Interpretation

## Study Design

### Principles

Genome-wide association studies (GWAS) take a hypothesis-free approach to identifying genetic variants associated with a trait or disease across the entire genome. The fundamental strategy is to compare allele frequencies of hundreds of thousands to millions of single nucleotide polymorphisms (SNPs) between cases and controls, or across the distribution of a quantitative trait. GWAS relies on linkage disequilibrium (LD), whereby tag SNPs on genotyping arrays serve as proxies for nearby causal variants that may not be directly genotyped. Standard genotyping platforms from Illumina and Affymetrix interrogate between 500,000 and 5,000,000 SNPs, and imputation using reference panels such as 1000 Genomes and TOPMed fills in the remaining gaps.

### Study Populations

The most common design is the case-control study, which requires careful matching for ancestry and population stratification. Quantitative trait designs analyze continuous phenotypes such as height, LDL cholesterol, or blood pressure across a cohort. Family-based designs like the transmission disequilibrium test (TDT) are robust to population stratification but have lower statistical power. Biobank-based GWAS using resources like the UK Biobank (approximately 500,000 participants), All of Us, BioBank Japan, and FinnGen provide large sample sizes with diverse phenotypes drawn from electronic health records.

### Quality Control

Rigorous quality control is essential at both the sample and variant levels. Sample-level QC removes individuals with low call rates, extreme heterozygosity, cryptic relatedness, or sex discrepancies. Variant-level QC excludes markers with low call rates, deviation from Hardy-Weinberg equilibrium in controls, or low minor allele frequency. Population stratification is corrected using principal component analysis (PCA), genomic control (lambda inflation factor), and mixed linear models such as BOLT-LMM and SAIGE.

## Statistical Framework

### Significance Threshold

The genome-wide significance threshold is set at p < 5 x 10^-8, reflecting a Bonferroni correction for approximately one million independent tests across the genome. A suggestive significance level of p < 1 x 10^-5 is often used for prioritizing variants for replication. Results are typically visualized using a Manhattan plot, which displays -log10(p-value) against chromosomal position, with peaks representing associated loci. The QQ plot of observed versus expected p-values serves as a quality diagnostic: early deviation from the diagonal suggests population stratification or other systematic bias, while deviation only in the tail reflects true associations.

### Effect Sizes

Effect sizes are reported as odds ratios for binary traits and beta coefficients for quantitative traits. Most GWAS hits for common diseases have small odds ratios in the range of 1.05 to 1.3. Larger effect sizes are typically found for pharmacogenomic loci, autoimmune HLA associations, and rare variant associations. Forest plots are used to display effect sizes across studies in meta-analyses, allowing assessment of consistency and heterogeneity.

### Power Considerations

GWAS requires very large sample sizes because effect sizes are small and the significance threshold is stringent. Statistical power depends on sample size, effect size, allele frequency, and the degree of LD between the tag SNP and the causal variant. Modern mega-GWAS and biobank studies now include hundreds of thousands to millions of participants. Meta-analysis of multiple cohorts further increases power while controlling for study-specific confounders.

## Linkage Disequilibrium

### Definition and Relevance

Linkage disequilibrium is the non-random association of alleles at different loci, measured by r-squared on a scale from 0 to 1. Because GWAS identifies tag SNPs in LD with causal variants rather than the causal variants themselves, the resolution of association mapping depends heavily on LD structure. LD blocks average approximately 50 to 100 kilobases in European populations but are substantially shorter in African populations, which have older and more extensively recombined genomes. The shorter LD in African populations provides better resolution for fine-mapping but requires denser arrays or larger sample sizes to achieve adequate power.

### Fine-Mapping

After GWAS identifies a locus, fine-mapping is performed to narrow the association to a credible set of candidate causal variants. Statistical fine-mapping methods such as FINEMAP, SuSiE, and PAINTOR generate probabilistic rankings of variants within a locus. Trans-ethnic fine-mapping exploits differences in LD structure across populations to narrow these credible sets further. Functional annotation data, including chromatin accessibility and expression quantitative trait locus (eQTL) information, is integrated with statistical evidence to prioritize the most likely causal variants.

## From GWAS Hits to Biology

### The Challenge of Functional Interpretation

Translating GWAS associations into biological understanding is a major challenge. Over 90% of GWAS-significant variants reside in noncoding regions -- either intergenic or intronic -- making their functional roles difficult to discern. The nearest gene to a GWAS hit is the causal gene only about 50% of the time, because regulatory variants may affect genes at considerable distances through enhancer-promoter interactions.

### Approaches to Identify Causal Genes

Several complementary approaches are used to bridge the gap from association to mechanism. eQTL mapping identifies variants associated with gene expression levels in relevant tissues, with the GTEx consortium providing a major resource. Chromatin interaction data from Hi-C and promoter capture Hi-C reveal physical contacts between GWAS loci and target gene promoters. Epigenomic annotation using ENCODE and Roadmap regulatory element data identifies overlap with enhancers, promoters, and DNase hypersensitive sites in disease-relevant cell types. CRISPR-based validation through CRISPRi/CRISPRa screens and targeted editing allows direct testing of regulatory variant function. Colocalization analysis using methods like COLOC, SMR, and TWAS tests whether a GWAS signal and an eQTL signal share the same causal variant. Mendelian randomization uses genetic variants as instrumental variables to infer causal relationships between exposures and outcomes.

## Missing Heritability

### The Problem

Twin and family studies have estimated high heritability for most common diseases -- approximately 80% for height and schizophrenia, and roughly 50% for type 2 diabetes. However, GWAS-identified common variants collectively explain only a fraction of this estimated heritability. Even with the largest available GWAS, common variants account for approximately 40 to 50% of heritability for height, about 25% for schizophrenia, and around 18% for type 2 diabetes.

### Potential Explanations

Several factors likely contribute to this "missing heritability." Many associated common variants remain undiscovered because GWAS sample sizes, while large, are still insufficient to detect variants with very small effects. Rare variants with minor allele frequencies below 1%, which may have larger individual effect sizes, are poorly captured by standard GWAS arrays and require sequencing approaches. Structural variants such as copy number variations, inversions, and complex rearrangements are also not well captured by SNP arrays. Gene-gene interactions (epistasis) are difficult to detect and remain largely unexplored in GWAS. Gene-environment interactions, where environmental modifiers alter genetic effects, are similarly not captured. There is also debate about whether twin-based heritability estimates themselves may be inflated by assumptions about shared environment. LD score regression is a method used to estimate SNP heritability -- the proportion of heritability explained by common SNPs genome-wide, including those below GWAS significance.

### Rare Variant Approaches

Whole exome and whole genome sequencing-based association studies, such as the UK Biobank whole-exome sequencing of 500,000 individuals, are beginning to address the rare variant contribution. Burden tests (SKAT, SKAT-O) aggregate rare variants within a gene to increase statistical power, and collapsing methods test whether the cumulative burden of rare variants in a gene differs between cases and controls. These rare variant GWAS are identifying new genes and mechanisms not captured by common variant GWAS.

## Key GWAS Findings Across Disease Areas

### Autoimmune/Inflammatory

The HLA region produces the strongest GWAS associations for many autoimmune conditions. Celiac disease has an odds ratio of approximately 6 for HLA-DQ2/DQ8, and strong HLA associations also underlie type 1 diabetes, rheumatoid arthritis, and ankylosing spondylitis. Over 200 loci have been identified for inflammatory bowel disease, highlighting the roles of immune function and barrier integrity pathways.

### Neuropsychiatric

Schizophrenia GWAS has identified over 270 loci with enrichment in synaptic and neuronal pathways, along with a notable HLA association that suggests a possible autoimmune component. For Alzheimer disease, APOE4 remains the strongest common variant risk factor, with additional loci in microglial and immune pathways including TREM2, BIN1, and CLU.

### Cardiometabolic

Over 160 loci have been identified for coronary artery disease, implicating LDL and lipid pathways, inflammation, and nitric oxide signaling. Type 2 diabetes GWAS has identified over 400 loci, with beta-cell function and insulin secretion pathways predominating over insulin resistance loci.

### Cancer

Breast cancer GWAS has identified over 200 loci, many located in regulatory regions that affect gene expression in mammary tissue. Prostate cancer has over 260 associated loci, with the 8q24 region containing multiple independent signals in a gene desert -- a region devoid of annotated genes.

| Disease Area | Example Condition | Number of Loci Identified | Strongest Associations | Key Biological Pathways |
|---|---|---|---|---|
| Autoimmune | Celiac disease | >40 | HLA-DQ2/DQ8 (OR ~6) | Immune function; barrier integrity |
| Neuropsychiatric | Schizophrenia | >270 | Multiple HLA loci | Synaptic; neuronal; immune |
| Cardiometabolic | Type 2 diabetes | >400 | TCF7L2 (OR ~1.4) | Beta-cell function; insulin secretion |
| Cardiometabolic | Coronary artery disease | >160 | 9p21 (CDKN2A/B locus) | LDL/lipid pathways; inflammation |
| Cancer | Breast cancer | >200 | Multiple regulatory regions | Mammary tissue gene regulation |
| Cancer | Prostate cancer | >260 | 8q24 gene desert | Unknown (multiple independent signals) |

## GWAS Catalog and Resources

The NHGRI-EBI GWAS Catalog is a curated database of all published GWAS, encompassing over 6,000 publications and more than 400,000 associations. Open Targets Genetics integrates GWAS, eQTL, and functional data to support drug target identification. The PheWAS (Phenome-Wide Association Study) takes a reverse approach, testing one variant against thousands of phenotypes to discover pleiotropic effects.

<image>A composite GWAS methodology figure with four panels. Panel 1: A Manhattan plot showing genome-wide results across all chromosomes, with the genome-wide significance threshold (red dashed line at p = 5 x 10^-8) and several peaks exceeding the threshold labeled with candidate gene names. Panel 2: A QQ plot showing observed vs. expected p-values, with a well-calibrated baseline and deviation at the tail indicating true associations. Panel 3: A regional association plot (LocusZoom) of a single significant locus, showing LD structure (color-coded by r-squared with the lead SNP) and nearby genes. Panel 4: A forest plot from a meta-analysis of four cohorts showing consistent effect sizes with a combined diamond at the bottom.</image>

<image>A schematic illustrating the challenge of translating GWAS hits to biological mechanism. A chromosomal region is shown with three genes (Gene A, Gene B, Gene C) and the GWAS lead SNP located in an intergenic enhancer element between genes A and B. Incorrect interpretation arrow points to the nearest gene (Gene A). Correct interpretation is shown through: (1) an eQTL analysis linking the variant to expression of Gene C (the distant gene), (2) a chromatin conformation capture (Hi-C) arc showing physical interaction between the enhancer and Gene C promoter, and (3) CRISPR deletion of the enhancer leading to reduced Gene C expression. The lesson is highlighted: the nearest gene is not always the causal gene.</image>

## Clinical Pearls

GWAS identifies statistical associations, not causal variants or genes. The lead SNP is a tag for a locus, and substantial downstream work -- including eQTL mapping, chromatin interaction studies, and functional validation -- is needed to identify the causal gene and mechanism. The nearest gene to a GWAS hit is the causal gene only about 50% of the time, which underscores the importance of functional data for proper gene assignment. GWAS common variant associations typically have very small effect sizes (odds ratios of 1.05 to 1.3) and are not useful for individual-level prediction; their value lies in pathway discovery, drug target identification, and aggregate polygenic risk scoring. The HLA region produces the strongest GWAS associations for autoimmune diseases, but its complex LD structure and extreme polymorphism require specialized imputation and analysis methods. Missing heritability is being progressively explained by larger sample sizes, rare variant sequencing studies, and improved statistical methods -- it is not evidence that genetics is unimportant for common disease. African-ancestry populations, with their shorter LD blocks and greater genetic diversity, are invaluable for fine-mapping GWAS loci, yet they remain severely underrepresented in GWAS, with significant equity implications. GWAS results should always be interpreted in the context of the study population's ancestry, as effect sizes and risk allele frequencies may differ across populations.

## References

- Uffelmann E et al. Genome-wide association studies. Nat Rev Methods Primers. 2021;1:59.
- Visscher PM et al. 10 years of GWAS discovery: biology, function, and translation. Am J Hum Genet. 2017;101(1):5-22.
- Buniello A et al. The NHGRI-EBI GWAS Catalog of published genome-wide association studies. Nucleic Acids Res. 2019;47(D1):D1005-D1012.
- Gallagher MD, Chen-Plotkin AS. The post-GWAS era: from association to function. Am J Hum Genet. 2018;102(5):717-730.
- Manolio TA et al. Finding the missing heritability of complex diseases. Nature. 2009;461(7265):747-753.
