# Structure of the Human Genome: From Nucleotides to Chromosomes

## Overview of Genome Organization

### Genome Size and Composition

The human genome contains approximately 3.2 billion base pairs distributed across 22 autosomes and 2 sex chromosomes. Only about 1.5% of the genome encodes proteins, corresponding to roughly 20,000 to 25,000 protein-coding genes. The remaining 98.5% consists of noncoding sequences that serve diverse functional and non-functional roles. The ENCODE project demonstrated that up to 80% of the genome may have some form of biochemical function, though the true biological relevance of much of this activity remains actively debated.

### Coding DNA

Exons constitute the protein-coding portions of genes, and the average gene contains approximately 8 to 10 exons. The exome, totaling about 30 to 40 megabases, represents roughly 1 to 2% of the genome but harbors approximately 85% of known disease-causing variants. Alternative splicing allows a single gene to produce multiple transcript isoforms, substantially expanding the proteome beyond what the gene count alone would suggest. Codon usage bias varies across species and can influence translation efficiency.

### Noncoding DNA

Introns comprise the majority of gene length and are far from inert; they contain regulatory elements, microRNAs, and small nucleolar RNAs. The untranslated regions of mRNA also serve critical roles: the 5' UTR regulates translation initiation, while the 3' UTR contains miRNA binding sites and polyadenylation signals that influence mRNA stability and localization. Beyond gene-associated sequences, the genome contains an elaborate set of regulatory elements including promoters, enhancers, silencers, and insulators that control where and when genes are expressed. Over 16,000 long noncoding RNAs (lncRNAs) have been identified, with roles in chromatin remodeling, transcriptional regulation, and imprinting, the most famous being XIST, which orchestrates X-chromosome inactivation. Additionally, approximately 2,600 mature microRNAs regulate post-transcriptional gene expression and are increasingly implicated in disease processes.

## Repetitive Elements

### Interspersed Repeats (Transposable Elements)

Transposable elements have colonized the human genome extensively over evolutionary time. Short interspersed nuclear elements (SINEs), particularly Alu elements, are the most abundant repetitive family, accounting for about 11% of the genome with over one million copies. Alu-mediated recombination is a clinically important source of pathogenic structural variants. Long interspersed nuclear elements (LINEs), specifically LINE-1 (L1) elements, comprise approximately 17% of the genome. Although most L1 copies are truncated and inactive, a small fraction remain retrotransposition-competent and can cause insertional mutagenesis. DNA transposons make up about 3% of the genome and are largely inactive in modern humans, though they historically played a major role in shaping genome architecture. Endogenous retroviruses (ERVs), remnants of ancestral retroviral integrations, account for roughly 8% of the genome.

| Repeat Class | Example | Genome Fraction | Copy Number | Clinical Relevance |
|---|---|---|---|---|
| SINEs | Alu | ~11% | >1,000,000 | NAHR causing pathogenic deletions/duplications |
| LINEs | LINE-1 (L1) | ~17% | ~500,000 | Insertional mutagenesis (retrotransposition) |
| DNA transposons | — | ~3% | — | Largely inactive; historical genome shaping |
| Endogenous retroviruses | ERVs | ~8% | — | Ancestral retroviral integrations |

### Tandem Repeats

Satellite DNA consists of large arrays of tandem repeats at centromeres (alpha-satellite) and pericentromeric regions, where it plays a critical role in chromosome segregation. Minisatellites, with repeat units of 10 to 60 base pairs, were historically used in DNA fingerprinting. Microsatellites, also known as short tandem repeats (STRs), consist of 1 to 6 base pair repeat units and form the basis of forensic DNA profiling. They are also the molecular substrate for triplet repeat expansion disorders such as Huntington disease and fragile X syndrome. Telomeric repeats, consisting of the hexameric sequence TTAGGG, cap chromosome termini. These repeats progressively shorten with each cell division and are maintained by the enzyme telomerase in stem cells and cancer cells.

## Chromatin Structure

### Levels of DNA Packaging

The fundamental unit of chromatin is the nucleosome, in which 147 base pairs of DNA wrap around a histone octamer composed of two copies each of histones H2A, H2B, H3, and H4. The linker histone H1 stabilizes higher-order chromatin structure. The existence of a 30-nm fiber as the next level of packaging has been controversial, with evidence for this structure debated in vivo. At a larger scale, chromatin forms loops of 100 to 1,000 kilobases organized by the cohesin complex and anchored by CTCF proteins; these loops are fundamental to gene regulation. Topologically associating domains (TADs) are self-interacting genomic regions spanning approximately 200 kilobases to 2 megabases that constrain enhancer-promoter interactions. At the highest level, individual chromosomes occupy distinct territories within the nucleus, with gene-rich chromosomes tending to localize toward the nuclear interior.

### Euchromatin vs. Heterochromatin

Euchromatin is the open, transcriptionally active form of chromatin, enriched for acetylated histones (H3K27ac, H3K9ac) and H3K4 methylation. Constitutive heterochromatin is permanently condensed, enriched in H3K9me3, and found at centromeres, telomeres, and pericentromeric regions. Facultative heterochromatin is conditionally silenced and enriched in H3K27me3, a mark deposited by the Polycomb repressive complex. The inactive X chromosome, or Barr body, is the most prominent example of facultative heterochromatin.

| Feature | Euchromatin | Constitutive Heterochromatin | Facultative Heterochromatin |
|---|---|---|---|
| Condensation state | Open | Permanently condensed | Conditionally condensed |
| Transcriptional activity | Active | Silent | Silenced (reversible) |
| Key histone marks | H3K27ac, H3K9ac, H3K4me | H3K9me3 | H3K27me3 |
| Genomic location | Gene-rich regions | Centromeres, telomeres, pericentromeric | Variable (e.g., inactive X) |
| Replication timing | Early S-phase | Late S-phase | Late S-phase |

### Histone Modifications and the Histone Code

Post-translational modifications of histone tails include acetylation, methylation, phosphorylation, ubiquitination, and SUMOylation. Specific combinations of these modifications, sometimes called the "histone code," recruit reader proteins that either activate or repress transcription depending on the modification pattern. Histone acetyltransferases (HATs) and histone deacetylases (HDACs) are of particular interest as therapeutic targets in cancer.

## Topologically Associating Domains (TADs)

### Structure and Function

TADs are megabase-scale genomic compartments that constrain regulatory interactions, ensuring that enhancers act on their intended target genes rather than on neighboring loci. TAD boundaries are enriched for CTCF binding sites in convergent orientation and for the cohesin complex. According to the loop extrusion model, cohesin extrudes chromatin loops until it is blocked by convergent CTCF sites, thereby defining TAD boundaries. TADs are largely conserved across cell types and species, suggesting they serve a fundamental structural role in genome organization.

### Clinical Significance of TAD Disruption

Structural variants that disrupt TAD boundaries can cause "enhancer hijacking," in which enhancers are repositioned to act on non-target genes. This mechanism has been demonstrated in limb malformations caused by TAD disruptions at the WNT6/IHH/EPHA4 locus, and in blepharophimosis-ptosis-epicanthus inversus syndrome (BPES) from disruption of the FOXL2 regulatory landscape. Copy number variants that encompass TAD boundaries may have more severe phenotypic effects than those contained within a single TAD. Understanding TAD architecture is becoming increasingly important for interpreting noncoding structural variants in clinical genomics.

## Functional Significance for Variant Interpretation

### Coding Variant Interpretation

Missense, nonsense, frameshift, and splice site variants are assessed using the ACMG/AMP classification criteria. Constraint metrics such as pLI and LOEUF from gnomAD indicate how tolerant a gene is to loss-of-function variation at the gene level. Conservation scores like GERP and PhyloP reflect evolutionary constraint at the nucleotide level, helping to prioritize variants in conserved positions.

### Noncoding Variant Interpretation

Interpreting noncoding variants remains challenging, particularly because most GWAS hits fall in noncoding regions. Regulatory element annotations from the ENCODE and Roadmap Epigenomics projects help prioritize noncoding variants for further evaluation. Disruption of CTCF or cohesin binding sites, splice regulatory elements, and miRNA binding sites represent emerging classes of pathogenic noncoding variants. Deep intronic variants that create cryptic splice sites are increasingly recognized and can be detected through RNA sequencing.

<image>A multi-level diagram showing DNA packaging from the double helix (2 nm) through nucleosome beads-on-a-string (11 nm), the 30-nm chromatin fiber, chromatin loops and TADs, to the fully condensed metaphase chromosome. Each level is labeled with its approximate diameter and key associated proteins (histones, CTCF, cohesin, condensin). Color-coded histone modifications (green acetylation marks on open euchromatin, red methylation marks on condensed heterochromatin) are shown at the nucleosome level.</image>

<image>A schematic representation of a topologically associating domain (TAD) showing a chromatin loop anchored by convergent CTCF binding sites (depicted as red and blue arrows indicating orientation) with cohesin ring complex at the base. Inside the TAD, an enhancer element (yellow oval) is shown interacting with its target gene promoter (green arrow). Outside the TAD boundary, a neighboring gene is shielded from the enhancer. A second panel shows the pathological state where a deletion removes the TAD boundary, allowing the enhancer to aberrantly activate the neighboring gene, indicated by a red dashed arrow.</image>

<image>A pie chart of human genome composition showing: protein-coding exons (1.5%), introns (25%), intergenic unique sequences (15%), SINEs including Alu (13%), LINEs including L1 (20%), DNA transposons (3%), LTR retrotransposons/ERVs (8%), tandem repeats and satellite DNA (3%), other sequences (11.5%). Each slice is distinctly colored with a legend identifying the repeat family. An inset shows the structure of a typical Alu element with A and B boxes and a poly-A tail.</image>

## Clinical Pearls

The exome represents only about 1.5% of the genome but harbors approximately 85% of known disease-causing variants, which explains why exome sequencing remains a cost-effective first-tier genomic test for many clinical indications. Alu-mediated non-allelic homologous recombination is a common mechanism underlying pathogenic deletions and duplications, particularly in gene-rich regions with high Alu density. TAD boundary disruption is an emerging mechanism of disease, and structural variants spanning TAD boundaries may have unexpectedly severe phenotypic consequences through enhancer hijacking. Triplet repeat expansions in short tandem repeats cause a growing list of neurogenetic disorders and require specialized assays such as repeat-primed PCR, Southern blot, or long-read sequencing, as they are not reliably detected by standard next-generation sequencing. Noncoding variant interpretation remains a major frontier in genomic medicine, and RNA sequencing is an increasingly valuable adjunct for resolving variants of uncertain significance that fall in noncoding regions.

## References

- Lander ES et al. Initial sequencing and analysis of the human genome. Nature. 2001;409:860-921.
- ENCODE Project Consortium. An integrated encyclopedia of DNA elements in the human genome. Nature. 2012;489:57-74.
- Lupiáñez DG et al. Disruptions of topological chromatin domains cause pathogenic rewiring of gene-enhancer interactions. Cell. 2015;161(5):1012-1025.
- Spielmann M, Lupiáñez DG, Mundlos S. Structural variation in the 3D genome. Nat Rev Genet. 2018;19(7):453-467.
- Karczewski KJ et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature. 2020;581:434-443.
