Medical School · Year 1 · Foundations · includes a quiz and discussion video
Lecture 11: Transcription and RNA Processing
Unit 1.1: Foundations of Medicine & Medical Sciences
Learning Objectives
By the end of this lecture, students will be able to:
- Describe the process of transcription including initiation, elongation, and termination
- Identify the key components of transcription machinery including RNA polymerase and transcription factors
- Explain the mechanisms of RNA splicing including the role of snRNPs and the spliceosome
- Describe post-transcriptional modifications: 5' capping, 3' polyadenylation, and editing
- Identify the different types of RNA and their functions
- Explain the regulation of gene expression at the transcriptional level
Overview of Transcription
The Central Dogma
The central dogma of molecular biology describes the flow of genetic information from DNA to RNA to protein. Transcription represents the first step in gene expression—the synthesis of an RNA copy from a DNA template. This RNA transcript may serve as a messenger (mRNA) carrying instructions for protein synthesis, or it may function directly as transfer RNA (tRNA), ribosomal RNA (rRNA), or various regulatory RNAs. Understanding transcription is fundamental to understanding how genes are expressed, how that expression is regulated, and how disruptions in this process contribute to disease.
The Transcription Process
During transcription, one strand of the DNA double helix serves as a template for RNA synthesis. RNA polymerase reads this template strand in the 3' to 5' direction while synthesizing the new RNA strand in the 5' to 3' direction. The resulting RNA is complementary and antiparallel to the template strand, meaning it has the same sequence as the non-template (coding) strand—except with uracil in place of thymine.
Comparing Replication and Transcription
Although both replication and transcription copy genetic information from DNA, they differ in fundamental ways. Replication produces DNA copies of both strands and requires primers to initiate synthesis; transcription produces RNA from only one strand and does not require primers. Replication uses deoxyribonucleotides containing deoxyribose sugar; transcription uses ribonucleotides containing ribose. In replication, adenine pairs with thymine; in transcription, adenine in the template pairs with uracil in the RNA product.
<image>Panel A: DNA replication showing a double helix with both strands being copied simultaneously, producing two complete double-stranded DNA molecules with semi-conservative distribution. Panel B: Transcription showing a DNA double helix with only the template strand (blue) being used to synthesize a single-stranded RNA molecule (red), with the non-template coding strand (gray) displaced. Panel C: Arrows indicating direction of synthesis (5' to 3') for both processes, highlighting antiparallel orientation of template reading versus new strand synthesis. Panel D: Comparison table summarizing key differences between replication and transcription: template (both strands vs. one strand), product (DNA vs. RNA), primer (required vs. not required), and sugar (deoxyribose vs. ribose).</image>
The Transcription Machinery
RNA Polymerases
The enzymes responsible for transcription differ between prokaryotes and eukaryotes. Prokaryotes employ a single RNA polymerase for all transcription. This core enzyme consists of five subunits (α₂ββ'ω) and associates with sigma factors that provide specificity for different promoter sequences. The sigma factor enables recognition of the promoter and is released after transcription initiates.
Eukaryotes have evolved three distinct nuclear RNA polymerases, each dedicated to transcribing different classes of genes. RNA polymerase I operates in the nucleolus, synthesizing the large ribosomal RNA precursor that will be processed into 28S, 18S, and 5.8S rRNAs. RNA polymerase II transcribes all protein-coding genes (producing messenger RNA), as well as small nuclear RNAs (snRNAs) and microRNAs. RNA polymerase III synthesizes transfer RNAs, 5S ribosomal RNA, and other small RNAs.
RNA polymerase II is particularly important for medical students to understand because it transcribes all protein-coding genes. This polymerase is uniquely sensitive to α-amanitin, a toxin from the death cap mushroom (Amanita phalloides), which explains why ingestion of this mushroom causes liver failure—hepatocytes have high transcriptional activity and are devastated when mRNA synthesis halts.
General Transcription Factors
Unlike bacterial RNA polymerase, eukaryotic RNA polymerase II cannot recognize promoters or initiate transcription on its own. It requires a set of general transcription factors (GTFs) that assemble at the promoter to form the pre-initiation complex (PIC). These factors are named with the convention TFII (transcription factor for polymerase II) followed by a letter.
The assembly process is sequential and ordered. TFIID arrives first, containing the TATA-binding protein (TBP) that recognizes the TATA box and associated factors (TAFs) that recognize other promoter elements. TFIIB then joins, helping to position the polymerase correctly and aiding in start site selection. TFIIF brings RNA polymerase II to the complex. TFIIE recruits TFIIH, which has two critical enzymatic activities: its helicase activity melts the DNA to create an open complex, and its kinase activity phosphorylates the carboxy-terminal domain (CTD) of RNA polymerase II, triggering the transition to active elongation.
<image>Panel A: Bare DNA with the TATA box highlighted approximately 25 bp upstream of the transcription start site (+1), followed by TFIID (saddle-shaped TBP with associated TAFs) binding and bending the DNA at the TATA box. Panel B: TFIIB joining the complex and positioning for start site selection. Panel C: TFIIF escorting RNA Polymerase II (large blue oval) to the promoter, followed by TFIIE recruitment. Panel D: The complete pre-initiation complex (PIC) with TFIIH bound, showing its helicase domain preparing to melt DNA, with all general transcription factors distinctly colored and labeled.</image>
Promoter Elements
The Core Promoter
The core promoter encompasses the minimal DNA sequences required for accurate transcription initiation. The most well-known element is the TATA box, located approximately 25 base pairs upstream of the transcription start site and having the consensus sequence TATAAAA. The TATA-binding protein (TBP) recognizes this sequence, inserting into the minor groove and bending the DNA dramatically—this distortion helps nucleate assembly of the remaining transcription machinery. Not all genes contain TATA boxes; many housekeeping genes use alternative core promoter elements.
The initiator element (Inr) spans the transcription start site itself and helps specify exactly where transcription begins. Some promoters also contain a downstream promoter element (DPE), which works in conjunction with or instead of the TATA box. These elements collectively ensure that RNA polymerase II starts at the correct position.
Proximal Promoter Elements
Immediately upstream of the core promoter, typically within 200 base pairs of the start site, lie proximal promoter elements that modulate transcription efficiency. The GC box (consensus GGGCGG) is bound by the transcription factor Sp1 and is common in housekeeping genes. The CAAT box (consensus CCAAT) is recognized by various transcription factors including NF-Y. These elements boost transcription above basal levels established by the core promoter alone.
Distal Regulatory Elements
Gene expression is fine-tuned by regulatory elements that can lie thousands of base pairs away from the promoter, either upstream or downstream, or even within introns. Enhancers are sequences that increase transcription when bound by activator proteins. Silencers decrease transcription when bound by repressor proteins. Insulators are boundary elements that block the influence of enhancers or silencers on neighboring genes.
The ability of these distant elements to influence transcription is explained by DNA looping—proteins bound at enhancers make direct contact with proteins at the promoter by looping out the intervening DNA. This three-dimensional organization of the genome is increasingly recognized as critical to proper gene regulation.
<image>Panel A: Core promoter region magnified showing the TATA box (~-25), initiator element (spanning +1), and downstream promoter element at the transcription start site. Panel B: Proximal promoter elements (GC box, CAAT box) within 200 bp upstream, with transcription factors Sp1 and NF-Y bound. Panel C: Enhancer element several kilobases upstream with activator proteins bound, and DNA looping (curved dashed line) bringing enhancer-bound activators into contact with the promoter machinery. Panel D: Silencer element downstream with a repressor protein bound and an insulator element with bound CTCF protein blocking the spread of enhancer effects, with color coding distinguishing activating (green), repressive (red), and neutral (blue) elements.</image>
The Transcription Cycle
Initiation
Transcription begins when the pre-initiation complex assembles at the promoter. Once the complete PIC forms, TFIIH uses its helicase activity to unwind approximately 17 base pairs of DNA around the transcription start site, creating an "open complex" where the template strand is accessible to the polymerase active site. TFIIH's kinase activity then phosphorylates serine residues at position 5 of the heptapeptide repeats that make up the carboxy-terminal domain (CTD) of RNA polymerase II.
This phosphorylation triggers synthesis of the first few nucleotides and promoter clearance—the polymerase breaks its contacts with the promoter and general transcription factors, committing to productive elongation. Some transcription factors (TBP, TFIID) remain at the promoter and facilitate reinitiation, allowing multiple polymerase molecules to transcribe the same gene.
Elongation
Once promoter clearance occurs, RNA polymerase II moves along the template strand, maintaining a transcription bubble of approximately 17 base pairs where the DNA is locally unwound. The growing RNA chain is about 8-9 nucleotides long within the polymerase before it emerges through an exit channel. Elongation proceeds at approximately 20-40 nucleotides per second, though this rate varies depending on sequence and chromatin context.
During elongation, the CTD becomes phosphorylated on serine 2 residues, which recruits different processing factors. Several elongation factors assist the polymerase. TFIIS helps the polymerase overcome pause sites and recover from backtracking events. P-TEFb (positive transcription elongation factor b) phosphorylates serine 2 of the CTD and overcomes promoter-proximal pausing. The Mediator complex serves as a bridge between gene-specific transcription factors bound at enhancers and the general transcription machinery.
<image>Panel A: Cross-sectional view of RNA Polymerase II as a large multi-subunit enzyme (light to dark blue gradient) with DNA entering as a double helix from the left. Panel B: The transcription bubble within the enzyme showing unwound single strands, with the template strand (dark gray) passing through the active site where ribonucleotides are added. Panel C: The nascent RNA strand (red), approximately 8 nucleotides long, emerging through an exit channel at the top of the polymerase. Panel D: The CTD extending from the polymerase with phosphate groups on serine residues (S5-P near promoter, S2-P during elongation), with arrows indicating direction of polymerase movement (5' to 3' on RNA).</image>
Termination
Termination of transcription by RNA polymerase II is less precisely defined than in prokaryotes and is closely linked to 3' end processing of the mRNA. Two models describe how the polymerase eventually disengages from the template.
In the torpedo model, the polymerase transcribes through the polyadenylation signal (AAUAAA), which triggers cleavage of the nascent RNA and addition of the poly(A) tail. However, the polymerase continues transcribing beyond this site. A 5' to 3' exonuclease (Xrn2 in humans, Rat1 in yeast) loads onto the free 5' end of the remaining RNA attached to the polymerase and degrades it. When the exonuclease catches up with the slowly transcribing polymerase, it triggers dissociation—like a torpedo sinking a ship.
In the allosteric model, passage through the polyadenylation signal causes conformational changes in the polymerase that reduce its processivity, eventually causing it to release from the template without requiring exonuclease activity. These models are not mutually exclusive and may both contribute to termination.
RNA Processing: mRNA Maturation
Co-transcriptional Processing
A remarkable feature of eukaryotic gene expression is that RNA processing occurs while transcription is still ongoing, not after it completes. The CTD of RNA polymerase II serves as a landing pad that recruits processing factors at appropriate times: capping enzymes arrive early (when the CTD is phosphorylated on Ser5), splicing factors come during elongation, and cleavage/polyadenylation factors are recruited later (when Ser2 phosphorylation predominates). This coupling ensures efficient and coordinated processing.
5' Capping
The 5' cap is added to nascent mRNA very early, when the transcript is only about 25 nucleotides long. This cap consists of a 7-methylguanosine connected to the first transcribed nucleotide through an unusual 5'-5' triphosphate linkage (rather than the standard 3'-5' phosphodiester bond). Additional methyl groups may be added to the first one or two nucleotides of the transcript.
The cap serves multiple essential functions. It protects the mRNA from degradation by 5' exonucleases. It is required for efficient translation initiation, as the ribosome recognizes the cap through initiation factors. It promotes splicing of the first intron. It serves as a nuclear export signal, helping direct the mature mRNA to the cytoplasm.
The capping process involves three enzymatic activities. First, an RNA triphosphatase removes the terminal phosphate from the nascent transcript. Second, guanylyltransferase adds GMP in the unusual 5'-5' linkage. Third, methyltransferases add methyl groups to the guanine (at N7) and potentially to the ribose of the first nucleotides.
<image>Panel A: Chemical structure of the 7-methylguanosine (m7G) with its methyl group highlighted in green attached to the N7 position, connected via a triphosphate bridge (three phosphate groups in orange). Panel B: The unusual 5'-5' orientation of the triphosphate linkage to the first nucleotide of the mRNA, with arrows emphasizing the backward orientation compared to standard 3'-5' bonds. Panel C: Inset comparing normal 3'-5' phosphodiester linkage to the 5'-5' cap linkage, with ribose sugars clearly drawn with numbered carbons. Panel D: Cap1 and cap2 structures showing 2'-O-methyl groups on the first and second nucleotides of the transcript.</image>
3' Polyadenylation
The 3' end of most eukaryotic mRNAs is not defined by transcription termination but rather by endonucleolytic cleavage followed by addition of a poly(A) tail. The cleavage site is typically located about 10-30 nucleotides downstream of the polyadenylation signal AAUAAA. A downstream GU-rich element helps position the cleavage complex.
The cleavage and polyadenylation specificity factor (CPSF) recognizes AAUAAA, while cleavage stimulation factor (CstF) binds the GU-rich element. Together with other factors, they cleave the pre-mRNA. Poly(A) polymerase then adds approximately 100-250 adenosine residues without using a template.
The poly(A) tail protects mRNA from 3' exonucleases, with deadenylation being a major pathway for mRNA decay. The tail also promotes nuclear export and enhances translation efficiency by enabling circularization of the mRNA (through interactions between poly(A)-binding proteins and cap-binding proteins). The tail length can be regulated to control mRNA stability—shortening the tail often triggers mRNA degradation.
RNA Splicing
The Need for Splicing
One of the most surprising discoveries about eukaryotic genes was that their coding sequences are interrupted by non-coding sequences. Exons (expressed sequences) contain the protein-coding information; introns (intervening sequences) lie between exons and must be precisely removed from the primary transcript. The average human gene contains 8-10 introns, though this varies enormously. The dystrophin gene, for example, spans 2.4 million base pairs—making it one of the largest human genes—yet its mRNA is only about 14,000 nucleotides because most of the gene consists of introns.
Splice Site Signals
Intron removal requires precise recognition of splice sites. The 5' splice site (donor site) almost invariably contains the dinucleotide GU at the intron's beginning. The 3' splice site (acceptor site) almost invariably contains AG at the intron's end. Within the intron, a branch point sequence contains a critical adenosine residue, typically located 18-40 nucleotides upstream of the 3' splice site. A polypyrimidine tract (a stretch of pyrimidines, mostly uracil) lies between the branch point and the 3' splice site.
These sequences are necessary but not sufficient for splicing. Exonic and intronic sequences also contain splicing enhancers and silencers that help define which sequences are true exons and which potential splice sites should be used.
<image>Panel A: Two exons (blue boxes) flanking an intron (gray line), with the 5' splice site showing the exon-intron boundary where the intron begins with the nearly invariant GU dinucleotide (red bold). Panel B: The 3' splice site where the intron ends with the nearly invariant AG dinucleotide (red bold), followed by the downstream exon. Panel C: The branch point adenosine (A in a circle) within the intron, approximately 30 nucleotides from the 3' end, with surrounding consensus sequence and the polypyrimidine tract (series of U's and C's) between the branch point and 3' splice site. Panel D: Arrows indicating the two cleavage sites that will form during splicing, with labeled brackets identifying each splice signal element.</image>
The Spliceosome
Splicing is catalyzed by the spliceosome, one of the most complex molecular machines in the cell. The spliceosome contains five small nuclear RNAs (snRNAs)—U1, U2, U4, U5, and U6—each complexed with proteins to form small nuclear ribonucleoproteins (snRNPs, pronounced "snurps"). Over 150 proteins participate in the spliceosome. Remarkably, the catalytic center of the spliceosome consists of RNA, making splicing fundamentally a ribozyme-catalyzed reaction—evidence for the ancient "RNA world" hypothesis.
Spliceosome assembly occurs stepwise on each intron. First, U1 snRNP recognizes and binds the 5' splice site through base pairing between U1 snRNA and the splice site sequence. U2 snRNP then binds the branch point, assisted by the U2 auxiliary factor (U2AF) that recognizes the polypyrimidine tract and 3' splice site. The U4/U6·U5 tri-snRNP joins to form the complete spliceosome. Extensive rearrangements then occur: U1 and U4 are displaced, and new base-pairing interactions form between U6 and the 5' splice site, and between U2 and U6. These rearrangements create the active site that catalyzes the splicing reaction.
The Splicing Mechanism
Splicing proceeds through two transesterification reactions—phosphodiester bonds are broken and formed without net energy input.
In the first reaction, the 2'-OH of the branch point adenosine attacks the phosphate at the 5' splice site. This creates a lariat structure in which the 5' end of the intron is attached to the branch point adenosine through a 2'-5' phosphodiester bond. Simultaneously, the first exon is released and its 3'-OH becomes free.
In the second reaction, the newly freed 3'-OH of the first exon attacks the phosphate at the 3' splice site. This joins the two exons together and releases the intron as a lariat, which is subsequently linearized and degraded.
<image>Panel A: Pre-mRNA substrate showing Exon 1 (blue box) connected to the intron (gray) connected to Exon 2 (blue box), with the branch point adenosine (A in a circle) marked within the intron. Panel B: First transesterification reaction with a curved arrow showing the 2'-OH of the branch point A attacking the 5' splice site phosphate, producing the freed Exon 1 with a 3'-OH group and the lariat intermediate. Panel C: Second transesterification reaction with the 3'-OH of Exon 1 attacking the 3' splice site phosphate (curved arrow), joining the two exons together. Panel D: Final products showing the ligated exons (Exon 1 - Exon 2 connected) and the released lariat intron (loop with a tail), with bonds broken and formed color-coded.</image>
Alternative Splicing
Expanding the Proteome
Alternative splicing is a powerful mechanism for generating protein diversity from a limited number of genes. By including or excluding different combinations of exons, a single gene can produce multiple distinct mRNA variants encoding different protein isoforms. This explains how humans, with only about 20,000 protein-coding genes, can produce a proteome estimated at over 100,000 distinct proteins. Approximately 95% of human multi-exon genes undergo alternative splicing.
Patterns of Alternative Splicing
Several patterns of alternative splicing are recognized. Exon skipping (or cassette exon) is the most common type, where an exon is either included or excluded from the mature mRNA. Mutually exclusive exons represent a choice between two or more exons, where exactly one is included. Alternative 5' or 3' splice sites allow selection of different splice junctions within an exon, changing its boundaries. Intron retention, where an intron remains in the mature mRNA, is less common in vertebrates but frequent in plants and fungi.
Regulation of Alternative Splicing
Which splice sites are used depends on regulatory proteins that bind to specific sequences in the pre-mRNA. SR proteins (serine/arginine-rich proteins) generally promote splicing and bind to exonic splicing enhancers (ESEs) and intronic splicing enhancers (ISEs). Heterogeneous nuclear ribonucleoproteins (hnRNPs) often antagonize SR proteins, typically repressing splicing when bound to exonic splicing silencers (ESSs) or intronic splicing silencers (ISSs).
The balance between activating and repressing factors determines splice site selection. This balance varies between cell types and developmental stages, producing tissue-specific isoforms of many proteins. Neuronal cells, for example, express specific splicing regulators that generate brain-specific variants of ion channels, receptors, and synaptic proteins.
<image>Panel A: Exon skipping pattern showing pre-mRNA with exons 1-2-3 producing either the full 1-2-3 transcript or 1-3 with exon 2 skipped, and mutually exclusive exons showing exons 1-2a-2b-3 producing either 1-2a-3 or 1-2b-3. Panel B: Alternative 5' splice site pattern where exon 2 has two possible 5' boundaries producing long or short versions, and alternative 3' splice site where exon 2 has two possible 3' boundaries. Panel C: Intron retention pattern where the intron between exons 1 and 2 is either removed or retained in the mature mRNA. Panel D: Regulatory elements with SR proteins binding exonic splicing enhancers to promote inclusion and hnRNPs binding exonic splicing silencers to promote exclusion, illustrating how the balance of factors determines splice site selection.</image>
Other RNA Processing
RNA Editing
RNA editing changes the nucleotide sequence of an mRNA after transcription, creating a product that differs from what the gene encodes. The two main types in mammals are A-to-I editing and C-to-U editing.
A-to-I editing is catalyzed by ADAR (adenosine deaminase acting on RNA) enzymes, which convert adenosine to inosine by deamination. Inosine is read as guanosine by the translation machinery, effectively creating an A-to-G change. This editing is particularly important in the nervous system. A dramatic example is the glutamate receptor B (GluR-B), where editing changes a glutamine codon (CAG) to an arginine codon (CIG read as CGG). This single amino acid change dramatically alters the receptor's calcium permeability—failure of this editing in motor neurons may contribute to ALS.
C-to-U editing is catalyzed by APOBEC enzymes. The classic example is apolipoprotein B mRNA in the intestine, where editing creates a stop codon that produces a truncated protein (apoB-48) with different properties than the liver version (apoB-100) made from unedited mRNA.
mRNA Export
Before translation can occur, mature mRNA must be transported from the nucleus to the cytoplasm through nuclear pore complexes. This export is coupled to successful processing—the cap-binding complex, spliced exon-junction complexes, and the poly(A) tail all contribute to marking an mRNA as mature and ready for export. The TREX (transcription/export) complex links transcription elongation and splicing to export.
Quality control mechanisms prevent export of improperly processed mRNAs. Transcripts that have not been properly capped, spliced, or polyadenylated are retained in the nucleus and degraded. This surveillance prevents potentially harmful aberrant proteins from being produced.
Types of RNA
The transcriptome encompasses many RNA types beyond messenger RNA. Ribosomal RNAs (rRNAs) form the structural and catalytic core of ribosomes; in humans, these include 28S, 18S, and 5.8S rRNAs (transcribed as a single precursor by RNA polymerase I) and 5S rRNA (transcribed by RNA polymerase III). Transfer RNAs (tRNAs), transcribed by RNA polymerase III, are approximately 75 nucleotides long and carry amino acids to the ribosome.
Small nuclear RNAs (snRNAs) are the RNA components of snRNPs and are essential for splicing; they range from about 100 to 200 nucleotides. Small nucleolar RNAs (snoRNAs) guide chemical modifications of rRNAs during ribosome biogenesis. MicroRNAs (miRNAs), only about 22 nucleotides long, regulate gene expression by base-pairing with target mRNAs, usually in the 3' untranslated region, leading to translation repression or mRNA degradation.
Long non-coding RNAs (lncRNAs), defined as longer than 200 nucleotides and not encoding proteins, have diverse functions in gene regulation, chromatin modification, and nuclear organization. Examples include XIST, which coats and silences one X chromosome in female mammals.
Regulation of Transcription
Transcription Factors
Gene expression is controlled largely at the level of transcription initiation, primarily through the action of sequence-specific transcription factors. These proteins bind to specific DNA sequences (enhancers, silencers, or proximal promoter elements) and either activate or repress transcription of nearby genes.
Transcription factors typically have modular structures with distinct domains. The DNA-binding domain recognizes specific sequences; common motifs include helix-turn-helix, zinc finger, leucine zipper, and helix-loop-helix structures. The activation or repression domain interacts with coactivators, corepressors, or the general transcription machinery to influence transcription rates.
The tumor suppressor p53 illustrates how transcription factors respond to cellular signals. In response to DNA damage or other stresses, p53 is stabilized and activates genes involved in cell cycle arrest, DNA repair, and apoptosis. NF-κB responds to inflammatory signals, activating genes for cytokines, adhesion molecules, and immune mediators. Steroid hormone receptors bind hormones like cortisol or estrogen in the cytoplasm, then translocate to the nucleus to activate hormone-responsive genes.
Chromatin and Epigenetic Regulation
In eukaryotes, DNA is packaged with histones into chromatin, and this packaging profoundly affects transcription. Genes in open, accessible chromatin (euchromatin) can be actively transcribed, while genes in condensed chromatin (heterochromatin) are generally silenced.
Histone modifications regulate this accessibility. Acetylation of histone lysine residues, catalyzed by histone acetyltransferases (HATs), neutralizes positive charges and loosens histone-DNA contacts, promoting transcription. Histone deacetylases (HDACs) remove acetyl groups, facilitating chromatin condensation and gene silencing. Histone methylation has context-dependent effects: methylation of histone H3 at lysine 4 (H3K4me3) marks active promoters, while methylation at lysine 9 or 27 (H3K9me3, H3K27me3) is associated with gene silencing.
DNA methylation, primarily at CpG dinucleotides, typically represses transcription. Methylated CpG islands in promoters are associated with stable gene silencing. DNA methylation patterns are heritable through cell division and are crucial for processes like X-chromosome inactivation and genomic imprinting.
<image>Panel A: Euchromatin state with nucleosomes (histones as colored balls with DNA wrapped around them) spaced apart, acetyl groups (Ac) attached to histone tails, and RNA Polymerase II actively transcribing. Panel B: Heterochromatin state with nucleosomes tightly packed and condensed, methyl groups (Me) on histone H3K9 and H3K27, and DNA methylation (mCpG) indicated on the DNA, with transcription silenced. Panel C: Transition between states showing how histone acetyltransferases (HATs) promote the open euchromatin configuration and histone deacetylases (HDACs) facilitate condensation to heterochromatin. Panel D: Key showing activating modifications (H3K4me3, acetylation) and repressive modifications (H3K9me3, H3K27me3, DNA methylation) with their associated chromatin states.</image>
Clinical Correlations
Splicing Mutations and Disease
Mutations affecting splicing are responsible for 15-50% of all disease-causing mutations, making this the most common mechanism by which point mutations cause disease. A mutation in a splice site sequence can cause exon skipping, intron retention, or activation of cryptic splice sites, producing abnormal proteins.
Beta-thalassemia, a common inherited hemoglobin disorder, is frequently caused by splicing mutations in the β-globin gene. Some mutations destroy normal splice sites, while others create new splice sites within introns or exons. The resulting abnormal mRNAs encode little or no functional β-globin, causing anemia.
Spinal muscular atrophy (SMA) provides an instructive example of splicing biology and therapy. Humans have two nearly identical genes, SMN1 and SMN2. SMA is caused by loss of SMN1, but SMN2 cannot compensate because a single nucleotide difference causes exon 7 to be frequently skipped, producing an unstable protein. The antisense oligonucleotide drug nusinersen (Spinraza) binds to SMN2 pre-mRNA and promotes exon 7 inclusion, restoring functional SMN protein production and dramatically improving outcomes in affected children.
α-Amanitin Poisoning
The death cap mushroom (Amanita phalloides) produces α-amanitin, a cyclic peptide that binds tightly to RNA polymerase II and blocks translocation during elongation. Ingestion causes delayed but severe hepatotoxicity because liver cells have high transcriptional demands and cannot synthesize essential proteins when mRNA production stops. Symptoms begin 6-12 hours after ingestion; without treatment, liver failure progresses over several days and is often fatal.
Therapeutic Targeting of Transcription
Understanding transcription has enabled development of targeted therapies. Besides antisense oligonucleotides for splicing modulation, researchers are developing inhibitors of specific transcription factors (such as those targeting NF-κB in inflammatory diseases) and drugs that modify chromatin state (HDAC inhibitors approved for certain cancers). The ability to manipulate gene expression at the transcriptional level opens vast therapeutic possibilities.
Summary
- Transcription produces RNA from DNA template using RNA polymerase
- Eukaryotic transcription requires assembly of pre-initiation complex with GTFs
- mRNA undergoes processing: 5' capping, splicing, 3' polyadenylation
- Splicing removes introns via spliceosome; alternative splicing increases diversity
- Transcription is regulated by transcription factors and chromatin state
- Splicing defects cause numerous human diseases
Key Terms
| Term | Definition |
|---|---|
| Promoter | DNA sequence where transcription machinery assembles |
| CTD | Carboxy-terminal domain of RNA Pol II; coordinates processing |
| Spliceosome | Complex that removes introns from pre-mRNA |
| Exon | Coding sequence retained in mature mRNA |
| Intron | Intervening sequence removed by splicing |
| Alternative splicing | Production of different mRNAs from same gene |
This content is subject to the MIT License. © 2024–2026 Hibbert School of Medicine.








