# Lecture 10: Transcription and RNA Processing

## Genetics

---

## Learning Objectives

By the end of this lecture, students will be able to:

1. Compare transcription in prokaryotes and eukaryotes
2. Describe the roles of RNA polymerases and transcription factors
3. Explain the three stages of transcription: initiation, elongation, and termination
4. Describe eukaryotic pre-mRNA processing: 5' capping, 3' polyadenylation, and splicing
5. Explain the mechanism of spliceosome-mediated splicing and alternative splicing
6. Discuss the significance of RNA editing and non-coding RNAs

---

## Lecture Content

### I. Overview of Transcription

Transcription is the synthesis of an RNA copy from a DNA template and represents the first step of the central dogma: DNA to RNA to Protein. The cell produces several types of RNA: **mRNA** (messenger RNA) encodes proteins; **rRNA** (ribosomal RNA) serves as a structural and catalytic component of ribosomes; **tRNA** (transfer RNA) acts as an adaptor molecule delivering amino acids during translation; **snRNA** (small nuclear RNA) forms part of the spliceosome; and various non-coding RNAs including miRNA, siRNA, and lncRNA perform regulatory functions.

Several key features distinguish transcription from replication. Transcription uses one DNA strand as the template (the template or antisense strand, read in the 3' to 5' direction). RNA is synthesized in the 5' to 3' direction and has the same sequence as the coding (sense) strand, except that uracil replaces thymine. Unlike DNA replication, transcription does not require a primer and uses ribonucleoside triphosphates (ATP, GTP, CTP, and UTP) as substrates.

### II. Prokaryotic Transcription

In prokaryotes, a single RNA polymerase (RNAP) transcribes all types of RNA. The **RNA polymerase holoenzyme** consists of the core enzyme (with subunit composition alpha2-beta-beta'-omega) plus a sigma factor. The sigma factor, most commonly sigma-70, recognizes promoter sequences and enables transcription initiation, while the core enzyme carries out elongation.

Prokaryotic promoters in *E. coli* contain several key elements: the **-10 region (Pribnow box)** with a consensus sequence of TATAAT located approximately 10 base pairs upstream of the transcription start site (+1), the **-35 region** with a consensus of TTGACA, and a critical spacing of approximately 17 base pairs between these two elements. Some strong promoters also possess an UP element upstream of the -35 region.

During **initiation**, the sigma factor binds the promoter to form a closed complex. The DNA then unwinds locally over approximately 12-15 base pairs to create an open complex (the transcription bubble), the first two ribonucleotides are joined, and RNA synthesis begins, usually with a purine. After 8-10 nucleotides have been synthesized, the sigma factor is released and the core enzyme proceeds to the elongation phase. During **elongation**, RNAP moves along the template in the 3' to 5' direction, synthesizing RNA in the 5' to 3' direction at a rate of approximately 40-80 nucleotides per second, with a transcription bubble of 12-15 base pairs and an RNA-DNA hybrid of 8-9 base pairs.

**Termination** occurs by one of two mechanisms. In **rho-independent (intrinsic) termination**, a GC-rich palindrome in the RNA forms a hairpin stem-loop followed by a poly-U stretch, which destabilizes the RNA-DNA hybrid and causes RNAP to dissociate. In **rho-dependent termination**, the Rho protein (a hexameric helicase) binds a rut (Rho utilization) site on the RNA, translocates in the 5' to 3' direction along the RNA, catches up to a paused RNAP, and unwinds the RNA-DNA hybrid to trigger termination.

<image>Panel A: Diagram of prokaryotic transcription showing the RNA polymerase holoenzyme bound to a promoter with -35 and -10 regions labeled, the open complex with transcription bubble, elongation with the growing RNA strand, and termination (both rho-independent with hairpin/poly-U and rho-dependent with Rho hexamer). Panel B: Detailed view of the E. coli promoter showing the -35 element, -10 element (Pribnow box), spacer region, +1 transcription start site, and the consensus sequences. Panel C: Structure of rho-independent terminator showing the GC-rich palindrome forming a hairpin stem-loop in the RNA, followed by the poly-U tail, with the destabilized RNA-DNA hybrid region.</image>

### III. Eukaryotic Transcription

Eukaryotes employ three RNA polymerases with distinct functions. **RNA Pol I** transcribes rRNA genes (28S, 18S, and 5.8S) in the nucleolus. **RNA Pol II** transcribes mRNA, most snRNAs, and miRNAs, making it the most relevant polymerase for gene expression. **RNA Pol III** transcribes tRNA, 5S rRNA, and other small RNAs. These polymerases can be distinguished by their sensitivity to alpha-amanitin: Pol II is highly sensitive, Pol I is resistant, and Pol III is moderately sensitive.

The **RNA Pol II promoter** contains several elements organized hierarchically. The **core promoter** includes the TATA box (consensus TATAAA, located approximately 25-30 base pairs upstream of the start site) recognized by TBP (TATA-binding protein, a subunit of TFIID), the Initiator element (Inr) at +1, and the downstream promoter element (DPE) found in TATA-less promoters. **Proximal promoter elements** (approximately 50-200 base pairs upstream) include the CAAT box and GC box, which are bound by specific transcription factors such as Sp1 and CTF/NF-1. **Enhancers** can be located thousands of base pairs away in either direction and function independently of orientation; they are bound by activator proteins and contribute to tissue-specific and developmental stage-specific gene expression. **Silencers** repress transcription through binding of repressor proteins.

**General (basal) transcription factors** are required for all Pol II transcription and assemble at the promoter in a specific order: TFIID (containing TBP) binds first, followed by TFIIA and TFIIB, then RNA Pol II with TFIIF, then TFIIE, and finally TFIIH. Together they form the **pre-initiation complex (PIC)**. TFIIH is particularly important because it has both helicase activity (to open the DNA) and kinase activity (to phosphorylate the Pol II CTD).

The **CTD (C-terminal domain)** of Pol II contains heptad repeats of the sequence Tyr-Ser-Pro-Thr-Ser-Pro-Ser, with 52 repeats in humans. Phosphorylation of Ser5 by TFIIH/CDK7 triggers promoter clearance and initiates capping, while phosphorylation of Ser2 by P-TEFb/CDK9 drives productive elongation. The phosphorylation state of the CTD coordinates co-transcriptional RNA processing.

### IV. 5' Capping

Capping occurs co-transcriptionally when the nascent mRNA is approximately 20-30 nucleotides long. The cap structure consists of a 7-methylguanosine linked via a 5'-5' triphosphate bridge to the first nucleotide of the mRNA (m7G(5')ppp(5')N). The process involves three enzymatic steps: RNA triphosphatase removes the terminal phosphate from the 5' end, guanylyltransferase adds a GMP in a 5'-5' linkage, and methyltransferase methylates the guanine at the N7 position. Additional methylations on the ribose of the first and second nucleotides create Cap 1 and Cap 2 structures, respectively.

The 5' cap serves multiple critical functions: it protects mRNA from degradation by 5' exonucleases, is required for efficient splicing of the first intron, promotes nuclear export, and is essential for efficient translation through recognition by the initiation factor eIF4E.

### V. 3' Polyadenylation

Nearly all eukaryotic mRNAs, with the exception of histone mRNAs, receive a poly(A) tail at their 3' end. The **polyadenylation signal** is the hexamer AAUAAA, located approximately 10-30 nucleotides upstream of the cleavage site, along with a downstream GU-rich or U-rich element. **CPSF** (Cleavage and Polyadenylation Specificity Factor) binds the AAUAAA signal, and **CstF** (Cleavage Stimulation Factor) binds the downstream element. Endonucleolytic cleavage occurs 10-30 nucleotides downstream of AAUAAA, and then **Poly(A) polymerase (PAP)** adds approximately 200-250 adenine residues without a template. **PABPN1** (nuclear poly(A)-binding protein) binds the growing tail and regulates its length.

The poly(A) tail protects mRNA from 3' exonuclease degradation, is required for nuclear export, promotes translation by synergizing with the 5' cap through eIF4G bridging, and regulates mRNA stability since poly(A) tail shortening triggers mRNA degradation. Polyadenylation is coupled with transcription termination: after cleavage, a 5' to 3' exonuclease (Rat1/Xrn2) degrades the downstream RNA still associated with Pol II and triggers Pol II release through the "torpedo model."

### VI. Pre-mRNA Splicing

Eukaryotic genes contain **exons** (expressed sequences) and **introns** (intervening sequences). The average human gene has approximately 8-9 exons, and introns can be extremely large, with some exceeding 100 kb. Intron removal and exon joining is catalyzed by the **spliceosome**, and the process depends on conserved sequences at intron boundaries: the **5' splice site (donor)** contains an almost invariant GU, the **3' splice site (acceptor)** contains an almost invariant AG, the **branch point** is an adenine residue located approximately 18-40 nucleotides upstream of the 3' splice site (consensus YNYURAY in mammals), and a **polypyrimidine tract** lies between the branch point and the 3' splice site.

The spliceosome is a large ribonucleoprotein complex composed of five snRNPs (U1, U2, U4, U5, and U6) plus many associated proteins. Splicing proceeds through two transesterification reactions. In Step 1, the 2'-OH of the branch point adenine attacks the phosphodiester bond at the 5' splice site, forming a lariat intermediate in which the intron is looped with a 2'-5' bond at the branch point; exon 1 is released. In Step 2, the free 3'-OH of exon 1 attacks the phosphodiester bond at the 3' splice site, joining the exons and releasing the lariat intron for degradation.

Spliceosome assembly proceeds stepwise: U1 snRNP binds the 5' splice site to form the E complex; U2 snRNP binds the branch point to form the A complex; the U4/U6.U5 tri-snRNP joins to form the B complex; then rearrangements occur in which U1 and U4 leave and U6 replaces U1 at the 5' splice site, creating the catalytically active C complex. The U2 and U6 snRNAs together form the catalytic center, functioning as a ribozyme.

<image>Panel A: Diagram of the two-step transesterification mechanism of pre-mRNA splicing, showing: the pre-mRNA with exon 1, intron (with 5' GU, branch point A, polypyrimidine tract, 3' AG), and exon 2; Step 1 forming the lariat intermediate; Step 2 joining the exons and releasing the lariat. Panel B: Spliceosome assembly pathway showing the stepwise binding of U1, U2, U4/U6.U5 snRNPs to form E, A, B, and C complexes, with catalytic activation indicated. Panel C: Diagram of co-transcriptional mRNA processing showing RNA Pol II with phosphorylated CTD, and the three processing events occurring simultaneously: 5' capping at the emerging 5' end, splicing of introns as they emerge, and 3' polyadenylation at the end.</image>

### VII. Alternative Splicing

A single gene can produce multiple mRNA variants and thus multiple protein isoforms by including or excluding specific exons through alternative splicing. The major types include **exon skipping (cassette exon)**, the most common form in which an exon is either included or excluded; **alternative 5' splice site** and **alternative 3' splice site** usage, which employ different splice sites within the same intron; **intron retention**, in which an intron remains in the mature mRNA; and **mutually exclusive exons**, where one of two exons is included but never both.

It is estimated that over 95% of human multi-exon genes undergo alternative splicing, making this a major source of proteomic diversity. The *DSCAM* gene in *Drosophila* can produce over 38,000 different mRNA isoforms. The calcitonin/CGRP gene produces calcitonin in thyroid cells but CGRP in neurons through tissue-specific splicing. Immunoglobulin heavy chain genes produce membrane-bound versus secreted antibody forms by the same mechanism.

Alternative splicing is regulated by **SR proteins** (serine/arginine-rich proteins), which bind exonic splicing enhancers (ESEs) and promote exon inclusion, and by **hnRNPs**, which bind exonic or intronic splicing silencers (ESSs, ISSs) and promote exon skipping. Tissue-specific splicing factors such as NOVA, RBFOX, and MBNL add further layers of regulation.

### VIII. RNA Editing

RNA editing is the post-transcriptional alteration of the nucleotide sequence of RNA, producing changes not encoded in the DNA template. **Adenosine-to-inosine (A-to-I) editing** is catalyzed by ADAR enzymes (Adenosine Deaminases Acting on RNA). Because inosine is read as guanosine by the translation machinery, this editing can change codons, alter splicing patterns, or modify miRNA targets. A notable example is the GluR-B glutamate receptor subunit, where editing at the Q/R site changes a glutamine codon to an arginine codon, fundamentally altering the ion channel's permeability properties.

**Cytidine-to-uridine (C-to-U) editing** is catalyzed by APOBEC enzymes. The most well-known example involves apolipoprotein B mRNA: in intestinal cells, editing creates a premature stop codon that produces a shorter protein (apoB-48), while the unedited version in the liver produces the full-length protein (apoB-100). This single editing event thus generates two functionally distinct proteins from the same gene.

---
