# Lecture 16: PCR, Cloning, and Sequencing

## Genetics

---

## Learning Objectives

By the end of this lecture, students will be able to:

1. Explain the principle and steps of the polymerase chain reaction (PCR)
2. Describe variations of PCR (RT-PCR, qPCR, multiplex PCR) and their applications
3. Explain the Sanger dideoxy sequencing method
4. Describe expression cloning and site-directed mutagenesis
5. Explain the principles of molecular cloning strategies beyond restriction enzyme-based methods
6. Discuss applications of PCR and sequencing in research and clinical diagnostics

---

## Lecture Content

### I. Polymerase Chain Reaction (PCR)

The polymerase chain reaction, developed by Kary Mullis in 1983 (Nobel Prize in 1993), enables the in vitro amplification of a specific DNA sequence, generating millions of copies from a minute amount of starting template. The reaction requires template DNA (even in trace quantities), two oligonucleotide primers (forward and reverse, each approximately 18-25 nucleotides long) flanking the target region, a thermostable DNA polymerase (most commonly Taq polymerase from *Thermus aquaticus*, with an optimal temperature of 72 degrees Celsius), dNTPs (dATP, dGTP, dCTP, and dTTP), and a buffer containing MgCl2 (since Mg2+ is a cofactor for the polymerase).

Each PCR cycle consists of three temperature-controlled steps, repeated 25-40 times. **Denaturation** (approximately 94-98 degrees Celsius for 15-30 seconds) separates the double-stranded template DNA into single strands. **Annealing** (approximately 50-65 degrees Celsius for 15-30 seconds) allows the primers to bind to their complementary sequences on the template strands, with the annealing temperature depending on the primer melting temperature (Tm). **Extension** (approximately 72 degrees Celsius for 30 seconds to several minutes) allows Taq polymerase to extend the primers in the 5' to 3' direction, synthesizing new complementary strands.

Amplification is exponential: after n cycles, there are theoretically 2^n copies of the target sequence. After 30 cycles, this yields approximately 10^9 copies, a billion-fold amplification. In practice, efficiency decreases in later cycles due to reagent depletion and product inhibition. Taq polymerase is heat-stable (surviving the denaturation step) but lacks 3' to 5' proofreading exonuclease activity, giving it an error rate of approximately 1 in 10^4 nucleotides. For high-fidelity applications, proofreading polymerases such as Pfu or Phusion are preferred.

<image>Panel A: Diagram of three PCR cycles showing template DNA, primer annealing, extension, and exponential amplification — cycle 1 produces 2 copies, cycle 2 produces 4, cycle 3 produces 8, with the target-length amplicons highlighted from cycle 3 onward. Panel B: Temperature profile graph of a PCR reaction showing the three temperature steps (denaturation at 94-98C, annealing at 50-65C, extension at 72C) repeated over multiple cycles, with time on the x-axis and temperature on the y-axis. Panel C: Agarose gel image diagram showing PCR products — lane 1: DNA ladder, lane 2: PCR product (bright band at expected size), lane 3: negative control (no template, no band), lane 4: positive control.</image>

### II. PCR Variations

**Reverse Transcription PCR (RT-PCR)** detects and amplifies RNA sequences. The first step uses reverse transcriptase to convert mRNA into cDNA, and the second step uses standard PCR to amplify the cDNA. This technique is used to detect gene expression and viral RNA, as in the SARS-CoV-2 diagnostic test.

**Quantitative real-time PCR (qPCR / RT-qPCR)** monitors amplification in real time using fluorescent reporters. **SYBR Green** is a non-specific fluorescent dye that binds double-stranded DNA. **TaqMan probes** are sequence-specific hydrolysis probes carrying a 5' fluorophore and a 3' quencher; during extension, Taq's 5' to 3' exonuclease activity cleaves the probe, releasing the fluorophore and generating fluorescence. The **Ct (cycle threshold) value** is the cycle number at which fluorescence crosses a set threshold; a lower Ct indicates more initial template. qPCR can quantify absolute copy number using a standard curve or relative expression using the delta-delta Ct method. Applications include gene expression analysis, viral load quantification, and copy number variation detection.

**Multiplex PCR** uses multiple primer pairs in a single reaction to amplify several targets simultaneously, commonly employed in forensics, pathogen detection panels, and genetic testing. **Nested PCR** uses two rounds of amplification, with the second round using primers internal to the first product to increase specificity. **Inverse PCR** amplifies unknown sequences flanking a known sequence. **Digital PCR (dPCR)** partitions the sample into thousands of micro-reactions for absolute quantification without a standard curve.

### III. Sanger Dideoxy Sequencing

Developed by Frederick Sanger in 1977 (Nobel Prize in 1980), Sanger sequencing operates on the principle of chain termination using dideoxynucleotides (ddNTPs). Because ddNTPs lack the 3'-OH group, no further phosphodiester bonds can form once a ddNTP is incorporated, and the growing chain terminates.

The method begins with single-stranded or denatured template DNA to which a primer is annealed. DNA polymerase extends the primer in the presence of normal dNTPs and a small proportion of fluorescently labeled ddNTPs (ddATP, ddGTP, ddCTP, and ddTTP, each carrying a different fluorophore). The polymerase occasionally incorporates a ddNTP, causing random termination at every possible position. This produces a population of fragments of every possible length, each terminated with a fluorescent ddNTP indicating the identity of the terminal base. The fragments are separated by capillary electrophoresis at single-nucleotide resolution, and a laser excites the fluorophores as fragments pass the detector, generating a chromatogram.

Sanger sequencing produces read lengths of approximately 700-1,000 base pairs per run with accuracy exceeding 99.999% for high-quality reads. It is used for validating cloned inserts, confirming mutations, and targeted sequencing of specific genes. Although largely replaced by next-generation sequencing for large-scale projects, Sanger sequencing remains the gold standard for validation.

### IV. Molecular Cloning Strategies

**TA cloning** exploits the fact that Taq polymerase adds a single 3' adenine overhang to PCR products. Vectors linearized with single 3' thymine overhangs allow A-T base pairing to facilitate ligation without restriction enzyme digestion, making this a quick and easy method for cloning PCR products.

**Gateway cloning (recombinational cloning)** uses site-specific recombination at att sites instead of restriction enzymes and ligase. A BP reaction recombines an insert flanked by attB sites with a donor vector carrying attP sites to produce an entry clone with attL sites. An LR reaction then transfers the insert from the entry clone to a destination vector, producing an expression clone. Advantages include high efficiency, directional cloning, and easy shuttling between multiple destination vectors.

**Gibson Assembly** joins multiple DNA fragments with overlapping ends in a single isothermal reaction at 50 degrees Celsius. Three enzymes work together: a 5' exonuclease creates overhangs, a polymerase fills gaps, and a ligase seals nicks. No restriction enzymes are needed, the assembly is seamless, and 2-6 or more fragments can be joined simultaneously. **Golden Gate Assembly** uses Type IIS restriction enzymes, which cut outside their recognition site, to generate unique 4-nucleotide overhangs for directional, scarless assembly of multiple fragments. **In-Fusion cloning** uses PCR products with 15 base pair overlaps to the vector for recombination-based joining.

### V. Site-Directed Mutagenesis

Site-directed mutagenesis introduces specific, predetermined mutations into a cloned gene. The most common approach, the **QuikChange method**, involves designing complementary primers containing the desired mutation, using these primers to amplify the entire plasmid by PCR, digesting the parental (methylated) template DNA with DpnI (which cuts only methylated DNA), and transforming the mutated (unmethylated) plasmid into bacteria. Applications include studying the function of specific amino acids in a protein, creating disease-associated mutations in model systems, engineering enzymes with altered properties, and altering regulatory elements to study gene expression.

### VI. Expression Systems and Protein Production

**Expression vectors** are designed to produce recombinant protein from a cloned gene and include a strong promoter (such as the T7 promoter for *E. coli* or the CMV promoter for mammalian cells), a ribosome binding site, an affinity tag for purification (His-tag, GST-tag, FLAG-tag, or MBP-tag), and a multiple cloning site downstream of the promoter.

Different **expression hosts** suit different needs. *E. coli* is fast, inexpensive, and produces high yields but lacks post-translational modifications and may form inclusion bodies. Yeast (*S. cerevisiae* or *Pichia pastoris*) provides some post-translational modifications and secretion capability. Insect cells using the baculovirus system offer eukaryotic modifications with high yield. Mammalian cells (CHO or HEK293) provide the full complement of post-translational modifications and are essential for producing therapeutic proteins. **Protein purification** typically begins with affinity chromatography (such as Ni-NTA for His-tagged proteins) followed by ion exchange and size exclusion chromatography.

<image>Panel A: Sanger sequencing workflow diagram showing: template + primer + dNTPs + fluorescently labeled ddNTPs, extension and random termination, capillary electrophoresis separation, laser detection of fluorescent labels, and the resulting chromatogram with colored peaks (A, T, G, C) and the deduced sequence. Panel B: Gibson Assembly diagram showing three DNA fragments with 15-40 bp overlapping ends, the three-enzyme reaction (exonuclease creating overhangs, polymerase filling gaps, ligase sealing), and the final assembled construct. Panel C: qPCR amplification plot showing fluorescence (y-axis) vs. cycle number (x-axis) for samples with high, medium, and low starting template amounts, with the threshold line and Ct values marked for each curve.</image>

---
