# Lecture 26: Introduction to Biotechnology and Recombinant DNA

## General Biology I — Molecular & Cellular

---

## Learning Objectives

By the end of this lecture, students will be able to:

1. Explain the principles of recombinant DNA technology, including the roles of restriction enzymes, vectors, and DNA ligase
2. Describe cloning strategies for producing recombinant proteins and genomic/cDNA libraries
3. Explain the polymerase chain reaction (PCR) and its applications
4. Describe gel electrophoresis, Southern blotting, and DNA sequencing methods
5. Explain the principles and applications of CRISPR-Cas9 gene editing

---

## Lecture Content

### I. Overview of Biotechnology

**Biotechnology** — the use of biological systems, organisms, or their components to develop products and technologies. Modern biotechnology is founded on **recombinant DNA technology** — the ability to combine DNA from different sources into a single molecule. Key enabling discoveries: **Restriction enzymes** — discovered in the late 1960s-1970s (Werner Arber, Daniel Nathans, Hamilton Smith — Nobel Prize 1978) **DNA ligase** — enzyme that joins DNA fragments. **Vectors** (plasmids, phages) — vehicles for carrying foreign DNA into host cells. **DNA polymerase and reverse transcriptase** — for copying DNA and converting RNA to DNA.

### II. Tools of Recombinant DNA Technology

#### Restriction Enzymes (Restriction Endonucleases)

Bacterial enzymes that cut DNA at specific **recognition sequences** (restriction sites) Part of the bacterial **restriction-modification system** — a defense against bacteriophages. The bacterium methylates its own DNA at restriction sites (preventing self-cleavage) Foreign (unmethylated) phage DNA is recognized and cleaved. **Type II restriction enzymes** — most commonly used in the lab; cut within or near the recognition site. Recognition sites are typically **4-8 bp palindromes** (the sequence reads the same on both strands in the 5' to 3' direction) Two types of cuts: **Sticky ends (cohesive ends)** — staggered cuts that produce single-stranded overhangs. Example: **EcoRI** recognizes 5'-GAATTC-3' and cuts between G and A on both strands, producing 4-nt 5' overhangs. Sticky ends from the same enzyme are **complementary** — they can base-pair with each other, facilitating ligation of fragments from different sources. **Blunt ends** — cuts straight across both strands. Example: **SmaI** recognizes 5'-CCCGGG-3' and cuts in the center. Blunt ends can be ligated but less efficiently than sticky ends. The number of fragments produced depends on the number of restriction sites in the DNA. **Restriction mapping** — determining the positions of restriction sites in a DNA molecule by analyzing fragment sizes.

#### Vectors

A **vector** is a DNA molecule used to carry foreign DNA into a host cell and ensure its replication. Properties of a good vector: **Origin of replication (ori)** — allows autonomous replication in the host. **Selectable marker** — allows identification of cells that have taken up the vector (e.g., antibiotic resistance gene) **Multiple cloning site (MCS / polylinker)** — a short region containing recognition sites for many different restriction enzymes; the site where foreign DNA is inserted. Types of vectors: **Plasmids** — small, circular, extrachromosomal DNA; carry inserts of ~1-10 kb. Example: **pBR322**, **pUC19**. **Bacteriophage lambda** — modified phage; inserts of ~10-20 kb. **Cosmids** — hybrid plasmid-phage vectors; inserts of ~35-45 kb. **BACs (Bacterial Artificial Chromosomes)** — inserts of ~100-300 kb; used for large genomic projects. **YACs (Yeast Artificial Chromosomes)** — inserts of ~200 kb to >1 Mb; contain telomeres, centromere, and ori for replication in yeast. **Expression vectors** — contain strong promoters, ribosome binding sites, and other regulatory elements to drive **high-level expression** of the cloned gene as protein.

#### DNA Ligase

**T4 DNA ligase** — joins DNA fragments by catalyzing the formation of **phosphodiester bonds** between the 3'-OH and 5'-phosphate of adjacent nucleotides. Requires ATP as a cofactor. Most efficient with complementary sticky ends; can also join blunt ends (at higher concentrations or with additives like PEG).

<image>A step-by-step diagram of molecular cloning using recombinant DNA technology. Step 1: A gene of interest is cut from source DNA using EcoRI restriction enzyme, producing fragments with AATT sticky ends. The plasmid vector (showing an origin of replication, an ampicillin resistance gene, and a multiple cloning site containing an EcoRI site within a lacZ gene) is also cut with EcoRI. Step 2: The gene of interest and the linearized plasmid are mixed together. Complementary sticky ends base-pair. T4 DNA ligase seals the phosphodiester bonds, creating a recombinant plasmid with the gene of interest inserted into the MCS. Step 3: The recombinant plasmid is introduced into E. coli by transformation (heat shock or electroporation). Step 4: Bacteria are plated on ampicillin plates with X-gal. Colonies with the recombinant plasmid are ampicillin-resistant and white (lacZ disrupted); colonies with re-ligated empty vector are ampicillin-resistant and blue (lacZ intact). Non-transformed bacteria do not grow.</image>

### III. Introducing DNA into Host Cells

**Transformation** — uptake of free DNA by bacteria. **Heat shock** — CaCl2 treatment makes cells competent; brief heat pulse (42 degrees C) triggers DNA uptake. **Electroporation** — brief electrical pulse creates temporary pores in the membrane. **Transfection** — introduction of DNA into eukaryotic cells. Chemical methods (calcium phosphate, liposomes/lipofection) Electroporation. Viral vectors (transduction) Microinjection. **Screening and selection:**. **Antibiotic selection** — only cells with the vector (carrying resistance gene) survive. **Blue-white screening** — the MCS is located within the *lacZ* gene in the vector. Insert disrupts lacZ -> white colonies (no functional beta-galactosidase) No insert -> blue colonies (functional lacZ cleaves X-gal to produce blue product) **Colony hybridization** — use a labeled probe to identify colonies carrying the gene of interest. **PCR screening** — amplify the insert directly from colonies.

### IV. DNA Libraries

#### Genomic Library

A collection of clones that collectively represent the **entire genome** of an organism. Made by: Extracting total genomic DNA. Cutting with restriction enzymes (partial digestion to generate overlapping fragments) Ligating fragments into vectors. Transforming bacteria — each colony carries one fragment. Contains **all sequences** — exons, introns, regulatory regions, repetitive elements. Used for studying gene organization, regulatory elements, and genome structure.

#### cDNA Library

A collection of clones representing only the **expressed genes** (mRNA) of a particular tissue or cell type at a specific time. Made by: Extracting **mRNA** from a tissue (selected using oligo-dT beads that bind poly-A tails) Using **reverse transcriptase** to synthesize first-strand **cDNA** (complementary DNA) Synthesizing the second strand (using DNA polymerase, RNase H) Ligating double-stranded cDNA into vectors; cDNA clones lack introns — they represent the processed mRNA. Tissue-specific — different tissues yield different cDNA libraries. Useful for: Expressing eukaryotic genes in bacteria (bacteria cannot splice introns) Studying gene expression profiles.

### V. Polymerase Chain Reaction (PCR)

Invented by **Kary Mullis (1983)** — Nobel Prize in 1993. An in vitro method for **amplifying a specific DNA sequence** exponentially. Requirements: **Template DNA** — the DNA containing the target sequence (can be a very small amount) **Two primers** — short, synthetic oligonucleotides (~18-25 nt) complementary to sequences flanking the target region (one for each strand) **Taq DNA polymerase** (from *Thermus aquaticus*) — a thermostable DNA polymerase that withstands repeated heating. **dNTPs** (dATP, dGTP, dCTP, dTTP) Buffer with MgCl2.

**Three-step cycle (repeated 25-40 times):**. **Denaturation** (~94-95 degrees C, 30 sec) — the double-stranded DNA is heated to separate the strands. **Annealing** (~50-65 degrees C, 30 sec) — temperature is lowered to allow primers to bind (anneal) to their complementary sequences on the template strands. **Extension** (~72 degrees C, 1-2 min) — Taq polymerase extends the primers, synthesizing new DNA strands complementary to the template.

**Exponential amplification** — each cycle doubles the number of target molecules: after *n* cycles, ~2^n copies; 30 cycles -> ~10^9 copies from a single starting molecule.

**Applications:**. Forensic DNA analysis (DNA fingerprinting from crime scene samples) Diagnostic detection of pathogens (e.g., SARS-CoV-2 testing by RT-qPCR) Cloning of specific genes. Prenatal genetic testing. Evolutionary studies using ancient DNA. Site-directed mutagenesis. Genotyping and SNP analysis.

**Variants:**. **RT-PCR (reverse transcription PCR)** — mRNA is first converted to cDNA by reverse transcriptase, then amplified; detects gene expression. **qPCR (quantitative / real-time PCR)** — uses fluorescent reporters to measure DNA amplification in real time; quantifies the amount of target DNA or mRNA. **Nested PCR** — two rounds of PCR with different primer sets for increased specificity.

<image>A diagram of three cycles of PCR. Cycle 1: The double-stranded template DNA is denatured (strands separate). Two primers (forward and reverse, shown as short arrows) anneal to complementary sequences flanking the target region. Taq polymerase extends both primers, producing two double-stranded copies. Cycle 2: The two copies are denatured (four single strands), primers anneal, and extension produces four copies. Cycle 3: Denaturation, annealing, and extension produce eight copies. A graph inset shows the exponential increase: 1 copy becomes 2, 4, 8, 16... 2^n after n cycles. The target-length product (defined by the two primer positions) first appears in cycle 3 and rapidly dominates in subsequent cycles.</image>

### VI. Gel Electrophoresis and Analysis

#### Agarose Gel Electrophoresis

Separates DNA fragments by **size** — smaller fragments migrate faster through the gel matrix. DNA is negatively charged (phosphate backbone) — migrates toward the **positive electrode (anode)** in an electric field. After electrophoresis, DNA is visualized with **ethidium bromide** (intercalates into DNA and fluoresces under UV) or safer alternatives (SYBR Safe) **DNA ladder (molecular weight marker)** — run alongside samples to estimate fragment sizes. Can also separate RNA and proteins (SDS-PAGE for proteins).

#### Southern Blotting

Developed by **Edwin Southern (1975)** — detects a specific DNA sequence within a mixture. Steps: DNA is digested with restriction enzymes. Fragments are separated by gel electrophoresis. DNA is transferred (**blotted**) from the gel to a nylon or nitrocellulose membrane. The membrane is incubated with a **labeled probe** (a single-stranded DNA or RNA complementary to the target sequence; labeled with radioactivity or fluorescence) The probe hybridizes to complementary sequences on the membrane. Excess probe is washed off; the bound probe is detected by autoradiography or fluorescence imaging. Applications: detecting gene rearrangements, RFLP analysis, confirming gene knockouts. Related techniques: **Northern blot** — same principle but for detecting specific **RNA** molecules. **Western blot** — detects specific **proteins** using antibodies.

### VII. DNA Sequencing

#### Sanger Sequencing (Chain-Termination Method)

Developed by **Frederick Sanger (1977)** — Nobel Prize (his second) in 1980. Principle: DNA synthesis is terminated at specific bases using **dideoxynucleotides (ddNTPs)** — lacking the 3'-OH group necessary for chain elongation. Modern Sanger sequencing (automated): A single-stranded DNA template and a primer are mixed with DNA polymerase, normal dNTPs, and a small amount of fluorescently labeled ddNTPs (each of the four ddNTPs has a different colored fluorescent tag) During synthesis, ddNTPs are randomly incorporated, terminating the chain at every possible position. The resulting fragments of all different lengths are separated by **capillary electrophoresis**. A laser detector reads the fluorescent color of each fragment as it passes — generating a chromatogram. Read lengths of ~700-1000 bases. Still used for targeted sequencing, validation, and small-scale projects.

#### Next-Generation Sequencing (NGS)

Massively parallel sequencing technologies that can sequence millions of DNA fragments simultaneously. Key platforms: **Illumina** (sequencing by synthesis), **Ion Torrent**, **PacBio** (long reads), **Oxford Nanopore** (real-time, long reads) Enables **whole-genome sequencing**, transcriptomics (RNA-seq), epigenomics, and metagenomics. Cost of sequencing a human genome has dropped from ~$3 billion (Human Genome Project, 2003) to under ~$200.

### VIII. CRISPR-Cas9 Gene Editing

**CRISPR** — Clustered Regularly Interspaced Short Palindromic Repeats. Originally discovered as a **bacterial adaptive immune system** against bacteriophages. Adapted as a revolutionary **gene-editing tool** by **Jennifer Doudna and Emmanuelle Charpentier (2012)** — Nobel Prize in Chemistry, 2020.

#### Mechanism

**Two components:**. **Cas9** — an endonuclease (from *Streptococcus pyogenes*) that makes a double-strand break in DNA. **Guide RNA (sgRNA)** — a synthetic single-guide RNA (~20 nt) complementary to the target DNA sequence; directs Cas9 to the correct genomic location. **PAM (protospacer adjacent motif)** — a short sequence (NGG for SpCas9) that must be present immediately downstream of the target site on the non-target strand; Cas9 will not cut without it. Cas9 unwinds the DNA, the guide RNA base-pairs with the target strand, and Cas9 cleaves **both strands** 3 bp upstream of the PAM.

#### Repair Outcomes

The cell repairs the double-strand break by one of two pathways: **NHEJ (Non-Homologous End Joining)** — error-prone; introduces small insertions or deletions (**indels**) that often disrupt the reading frame -> **gene knockout**. **HDR (Homology-Directed Repair)** — if a **donor template** with the desired sequence (flanked by homology arms) is provided, the cell can incorporate the new sequence precisely -> **gene correction, insertion of new sequences, or specific mutations**.

#### Applications

Basic research — studying gene function through knockouts and knock-ins. **Gene therapy** — correcting disease-causing mutations (e.g., sickle cell disease, beta-thalassemia) **Casgevy (exagamglogene autotemcel)** — first CRISPR-based therapy approved (2023) for sickle cell disease and transfusion-dependent beta-thalassemia. Agriculture — developing disease-resistant, drought-tolerant, or more nutritious crops without introducing foreign DNA (non-transgenic editing) **Gene drives** — using CRISPR to spread a desired gene through a wild population (e.g., to reduce malaria transmission by mosquitoes) Diagnostics — CRISPR-based detection of viral nucleic acids (SHERLOCK, DETECTR).

<image>A diagram of the CRISPR-Cas9 gene editing mechanism. At the top, the Cas9 protein (shown as a large bilobed structure) is complexed with a single-guide RNA (sgRNA). The sgRNA has a 20-nucleotide targeting sequence at its 5' end complementary to the genomic target, and a scaffold region that binds Cas9. The complex scans the DNA for a PAM sequence (NGG, highlighted). Upon finding the PAM and a complementary target, the sgRNA base-pairs with the target strand, forming an R-loop structure, and Cas9 cleaves both DNA strands (indicated by scissors icons). Below, two repair outcomes are shown. Left path (NHEJ): The broken ends are joined imprecisely, introducing insertions or deletions (indels) that disrupt the gene — resulting in a knockout. Right path (HDR): A donor DNA template with homology arms flanking a desired sequence is provided. The cell uses this template to repair the break, precisely inserting the new sequence — resulting in a gene correction or knock-in.</image>

### IX. Applications of Recombinant DNA Technology

**Recombinant protein production:**. Human **insulin** — the first recombinant pharmaceutical (1982); the human insulin gene is expressed in *E. coli* or yeast. **Human growth hormone**, **erythropoietin (EPO)**, **clotting factors (Factor VIII)**, **tissue plasminogen activator (tPA)**. **Transgenic organisms:**. **Transgenic bacteria** — engineered to produce pharmaceuticals, biofuels, or to degrade environmental pollutants (bioremediation) **Transgenic plants** — Bt crops (express insecticidal protein from *Bacillus thuringiensis*), Golden Rice (engineered to produce beta-carotene) **Transgenic animals** — used in research (knockout mice), pharmaceutical production (pharming) **DNA fingerprinting (DNA profiling):**. Analysis of **short tandem repeats (STRs)** — variable-number microsatellite loci. Used in forensics, paternity testing, and population genetics. The probability of two unrelated individuals sharing the same STR profile at 13+ loci is less than 1 in 10 billion. **Gene therapy:**. Replacement or repair of defective genes in patients. Approaches: ex vivo (cells removed, modified, returned) and in vivo (vector delivered directly to the patient) Successes: treatment of severe combined immunodeficiency (SCID), spinal muscular atrophy (Zolgensma), inherited retinal dystrophy (Luxturna) **Ethical considerations:**. Germline editing — modifications that are heritable; raises profound ethical concerns. Genetic privacy and discrimination. Environmental release of genetically modified organisms. Equitable access to genetic technologies.
