DNA Structure and Function
Learning Objectives
- Describe the double helix model of DNA and explain how the sugar-phosphate backbone and nitrogenous bases combine to form a stable, copyable molecule
- State the base-pairing rules and explain why they allow DNA to be both a stable information store and a template for accurate copying
- Outline the major steps of DNA replication and identify the key enzymes at each stage
- Summarize how transcription converts a DNA template into mRNA, and how translation converts mRNA into a protein
- Distinguish the roles of mRNA, tRNA, and rRNA in gene expression
- Explain why DNA is chemically suited to be the hereditary molecule while RNA is suited to being a working messenger
Quick Answer
DNA (deoxyribonucleic acid) is the molecule that stores the genetic instructions for building and running every living cell. It is a double helix made of two antiparallel strands of nucleotides, each built from a deoxyribose sugar, a phosphate group, and one of four nitrogenous bases — adenine (A), thymine (T), guanine (G), and cytosine (C). A always pairs with T, and G always pairs with C, held together by hydrogen bonds. This complementary pairing is the reason DNA matters: it lets the molecule be copied with near-perfect accuracy during replication, and it lets one strand serve as a template for building RNA during transcription. The information in DNA ultimately gets expressed as proteins through the central dogma: DNA → RNA → protein. Every technique in biotechnology, from PCR to CRISPR, is built on these structural rules.
The Double Helix Model
James Watson and Francis Crick proposed the double helix structure in 1953, building on Rosalind Franklin's X-ray diffraction data (Photo 51) and Erwin Chargaff's earlier observation that the amount of A always equals the amount of T, and G always equals C, in any organism's DNA. The double helix consists of two strands of nucleotides twisted around a common axis, running in opposite (antiparallel) directions — one strand runs 5' to 3', the other 3' to 5'.
Each nucleotide has three parts: a five-carbon deoxyribose sugar, a phosphate group, and a nitrogenous base. The sugar and phosphate groups alternate to form the backbone on the outside of the helix, while the bases point inward and pair across the two strands, like rungs on a twisted ladder.
Why it matters: The backbone being on the outside and the bases on the inside is not a cosmetic detail — it protects the information-carrying bases from chemical attack while keeping the whole molecule chemically uniform and stable enough to sit in a nucleus for a human lifetime.
Common misunderstanding: Students often picture DNA as a flat ladder. It is actually a right-handed helix that makes one complete turn roughly every 10.5 base pairs, creating a major groove and a minor groove — grooves that proteins (like transcription factors) use to "read" the sequence without unwinding the whole molecule.
Base Pairing Rules
Adenine pairs with thymine through two hydrogen bonds, and guanine pairs with cytosine through three hydrogen bonds. This is called complementary base pairing, and it is not arbitrary — it comes from the shapes and hydrogen-bond donor/acceptor patterns of the bases themselves. A and G are larger two-ring purines; T and C are smaller one-ring pyrimidines. Pairing a purine with a pyrimidine keeps the width of the helix constant at every position, which is essential for the regular, repeating helical structure.
Example: If one strand reads 5'-ATCGGCTA-3', the complementary strand must read 3'-TAGCCGAT-5'.
Real-world example: GC-rich DNA (more G-C pairs) melts (denatures) at a higher temperature than AT-rich DNA because three hydrogen bonds take more energy to break than two. This single fact is why PCR primer design software calculates a melting temperature (Tm) based on GC content — get it wrong and the PCR reaction fails.
Why it matters: Because each strand specifies the sequence of its partner, either strand alone contains enough information to reconstruct the whole molecule. This redundancy is what makes accurate replication and reliable repair possible.
DNA Replication (Overview)
Before a cell divides, it must copy its entire genome so each daughter cell receives an identical set of chromosomes. DNA replication is semiconservative: each new double helix contains one original (parental) strand and one newly made strand — a fact confirmed experimentally by Meselson and Stahl in 1958.
Key steps:
- Initiation — Helicase unwinds the double helix at the origin of replication, creating a replication fork; single-strand binding proteins keep the strands apart; topoisomerase relieves the twisting strain ahead of the fork.
- Priming — Primase lays down a short RNA primer, because DNA polymerase cannot start a new strand from nothing; it can only extend an existing 3'-OH end.
- Elongation — DNA polymerase adds nucleotides in the 5' to 3' direction. Because the template strands run in opposite directions, one new strand (the leading strand) is made continuously, while the other (the lagging strand) is made discontinuously in Okazaki fragments.
- Joining — DNA ligase seals the gaps between Okazaki fragments once RNA primers are removed and replaced with DNA.
Why it matters: DNA polymerase also proofreads as it goes (a 3' to 5' exonuclease activity), catching and removing mismatched bases immediately. This is why replication has an error rate of roughly one mistake per billion bases copied — accuracy that a chemistry textbook would call remarkable for a machine running at hundreds of nucleotides per second.
(For the full replication and repair mechanism, see DNA Replication and Repair.)
Transcription and Translation (Overview)
The information in DNA is expressed through two further steps, together known as gene expression.
- Transcription: RNA polymerase binds a promoter sequence and copies one gene's DNA template strand into a complementary mRNA molecule, using U in place of T.
- Translation: Ribosomes read the mRNA three bases (one codon) at a time and, with the help of tRNA molecules carrying specific amino acids, assemble the corresponding protein.
Common misunderstanding: Many students think DNA itself directly builds proteins. It never leaves the nucleus (in eukaryotes) and never touches a ribosome. DNA is the archived master copy; mRNA is the disposable working copy that actually reaches the ribosome.
(For the full mechanism of both processes, see Transcription and Translation.)
Central Dogma at a Glance
Applications of DNA Technology
Understanding DNA structure is the foundation for nearly every technique in modern biotechnology:
- PCR (Polymerase Chain Reaction): Exploits complementary base pairing and DNA polymerase to exponentially amplify a chosen DNA segment.
- DNA sequencing: Reads the exact order of bases, enabling genome projects, diagnostics, and forensic identification.
- Recombinant DNA technology: Uses restriction enzymes (which recognize specific base sequences) and ligase to cut and paste DNA from different sources.
- CRISPR-Cas9 gene editing: A guide RNA base-pairs with a target DNA sequence, directing the Cas9 enzyme to cut at a precise location.
Real-world example: DNA fingerprinting in forensic science relies on the fact that short tandem repeat (STR) regions of the genome vary in length between individuals — a direct consequence of DNA's sequence-based identity.
Key Terms
| Term | Definition | Related Concept |
|---|---|---|
| Nucleotide | Building block of DNA: a deoxyribose sugar, a phosphate group, and one nitrogenous base | Base pairing, polymer |
| Double helix | The twisted-ladder structure of two antiparallel DNA strands | Watson-Crick model |
| Antiparallel | The two DNA strands run in opposite 5'→3' directions | Replication, transcription |
| Complementary base pairing | A pairs with T; G pairs with C, via hydrogen bonds | Replication fidelity |
| Major/minor groove | Grooves on the helix surface where proteins read the DNA sequence | Transcription factor binding |
| Semiconservative replication | Each new DNA molecule retains one parental strand and one new strand | Meselson-Stahl experiment |
| Okazaki fragment | Short DNA segment synthesized discontinuously on the lagging strand | Lagging strand, DNA ligase |
| Origin of replication | Specific DNA site where the double helix unwinds to begin copying | Replication fork |
| Promoter | DNA sequence where RNA polymerase binds to start transcription | Transcription |
| Codon | Three-base sequence in mRNA specifying one amino acid or a stop signal | Genetic code, translation |
| Central dogma | The flow of genetic information: DNA → RNA → protein | Gene expression |
Common Mistakes
Misconception: DNA directly builds proteins. Why it's wrong: DNA never leaves the nucleus in eukaryotic cells and never physically interacts with a ribosome. Correct understanding: DNA is transcribed into mRNA, which then travels to the ribosome for translation. DNA is the permanent archive; mRNA is the temporary working copy that is actually read to build a protein.
Misconception: The two strands of DNA are identical. Why it's wrong: The strands are complementary, not identical — where one strand has an A, the other has a T at that position, and vice versa for G and C. Correct understanding: Complementarity, not identity, is what makes replication possible: each strand acts as a template that specifies the exact sequence of a brand-new partner strand.
Misconception: All hydrogen bonds in DNA base pairs are the same strength, so any two bases could pair together. Why it's wrong: Only A-T (2 hydrogen bonds) and G-C (3 hydrogen bonds) pairings fit geometrically inside the helix; mismatched pairs (like A-G) would distort the helix's uniform width. Correct understanding: Base pairing is chemically and geometrically constrained. This is why GC-rich DNA is more thermally stable and why DNA polymerase and repair enzymes can detect and reject a mismatched base — it physically doesn't fit.
Comparison and Connections
| Feature | DNA | RNA |
|---|---|---|
| Strand structure | Double-stranded helix | Usually single-stranded |
| Sugar | Deoxyribose | Ribose |
| Bases | A, T, G, C | A, U, G, C |
| Stability | High (long-term storage) | Lower (short-lived, functional turnover) |
| Location (eukaryotes) | Nucleus | Made in nucleus, functions throughout cell |
| Primary role | Stores genetic information | Carries out gene expression (mRNA, tRNA, rRNA) |
Practice Questions
Recall
-
What are the four nitrogenous bases found in DNA? Answer guidance: Adenine, thymine, guanine, and cytosine.
-
Name the three components of a single nucleotide. Answer guidance: A five-carbon deoxyribose sugar, a phosphate group, and one nitrogenous base.
Understanding
-
Explain why A pairs specifically with T, and G specifically with C, rather than any other combination. Answer guidance: The purine-pyrimidine pairing keeps the helix width constant at every rung, and the hydrogen-bonding pattern (2 bonds for A-T, 3 for G-C) only lines up correctly for these specific pairs — any other combination would not fit the geometry of the double helix.
-
Why is DNA replication called "semiconservative"? Answer guidance: Each new double helix contains one original (parental) strand and one newly synthesized strand, because each old strand serves as the template for building one new complementary strand — demonstrated by the Meselson-Stahl experiment.
Application
-
A researcher is designing PCR primers for a GC-rich gene region. Why should the annealing temperature be set higher than for an AT-rich region? Answer guidance: GC pairs form three hydrogen bonds versus two for AT pairs, making GC-rich DNA more thermally stable and requiring more energy (higher temperature) to melt or keep primers correctly annealed.
-
A sample of double-stranded DNA is found to be 30% adenine. What percentage of the DNA is guanine? Answer guidance: By Chargaff's rule, %A = %T, so T is also 30%, leaving 40% for G and C combined; since %G = %C, guanine makes up 20%.
Analysis
-
Compare the consequences of a mutation in a gene's promoter region versus a mutation within the protein-coding sequence itself. Answer guidance: A promoter mutation can alter how much or when RNA polymerase binds, changing the gene's expression level without changing the protein sequence at all. A mutation in the coding sequence can change the amino acid sequence of the protein itself (or truncate it), potentially altering or destroying its function directly.
-
Explain why DNA is chemically better suited than RNA to be the long-term hereditary molecule, even though RNA is thought to have come first in evolution. Answer guidance: DNA's deoxyribose sugar lacks the reactive 2'-OH group that ribose has, making DNA chemically more stable and less prone to spontaneous degradation. DNA is also double-stranded, giving each strand a built-in backup template for repair — a self-correcting feature single-stranded RNA lacks.
FAQ
1. Is DNA the same thing as a gene? No. DNA is the entire molecule — the complete sequence organized into chromosomes. A gene is a specific segment of that DNA that codes for a functional product, usually a protein. Think of DNA as an entire library and a gene as a single book in it.
2. Why does DNA form a helix instead of a straight ladder? The twist minimizes steric strain between the stacked bases and allows the hydrophobic bases to pack tightly in the interior while the charged sugar-phosphate backbone faces the watery cell environment on the outside — the most energetically stable arrangement for this molecule.
3. How can DNA polymerase copy DNA so accurately? Two safeguards: DNA polymerase is highly selective about which incoming nucleotide it accepts (only a correctly paired base fits its active site well), and it has a built-in proofreading function that removes a mismatched base immediately after adding it, before continuing synthesis.
4. What's the difference between the leading and lagging strand? Both are made by the same enzyme reading template in the same 3'→5' direction, but because the two template strands are antiparallel, one new strand can be made continuously toward the replication fork (leading), while the other must be made in short, backward-pointing pieces called Okazaki fragments (lagging).
5. Why does understanding DNA structure matter for biotechnology, not just biology class? Every major lab technique — PCR, sequencing, cloning, CRISPR, gel electrophoresis — works by exploiting a specific structural property of DNA (base pairing, charge, size, or enzyme recognition sites). You cannot troubleshoot a failed PCR or design a cloning strategy without understanding why the molecule behaves the way it does.
Quick Revision
- DNA is a double helix of two antiparallel strands; deoxyribose + phosphate backbone on the outside, bases pointing inward
- Base pairing: A-T (2 H-bonds), G-C (3 H-bonds) — Chargaff's rule: %A = %T, %G = %C
- GC-rich DNA is more thermally stable (more hydrogen bonds to break)
- Replication is semiconservative (Meselson-Stahl experiment); needs a primer; leading strand continuous, lagging strand discontinuous (Okazaki fragments)
- DNA polymerase proofreads (3'→5' exonuclease), giving ~1 error per billion bases
- Central dogma: DNA → (transcription) → mRNA → (translation) → protein
- DNA never leaves the nucleus to build protein directly — mRNA is the intermediate
- Major and minor grooves let proteins read DNA sequence without unwinding the helix
- PCR, sequencing, cloning, and CRISPR all exploit base pairing and DNA-modifying enzymes
Related Topics
Prerequisites: Basic cell biology (nucleus, chromosomes), general chemistry (hydrogen bonding, polymers, functional groups)
Related Topics: RNA Structure and Function, DNA Replication and Repair, Transcription and Translation
Next Topics: DNA Replication and Repair, Transcription and Translation, Techniques in Molecular Biology