Skip to main content

Transcription and Translation

Learning Objectives

  • Describe the three stages of transcription and identify the enzyme responsible for each
  • Explain how eukaryotic pre-mRNA is processed (capping, polyadenylation, splicing) before translation
  • Describe the three stages of translation and the roles of mRNA, tRNA, and the ribosome in each
  • Explain the properties of the genetic code (triplet, degenerate, unambiguous) and their consequences for mutations
  • Distinguish where transcription and translation occur in prokaryotes versus eukaryotes and explain why the difference matters
  • Apply knowledge of transcription and translation to predict the effect of a mutation on the resulting protein

Quick Answer

Transcription and translation are the two steps that convert the genetic information stored in DNA into a functional protein — together they make up gene expression. In transcription, RNA polymerase copies one gene's DNA template strand into a complementary mRNA molecule. In eukaryotes, that pre-mRNA is processed (a protective cap is added, a poly-A tail is added, and introns are spliced out) before leaving the nucleus. In translation, ribosomes read the mature mRNA three bases (one codon) at a time, and tRNA molecules — each carrying a specific amino acid matched to its anticodon — deliver the correct amino acids in order, building a polypeptide chain. This DNA → mRNA → protein pathway is the central dogma of molecular biology, and understanding each step precisely is essential for predicting how mutations, drugs, and biotechnology tools (like CRISPR or antibiotics that block bacterial ribosomes) affect a cell.

Transcription: DNA to mRNA

Transcription is the process by which RNA polymerase synthesizes an RNA molecule using one strand of DNA as a template. It occurs in the nucleus in eukaryotes and in the cytoplasm in prokaryotes (which lack a nucleus).

Stages of Transcription

  1. Initiation: RNA polymerase binds to a promoter — a specific DNA sequence upstream of the gene. In eukaryotes, transcription factors must first assemble at the promoter (often at a TATA box) to recruit RNA polymerase II.
  2. Elongation: RNA polymerase unwinds the DNA locally and moves along the template strand (read 3' to 5'), synthesizing a complementary mRNA strand 5' to 3', using uracil in place of thymine.
  3. Termination: RNA polymerase reaches a terminator sequence and releases the completed RNA transcript.

Why it matters: Unlike DNA polymerase, RNA polymerase does not require a primer and has little to no proofreading ability. This is precisely why RNA has a higher error rate than DNA — but because many mRNA copies of a gene are made and each is short-lived, an occasional transcription error is far less costly to the cell than a permanent DNA replication error would be.

RNA Processing in Eukaryotes

The initial transcript, pre-mRNA, is processed before it can be translated:

  • 5' capping: A modified guanine nucleotide is added to the 5' end, protecting the mRNA from degradation and helping the ribosome recognize where to begin.
  • 3' polyadenylation: A string of adenine nucleotides (poly-A tail) is added to the 3' end, further stabilizing the mRNA and aiding its export from the nucleus.
  • Splicing: Introns (non-coding sequences) are removed and exons (coding sequences) are joined together by a complex called the spliceosome, guided by small nuclear RNAs (snRNAs).

Common misunderstanding: Students often assume the entire transcript codes for protein. In fact, most of a human gene's pre-mRNA is intronic and is spliced out before translation. Alternative splicing — joining exons in different combinations — is a major reason why roughly 20,000 human genes can produce well over 100,000 distinct proteins.

Real-world example: Some beta-thalassemia mutations disrupt a splice site rather than the coding sequence itself, causing intron retention or exon skipping — proof that splicing errors, not just coding mutations, can cause disease.

Translation: mRNA to Protein

Translation converts the sequence of codons in mature mRNA into a sequence of amino acids, forming a polypeptide chain. It occurs at ribosomes, which are complexes of rRNA and protein found in the cytoplasm (and on the rough endoplasmic reticulum for secreted proteins).

Stages of Translation

  1. Initiation: The small ribosomal subunit binds the mRNA and scans to the start codon (AUG, which codes for methionine). The first tRNA, carrying methionine, base-pairs with the start codon via its anticodon, and the large ribosomal subunit then joins to complete the functional ribosome.
  2. Elongation: Charged tRNAs enter the ribosome's A site, and their anticodon pairs with the mRNA codon currently being read. A peptide bond forms between the new amino acid and the growing chain, and the ribosome then translocates one codon further along the mRNA.
  3. Termination: When the ribosome reaches a stop codon (UAA, UAG, or UGA), no tRNA matches it. Instead, a release factor binds, triggering release of the completed polypeptide chain from the ribosome.

Example: For the mRNA sequence 5'-AUG-GGC-UUU-UAA-3', the ribosome reads: AUG (start, methionine) → GGC (glycine) → UUU (phenylalanine) → UAA (stop). The resulting polypeptide is Met-Gly-Phe.

Why it matters: The genetic code is read in non-overlapping three-base codons with no punctuation between them — a fixed "reading frame." This is why an insertion or deletion of one or two bases (but not three) is so damaging: it shifts every codon downstream of the mutation, garbling the entire rest of the protein.

The Genetic Code

The genetic code is the set of rules that maps each three-base mRNA codon to a specific amino acid (or a stop signal). It has three important properties:

  • Triplet: Codons are always read three bases at a time.
  • Degenerate: Most amino acids are specified by more than one codon (for example, leucine has six different codons), which provides some tolerance against mutation.
  • Unambiguous: Each individual codon specifies only one amino acid — there is no overlap in the other direction.
  • Nearly universal: The same genetic code is used (with rare exceptions) across nearly all known life, which is why genes from one organism (like a jellyfish's fluorescent protein gene) can be expressed correctly in a completely different organism (like bacteria or plants) — the basis of recombinant DNA technology.
CodonAmino Acid
AUGMethionine (Start)
UUU, UUCPhenylalanine
UUA, UUG, CUU, CUC, CUA, CUGLeucine
GGU, GGC, GGA, GGGGlycine
UAA, UAG, UGAStop (no amino acid)

Why it matters: Degeneracy means many single-base changes are "silent" — they change the codon but not the amino acid, so the protein is unaffected. This is a key reason not every DNA mutation causes disease.

Transcription and Translation Pathway

Prokaryotic vs Eukaryotic Gene Expression

In prokaryotes, transcription and translation happen in the same cellular compartment (there is no nucleus), and translation of an mRNA can begin even before transcription of that mRNA is finished — a process called coupled transcription-translation. In eukaryotes, transcription and RNA processing occur in the nucleus, and only mature, processed mRNA is exported to the cytoplasm for translation. This spatial and temporal separation gives eukaryotes an additional checkpoint (RNA processing and nuclear export) at which gene expression can be regulated.

Real-world example: Several classes of antibiotics (like tetracyclines and macrolides) selectively bind bacterial (prokaryotic) ribosomes without significantly affecting human (eukaryotic) ribosomes, because the two ribosome types differ structurally enough for selective drug targeting — a direct clinical application of understanding translation machinery differences.

Key Terms

TermDefinitionRelated Concept
PromoterDNA sequence where RNA polymerase (and transcription factors) bind to initiate transcriptionTranscription initiation
Template strandThe DNA strand read by RNA polymerase to synthesize a complementary RNACoding strand, transcription
Pre-mRNAThe initial, unprocessed RNA transcript, still containing intronsRNA processing, splicing
5' capModified guanine nucleotide added to the 5' end of mRNA, protecting it and aiding ribosome bindingmRNA stability
Poly-A tailString of adenines added to the 3' end of mRNA, stabilizing it and aiding nuclear exportmRNA stability
SpliceosomeComplex of snRNA and proteins that removes introns from pre-mRNASplicing, alternative splicing
CodonThree-base mRNA sequence specifying one amino acid or a stop signalGenetic code, translation
AnticodonThree-base sequence on tRNA that pairs with an mRNA codontRNA, translation
Start codonAUG; signals the beginning of translation and codes for methionineTranslation initiation
Stop codonUAA, UAG, or UGA; signals the end of translation, no matching tRNATranslation termination
Reading frameThe specific grouping of mRNA bases into consecutive, non-overlapping codonsFrameshift mutation
Degeneracy (genetic code)Multiple codons can specify the same amino acidSilent mutation

Common Mistakes

Misconception: The entire pre-mRNA transcript is translated into protein. Why it's wrong: Pre-mRNA contains introns, which are removed by the spliceosome before the mRNA is considered mature and ready for translation. Correct understanding: Only the spliced-together exons remain in mature mRNA and are translated. Alternative splicing of the same pre-mRNA can produce multiple different mature mRNAs — and therefore multiple different proteins — from a single gene.


Misconception: DNA polymerase and RNA polymerase work the same way, so RNA polymerase also needs a primer. Why it's wrong: RNA polymerase can initiate a new RNA strand de novo, directly at the promoter, without needing an existing 3'-OH end to extend. Correct understanding: Only DNA polymerase requires a primer (because it can only extend an existing strand). RNA polymerase's ability to start from scratch is exactly why primase (which makes RNA primers for DNA replication) exists — it borrows this exact capability.


Misconception: Any insertion or deletion mutation causes a frameshift. Why it's wrong: Whether a reading frame shifts depends on how many bases are inserted or deleted, not simply whether an insertion or deletion occurred. Correct understanding: Insertions or deletions in multiples of three bases add or remove whole codons (whole amino acids) without shifting the reading frame downstream. Only insertions or deletions not in multiples of three shift every subsequent codon, typically producing a severely altered or nonfunctional protein.

Comparison and Connections

FeatureTranscriptionTranslation
TemplateDNA (one gene's template strand)mRNA
ProductRNA (pre-mRNA in eukaryotes)Polypeptide chain
Key enzyme/machineRNA polymeraseRibosome (rRNA + protein), tRNA
Location (eukaryotes)NucleusCytoplasm (or rough ER for secreted proteins)
Primer needed?NoNo (but requires an initiator tRNA)
ProofreadingMinimalErrors caught by ribosome fidelity checks, not repair enzymes
Direction of synthesis5' to 3' (reads template 3' to 5')Reads mRNA 5' to 3'; polypeptide grows N- to C-terminus

Practice Questions

Recall

  1. Name the three stages of transcription and the three stages of translation. Answer guidance: Transcription: initiation, elongation, termination. Translation: initiation, elongation, termination.

  2. What is the start codon, and which amino acid does it specify? Answer guidance: AUG, which specifies methionine.

Understanding

  1. Explain why eukaryotic pre-mRNA must be processed (capped, polyadenylated, and spliced) before it can be translated, while prokaryotic mRNA generally does not need this processing. Answer guidance: Eukaryotic genes contain introns that must be spliced out, and the cap and poly-A tail protect the mRNA during its journey out of the nucleus into the cytoplasm. Prokaryotes lack a nucleus and generally lack introns in most genes, so their mRNA can be translated immediately, often even while it is still being transcribed.

  2. Why is the genetic code described as "degenerate but unambiguous"? Answer guidance: Degenerate means multiple different codons can code for the same amino acid (for example, six codons all specify leucine). Unambiguous means that despite this redundancy, any single given codon always specifies only one particular amino acid — there is no codon that could mean two different things.

Application

  1. A single-nucleotide substitution changes a codon from CAG (glutamine) to CAA. Predict the effect on the resulting protein, and explain your reasoning using codon degeneracy. Answer guidance: Both CAG and CAA code for glutamine, so this is a silent mutation — the protein sequence does not change at all, because the genetic code's degeneracy allows different codons to specify the same amino acid.

  2. A drug is developed that selectively binds bacterial (70S) ribosomes but not human (80S) ribosomes. Explain how this selectivity could be therapeutically useful, and name a class of drug that works this way. Answer guidance: Because bacterial and human ribosomes are structurally different, a drug bound specifically to the bacterial ribosome can block bacterial protein synthesis (killing or inhibiting the bacteria) while leaving human translation, and therefore human cells, largely unaffected. Tetracyclines and macrolide antibiotics work by this mechanism.

Analysis

  1. Compare the consequence of a 3-base deletion versus a 1-base deletion occurring early in a gene's coding sequence. Answer guidance: A 3-base deletion removes exactly one whole codon (one amino acid) without shifting the reading frame, so the rest of the protein sequence downstream is unaffected — the protein is shorter by one amino acid but otherwise normal. A 1-base deletion shifts the reading frame for every codon downstream of the deletion, typically scrambling the entire remaining amino acid sequence and often introducing a premature stop codon, usually producing a severely truncated, nonfunctional protein.

  2. Explain why alternative splicing allows a single gene to produce multiple distinct proteins, and why this matters for the "one gene, one protein" idea historically taught in biology. Answer guidance: Alternative splicing allows the same pre-mRNA transcript to be spliced in different combinations of exons, producing multiple distinct mature mRNAs and therefore multiple distinct proteins (isoforms) from a single gene. This directly contradicts the older "one gene, one protein" hypothesis and explains how humans can have roughly 20,000 genes but well over 100,000 distinct proteins.

FAQ

1. Why doesn't RNA polymerase need a primer, but DNA polymerase does? RNA polymerase can catalyze the formation of the first phosphodiester bond between two free ribonucleotides directly at the promoter, whereas DNA polymerase's active site is structured to only extend an existing 3'-OH end, never to join two free nucleotides together.

2. What actually happens to the introns after they're spliced out? They are typically degraded rapidly within the nucleus. In some cases, however, intron sequences are processed further and give rise to functional small RNAs, such as some classes of microRNA — so "junk" introns aren't always entirely disposable.

3. How does the ribosome know which of the three possible reading frames to use? The ribosome always begins reading at the first AUG start codon it identifies (typically the first one after the 5' cap in a scanning process in eukaryotes), which fixes the reading frame for the rest of that particular translation event.

4. Can a mutation in an intron ever matter if introns get spliced out anyway? Yes. Mutations at or near splice sites — the exact boundary sequences the spliceosome uses to recognize where to cut — can disrupt splicing itself, causing intron retention or exon skipping, even though the mutation is technically located in "non-coding" intronic sequence.

5. Why do biotechnology students need to understand transcription and translation in this much detail? Nearly every genetic engineering technique depends on manipulating these exact steps: choosing the right promoter for a plasmid construct, designing primers around exon-intron boundaries, predicting whether a CRISPR-induced indel will cause a frameshift, or engineering codon-optimized genes for expression in a different host organism. Precision at this level is what separates a working construct from a failed experiment.

Quick Revision

  • Transcription: initiation (promoter binding) → elongation (RNA polymerase synthesizes RNA) → termination
  • Eukaryotic pre-mRNA processing: 5' cap, 3' poly-A tail, splicing (introns removed, exons joined)
  • Translation: initiation (start codon AUG, methionine) → elongation (tRNA delivers amino acids) → termination (stop codon, release factor)
  • Genetic code: triplet, degenerate (multiple codons per amino acid), unambiguous, nearly universal
  • Reading frame shifts occur only when indels are not in multiples of three bases
  • Silent mutations don't change the protein (codon degeneracy); missense changes one amino acid; nonsense causes premature truncation
  • Alternative splicing lets one gene produce multiple protein isoforms
  • Prokaryotes: transcription/translation in same compartment, can be coupled; Eukaryotes: transcription/processing in nucleus, translation in cytoplasm
  • Antibiotics like tetracyclines exploit structural differences between bacterial and human ribosomes

Prerequisites: DNA Structure and Function, RNA Structure and Function

Related Topics: Gene Regulation, DNA Replication and Repair, Techniques in Molecular Biology (Northern blotting, RT-PCR)

Next Topics: Gene Regulation, Genomic and Proteomic Approaches, Techniques in Molecular Biology