Skip to main content

Protein Structure and Function

Learning Objectives

  • Describe the four levels of protein structure — primary, secondary, tertiary, and quaternary — and how each builds on the last.
  • Explain the relationship between protein structure and protein function.
  • Identify the major functional categories of proteins: enzymes, structural proteins, transport proteins, signaling receptors, and antibodies.
  • Explain how protein misfolding leads to disease, using specific examples.
  • Name and describe key bioinformatics tools used to determine or predict protein structure (PDB, PyMOL, SWISS-MODEL, AlphaFold).

Quick Answer

Protein structure describes how a chain of amino acids folds into a specific three-dimensional shape, and that shape is what determines what the protein can actually do. Structure is organized in four levels: primary (the amino acid sequence itself), secondary (local folding patterns like alpha helices and beta sheets), tertiary (the overall 3D shape of one chain), and quaternary (how multiple chains assemble together). This matters because function follows form in biology — an enzyme's active site, an antibody's binding pocket, and a receptor's ligand-binding groove all depend on precise 3D shapes, not just the underlying sequence. Understanding structure is essential for drug design, understanding disease (many diseases stem from misfolded proteins), and tools like AlphaFold now let researchers predict structure directly from sequence with near-experimental accuracy.

Overview

A protein's amino acid sequence is like a set of instructions for origami: the same strip of paper (the sequence) will always fold into the same shape (the structure) under normal conditions, and that shape determines what the folded object can do. Unlike DNA, which mainly stores information, proteins are the molecules that actually do things in a cell — catalyzing reactions, providing structural support, transporting molecules, and transmitting signals. Understanding how a linear chain of amino acids becomes a functional 3D machine is one of the central problems bioinformatics helps solve.

Core Concepts

Primary and Secondary Structure

Definition: Primary structure is the linear sequence of amino acids in a polypeptide chain. Secondary structure is the local folding pattern within that chain, stabilized by hydrogen bonds between backbone atoms.

Explanation: Primary structure is determined directly by the genetic code — the DNA sequence dictates, via mRNA and translation, exactly which amino acid goes in which position. Secondary structure emerges as the chain begins folding locally: alpha helices form a right-handed coil stabilized by regular hydrogen bonding, while beta sheets form when strands of the chain lie alongside each other (parallel or antiparallel) and hydrogen-bond across strands.

Example: The sequence Methionine-Glycine-Alanine-Serine-Leucine is a primary structure; whether that stretch folds into an alpha helix or beta strand depends on the chemical properties of each amino acid and its neighbors.

Real-World Example: Silk fibers get their strength largely from repeating beta-sheet structures in fibroin protein, which pack tightly and resist stretching — a direct link between secondary structure and a macroscopic material property.

Why It Matters: Primary structure is the "source code" — a single amino acid change (as in sickle cell disease, where one substitution changes glutamate to valine in hemoglobin) can alter secondary and tertiary structure enough to cause disease.

Common Misunderstanding: Students often think secondary structure is fixed for a given protein. In reality, some proteins are "intrinsically disordered," lacking a stable secondary structure until they bind a partner molecule, which is a legitimate and functionally important structural category, not a failure to fold.

Tertiary and Quaternary Structure

Definition: Tertiary structure is the overall three-dimensional shape of a single folded polypeptide chain. Quaternary structure describes how multiple polypeptide chains (subunits) assemble into a functional complex.

Explanation: Tertiary structure forms as secondary-structure elements (helices, sheets, loops) pack together, stabilized by hydrophobic interactions (nonpolar side chains clustering away from water), hydrogen bonds, ionic bonds, and sometimes covalent disulfide bridges. Not every protein needs a quaternary structure — many function as single chains — but many others only work when multiple subunits combine.

Example: Hemoglobin's quaternary structure consists of four subunits (two alpha, two beta), and this assembly is what allows cooperative oxygen binding — binding oxygen at one subunit makes it easier for the next subunit to bind oxygen too.

Real-World Example: DNA polymerase, the enzyme that replicates DNA, functions as a multi-subunit complex; if any required subunit is missing, the whole complex fails to work even though each individual subunit's tertiary structure may be intact.

Why It Matters: Tertiary and quaternary structure are usually where a protein's actual functional surfaces — active sites, binding pockets, subunit interfaces — are formed; sequence alone doesn't tell you where these surfaces will end up in 3D space.

Common Misunderstanding: Students often assume more subunits automatically mean a "more complex" or "more important" protein. Quaternary structure is about functional requirements (like cooperative binding or structural stability), not a hierarchy of importance — many essential proteins function perfectly well as single chains.

Structure Determines Function

Definition: The principle that a protein's 3D shape dictates what molecules it can interact with and what chemical or mechanical work it can perform.

Explanation: Enzymes catalyze reactions because their active site's shape and chemistry precisely fit a specific substrate; structural proteins like collagen provide strength because of their organized fibrous packing; transport proteins like hemoglobin bind and release cargo because conformational changes open and close binding sites; antibodies recognize pathogens because their variable regions form a shape complementary to a specific antigen.

Example: An enzyme's active site is shaped almost like a lock, and only a substrate with a complementary shape (the "key") binds efficiently — this is the classic lock-and-key model, later refined by the "induced fit" model, where the enzyme's shape adjusts slightly as the substrate binds.

Real-World Example: Many antiviral and cancer drugs are designed to fit precisely into a specific protein's active site or binding pocket, blocking its normal function — this is only possible because researchers know the protein's 3D structure in detail.

Why It Matters: This principle is the reason bioinformatics invests so heavily in structure prediction — if you know a protein's shape, you can often infer or manipulate its function, even without directly testing it in a lab.

Common Misunderstanding: Students sometimes think sequence similarity alone guarantees similar function. Two proteins can have quite different sequences but very similar 3D structures and functions (or vice versa) — structure is often more evolutionarily conserved than sequence, since many different sequences can fold into a similar shape.

Protein Misfolding and Disease

Definition: The failure of a protein to fold into its correct, functional 3D structure, often producing a toxic or non-functional molecule.

Explanation: Folding is guided by the amino acid sequence and assisted by molecular chaperones, but errors — from mutations, cellular stress, or aging — can cause misfolding. Misfolded proteins can lose function, aggregate into clumps, or trigger harmful cellular responses.

Example: In Alzheimer's disease, amyloid-beta protein misfolds and aggregates into plaques between neurons, which is strongly associated with (though not the sole cause of) neurodegeneration.

Real-World Example: In cystic fibrosis, a specific mutation (deletion of one amino acid, phenylalanine, at position 508) causes the CFTR protein to misfold; it gets degraded before reaching the cell membrane, disrupting chloride ion transport in the lungs and digestive system.

Why It Matters: Many major diseases — Alzheimer's, Parkinson's, cystic fibrosis, some cancers — trace back to protein misfolding, making structural bioinformatics directly relevant to understanding and treating disease.

Common Misunderstanding: Students often assume misfolded proteins are simply "broken" and harmless once non-functional. In diseases like Alzheimer's, misfolded proteins can actively aggregate and become toxic to cells, meaning the danger isn't just losing the protein's normal function but gaining a new, damaging one.

From Sequence to Function

Key Terms

TermDefinition
Primary structureThe linear sequence of amino acids in a polypeptide chain.
Secondary structureLocal folding patterns (alpha helix, beta sheet) stabilized by backbone hydrogen bonds.
Tertiary structureThe overall 3D shape of a single folded polypeptide chain.
Quaternary structureThe arrangement of multiple polypeptide subunits into one functional complex.
Molecular chaperoneA helper protein that assists correct folding of other proteins.
Active siteThe specific region of an enzyme where substrate binding and catalysis occur.
AlphaFoldA deep-learning system that predicts 3D protein structure from amino acid sequence with high accuracy.
Protein Data Bank (PDB)A public database of experimentally determined 3D structures of proteins and nucleic acids.

Common Mistakes

Misconception 1: "A protein's function can be fully predicted just from its amino acid sequence alone, without any structural information." Why it's wrong: Function emerges from the 3D arrangement of atoms, especially at active sites and binding surfaces — two very different sequences can produce similar structures and functions, and sequence alone doesn't reveal spatial relationships between distant residues that come close together after folding. Correct understanding: Structure prediction (experimental or computational) adds critical information beyond sequence, especially for identifying functional sites and understanding how mutations affect function.

Misconception 2: "All proteins need a quaternary structure to function." Why it's wrong: Many essential, fully functional proteins consist of just a single polypeptide chain and never assemble into a multi-subunit complex. Correct understanding: Quaternary structure is required only when a protein's specific function (like cooperative binding, mechanical stability, or forming a channel) depends on multiple subunits working together — it's not a universal requirement.

Misconception 3: "Misfolded proteins are simply inactive and therefore harmless." Why it's wrong: In diseases like Alzheimer's, misfolded proteins actively aggregate into structures (like amyloid plaques) that are toxic to neurons, rather than just failing to work. Correct understanding: Misfolding can create a new, harmful gain-of-function (aggregation, toxicity) in addition to, or instead of, simply losing the protein's normal role.

Comparison and Connections

Structure LevelWhat It DescribesStabilized ByExample
PrimaryAmino acid sequencePeptide bondsInsulin's exact amino acid chain
SecondaryLocal folding (helix/sheet)Backbone hydrogen bondsAlpha helix in a transmembrane protein
TertiaryOverall 3D shape of one chainHydrophobic interactions, ionic/hydrogen bonds, disulfide bridgesLysozyme's folded active-site pocket
QuaternaryAssembly of multiple chainsSame forces as tertiary, but between subunitsHemoglobin's four-subunit complex

Practice Questions

Recall 1: List the four levels of protein structure in order. Answer guidance: Primary, secondary, tertiary, and quaternary.

Recall 2: What determines a protein's primary structure? Answer guidance: The genetic code, transcribed into mRNA and translated into a specific sequence of amino acids.

Understanding 1: Explain why a single amino acid substitution can sometimes cause a serious disease. Answer guidance: A substitution can change the local chemistry enough to disrupt secondary or tertiary folding, alter a functional binding site, or promote aggregation — as in sickle cell disease, where one substitution changes hemoglobin's shape and behavior under low oxygen.

Understanding 2: Why is structure often more evolutionarily conserved than sequence? Answer guidance: Many different amino acid sequences can fold into a similar overall shape, so as long as the shape (and therefore function) is preserved, the underlying sequence can drift substantially through evolution without being selected against.

Application 1: A researcher has the sequence of a newly discovered protein but no experimental structure. What computational tool could they use to predict its 3D structure, and why has this become more reliable in recent years? Answer guidance: AlphaFold (or similar deep-learning-based predictors) — it has become far more reliable because it was trained on vast amounts of known structures and evolutionary sequence data, allowing it to predict structures with accuracy approaching experimental methods for many proteins.

Application 2: A pharmaceutical company wants to design a drug that blocks a specific enzyme without affecting similar enzymes in the body. What structural information would be most useful, and why? Answer guidance: A detailed 3D structure of the enzyme's active site, ideally compared against the active sites of similar enzymes, so the drug can be designed to fit precisely into the target's unique binding pocket while avoiding off-target enzymes with a different-shaped active site.

Analysis 1: Compare the roles of hydrophobic interactions in tertiary structure versus hydrogen bonds in secondary structure, and explain why both are needed for a protein to reach its final functional shape. Answer guidance: Hydrogen bonds in secondary structure create regular, local folding patterns (helices, sheets) along the backbone; hydrophobic interactions in tertiary structure drive the overall global packing, pulling nonpolar side chains into the protein's core away from water. Secondary structure provides the local building blocks, while hydrophobic-driven tertiary folding assembles those blocks into the compact, functional 3D shape — neither alone would produce a properly folded, functional protein.

Analysis 2: A student argues that because AlphaFold can now predict most protein structures accurately, experimental methods like X-ray crystallography are becoming obsolete. Evaluate this claim using what you know about protein misfolding and disease. Answer guidance: The claim overstates AlphaFold's role — while it predicts likely folded structures well, it doesn't directly model misfolding, aggregation, or how a specific disease-causing mutation disrupts normal folding dynamics, which are often central to understanding diseases like Alzheimer's or cystic fibrosis. Experimental methods remain essential for validating predictions, studying dynamic or misfolded states, and characterizing protein complexes and disease-relevant conformations that a static prediction may miss.

FAQ

Do all proteins have all four levels of structure? Every protein has primary and secondary structure, and virtually all have tertiary structure once folded. Quaternary structure only applies to proteins made of more than one polypeptide chain — many proteins function as a single chain and simply don't have this level.

What's the difference between the lock-and-key model and induced fit? Lock-and-key assumes the enzyme's active site is a rigid shape that perfectly matches the substrate. Induced fit is the more accurate, updated model: the enzyme's active site adjusts its shape slightly as the substrate binds, improving the fit dynamically.

Why do chaperones exist if folding is supposedly determined by sequence alone? Sequence does determine the final folded shape thermodynamically, but the folding pathway can go wrong, especially in the crowded, busy environment of a cell. Chaperones prevent misfolding and aggregation along the way, acting like a folding assistant rather than dictating the final shape themselves.

Is AlphaFold's prediction always correct? No — accuracy is generally very high for well-studied protein families with many evolutionary relatives, but it's lower for proteins with few relatives, disordered regions, or complex multi-protein assemblies. AlphaFold also provides a per-residue confidence score, which should always be checked.

How do disulfide bridges differ from hydrogen bonds in stabilizing structure? Disulfide bridges are covalent bonds between cysteine side chains, making them much stronger and more permanent than the weaker, more easily broken hydrogen bonds — which is why disulfide-rich proteins (like antibodies) tend to be especially stable, including outside the cell.

Quick Revision

  • Four levels of protein structure: primary (sequence), secondary (local folding), tertiary (3D shape of one chain), quaternary (multi-chain assembly).
  • Primary structure is set by the genetic code; secondary structure includes alpha helices and beta sheets stabilized by hydrogen bonds.
  • Tertiary structure is stabilized by hydrophobic interactions, ionic/hydrogen bonds, and disulfide bridges.
  • Not all proteins have quaternary structure — it's required only when function depends on multiple subunits.
  • Function follows structure: enzymes, structural proteins, transport proteins, receptors, and antibodies all depend on precise 3D shapes.
  • Structure is often more evolutionarily conserved than sequence.
  • Misfolding can cause disease both by losing normal function and by creating toxic aggregates (e.g., Alzheimer's amyloid plaques, cystic fibrosis CFTR misfolding).
  • PDB stores experimentally determined structures; PyMOL visualizes them; SWISS-MODEL and AlphaFold predict structures computationally.
  • AlphaFold uses deep learning and evolutionary sequence data to predict structure with near-experimental accuracy for many proteins, though confidence varies.
  • Structural knowledge is central to rational drug design, targeting active sites or binding pockets precisely.

Prerequisites: Introduction to Bioinformatics, basic biochemistry (amino acids, peptide bonds).

Related Topics: Sequence Alignment and Analysis, Computational Biology, Bioinformatics Tools and Software.

Next Topics: Applications in Research, Computational Biology.