Epidemiology and Biostatistics
Learning Objectives
By the end of this chapter, you should be able to:
- Distinguish observational study designs (cross-sectional, case-control, cohort) from experimental designs (RCTs) and identify the correct measure of association for each.
- Calculate and interpret incidence, prevalence, relative risk, odds ratio, and attributable risk from a 2x2 table.
- Explain why case-control studies use odds ratios while cohort studies and RCTs use relative risk.
- Calculate sensitivity, specificity, PPV, and NPV from a screening test's 2x2 table, and explain how disease prevalence affects predictive values.
- Interpret a p-value and a 95% confidence interval correctly, and state what each does and does not mean.
- Choose an appropriate study design for a given research question and justify the choice.
Quick Answer
Epidemiology studies how disease is distributed in populations and what causes it; biostatistics gives epidemiology its mathematical tools. The three workhorse observational designs are cross-sectional (prevalence, one point in time), case-control (start with disease status, look backward for exposure, report odds ratio), and cohort (start with exposure status, look forward for outcome, report relative risk). Randomized controlled trials add random allocation to remove confounding and are the strongest design for proving causation. Screening tests are judged by sensitivity (catches true positives) and specificity (rejects true negatives), while predictive values tell you what a positive or negative result actually means for a given patient — and those depend heavily on how common the disease is in the population being tested.
Overview
Every clinical decision — "does this drug work," "is this exposure dangerous," "should we screen for this cancer" — ultimately rests on a study design and a set of numbers. Epidemiology is the discipline that designs the study to answer the question correctly; biostatistics is the discipline that turns the resulting data into a number you can trust and interpret.
The exam-relevant core is smaller than it looks: know the four main study designs and what each is good and bad at, know how to build and read a 2x2 table (both for disease-exposure questions and for screening test performance), and know what a p-value and confidence interval actually communicate. Almost every PSM/community medicine question is a variation on these skills applied to a clinical vignette.
Study Designs
Cross-Sectional Studies
Definition: A snapshot study — exposure and outcome are measured at the same point in time in a defined population.
Explanation: You survey a population once and record who has the exposure and who has the disease simultaneously. Because there's no time sequence, you cannot say the exposure came before the disease — only that they occur together. This design gives you prevalence, not incidence.
Example: Surveying 1,000 adults in a city on a single day to find what fraction currently has hypertension and whether it correlates with self-reported salt intake.
Real-World Example: The National Family Health Survey (NFHS) in India measures prevalence of anemia, stunting, and diabetes across states in a single survey round — it tells you "how much disease exists now," not "what caused it."
Why It Matters: Cross-sectional studies are cheap, fast, and useful for health planning (how many dialysis machines does a district need?), but they cannot establish causation.
Common Misunderstanding: Students often say a cross-sectional study shows "risk of developing disease." It cannot — there's no follow-up, so there's no incidence and no temporal sequence.
Case-Control Studies
Definition: A study that starts by selecting people who already have the disease (cases) and people who don't (controls), then looks backward to compare their prior exposure history.
Explanation: Because you begin with outcome status, you cannot calculate incidence in either group — you don't know the true size of the exposed and unexposed populations at risk. What you can calculate is the odds of exposure in cases versus controls, giving an odds ratio (OR) as the measure of association.
Example: Selecting 200 patients with lung cancer and 200 matched controls without lung cancer, then asking both groups about historical smoking habits.
Real-World Example: Doll and Hill's classic 1950 case-control study of lung cancer patients in London hospitals found heavy smokers had a markedly higher odds of being smokers than matched controls — one of the first studies to implicate smoking in lung cancer.
Why It Matters: Case-control studies are the design of choice for rare diseases because they don't require following huge cohorts for years to accumulate enough cases. They're also relatively quick and cheap.
Common Misunderstanding: Students frequently report "relative risk" from a case-control study. This is wrong — you cannot compute incidence or RR because the ratio of cases to controls is fixed by the investigator, not by nature. Only the odds ratio is valid, though it approximates RR reasonably well when the disease is rare (the "rare disease assumption").
Cohort Studies
Definition: A study that starts with a group free of disease, divides them by exposure status, and follows them forward in time to see who develops the outcome.
Explanation: Because you know the exact population at risk in each exposure group at the start, you can calculate incidence in the exposed and unexposed groups directly, and from that the relative risk (RR).
Example: Following 5,000 smokers and 5,000 non-smokers for 20 years and comparing the incidence of lung cancer in each group. This is a prospective cohort. A retrospective (historical) cohort uses existing records — e.g., factory employment logs from 1980 — to reconstruct exposure and then checks current disease registries for outcomes.
Real-World Example: The Framingham Heart Study has followed residents of Framingham, Massachusetts since 1948, establishing the link between hypertension, cholesterol, smoking, and cardiovascular disease — the foundation of modern cardiovascular risk prediction.
Why It Matters: Cohort studies can study multiple outcomes from one exposure, establish clear temporal sequence, and directly measure incidence. They're the best observational design for rare exposures (e.g., unusual occupational chemicals).
Common Misunderstanding: Students assume cohort studies are always prospective. Retrospective cohorts exist and are cheaper/faster, though they depend on the quality of historical records.
Randomized Controlled Trials (RCTs)
Definition: An experimental study in which the investigator randomly assigns participants to intervention or control groups, then follows both forward to compare outcomes.
Explanation: Randomization is what makes RCTs special — done properly, it distributes both known and unknown confounders evenly between groups, so any difference in outcome can be attributed to the intervention. Blinding (single, double) further reduces bias in outcome assessment.
Example: Randomly assigning 2,000 hypertensive patients to a new antihypertensive drug or placebo and comparing rates of stroke over 5 years.
Real-World Example: The UK ISIS-2 trial randomized thousands of heart attack patients to aspirin, streptokinase, both, or neither, and directly demonstrated that aspirin reduces mortality after MI — a finding that changed global practice.
Why It Matters: RCTs sit at the top of the evidence hierarchy for proving causation because randomization controls confounding that no observational design can fully eliminate.
Common Misunderstanding: Students think randomization guarantees the two groups are identical. It only guarantees confounders are, on average, evenly distributed — chance imbalance can still occur, especially in small trials, which is why baseline characteristic tables are reported.
Measures of Disease Frequency
Definition: Incidence is the number of new cases of a disease occurring in a population at risk over a defined time period. Prevalence is the number or proportion of existing cases (new + old) at a given point (point prevalence) or over a period (period prevalence).
Explanation: Incidence needs a denominator of people who are disease-free at the start and genuinely "at risk" (e.g., don't count men in the denominator for cervical cancer incidence). Prevalence and incidence are linked by disease duration: roughly, Prevalence ≈ Incidence × Average Duration of Disease (valid when prevalence is low and the disease is in steady state).
Example: In a town of 10,000 people, 50 new cases of diabetes are diagnosed in a year (incidence = 50/10,000/year), while 500 people currently live with diabetes at the start of the year (point prevalence = 500/10,000).
Real-World Example: A disease that is quickly fatal (like untreated rabies) has high incidence but very low prevalence because duration is short. A chronic, non-fatal disease with good treatment (like well-controlled diabetes) can have low incidence but high prevalence because people live with it for decades.
Why It Matters: Choosing the wrong measure changes the story completely — a disease can look like it's "increasing" in prevalence purely because treatment has improved and people are surviving longer with it, not because new cases are rising.
Common Misunderstanding: Students use "incidence" and "prevalence" interchangeably. They answer different questions: incidence answers "what's my risk of getting this," prevalence answers "how much of this disease exists right now" (relevant for resource planning).
Measuring Association: Relative Risk, Odds Ratio, and Attributable Risk
Build a standard 2x2 table with disease on one axis and exposure on the other:
| Disease + | Disease − | |
|---|---|---|
| Exposed | a | b |
| Unexposed | c | d |
- Relative Risk (RR) = [a/(a+b)] ÷ [c/(c+d)] — used in cohort studies and RCTs, where incidence in each group is known.
- Odds Ratio (OR) = (a×d) ÷ (b×c) — used in case-control studies, where incidence cannot be calculated.
- Attributable Risk (AR) = Incidence in exposed − Incidence in unexposed — the extra risk directly caused by the exposure, useful for public health impact (how many cases would be prevented by removing this exposure).
- Relative Risk Reduction (RRR) and Number Needed to Treat (NNT = 1/Absolute Risk Reduction) are the standard way trial results are reported for interventions.
Interpretation rule: RR/OR = 1 means no association; > 1 means increased risk with exposure; < 1 means the exposure is protective. An OR of 2.5 means the odds of the outcome are 2.5 times higher in the exposed/case group — it is not the same as "2.5 times more likely," a distinction that matters most when the outcome is common (OR overestimates RR as baseline risk rises).
Screening Tests: Sensitivity, Specificity, and Predictive Values
Build a 2x2 table with true disease status against test result:
| Disease + | Disease − | |
|---|---|---|
| Test Positive | a (TP) | b (FP) |
| Test Negative | c (FN) | d (TN) |
- Sensitivity = a/(a+c) — proportion of truly diseased people the test correctly flags as positive. A highly sensitive test, when negative, helps rule out disease (mnemonic: SnNout).
- Specificity = d/(b+d) — proportion of truly healthy people the test correctly flags as negative. A highly specific test, when positive, helps rule in disease (mnemonic: SpPin).
- Positive Predictive Value (PPV) = a/(a+b) — given a positive result, the probability the person actually has the disease.
- Negative Predictive Value (NPV) = d/(c+d) — given a negative result, the probability the person is actually disease-free.
The prevalence trap: sensitivity and specificity are intrinsic properties of the test and don't change with prevalence, but PPV and NPV depend heavily on it. In a low-prevalence population (e.g., mass screening for a rare cancer), even a very specific test produces many false positives relative to true positives, so PPV drops sharply. This is exactly why a positive screening test always needs a confirmatory diagnostic test before treatment begins.
Common Biostatistics Concepts
Definition: A p-value is the probability of observing a result as extreme as (or more extreme than) the one obtained, if the null hypothesis were actually true. A confidence interval (CI) is a range of values, calculated from sample data, that is likely to contain the true population parameter.
Explanation: A p-value < 0.05 (the conventional threshold) means the observed result would occur less than 5% of the time by chance alone if there were truly no effect — so we reject the null hypothesis. A 95% CI means that if we repeated the study infinitely many times, 95% of the calculated intervals would contain the true value. If a 95% CI for a relative risk includes 1.0, the result is not statistically significant, mirroring a p-value ≥ 0.05.
Example: A drug trial reports RR = 0.7 (95% CI: 0.5–0.9), p = 0.01. Since the CI excludes 1.0 and p < 0.05, the reduced risk is statistically significant.
Real-World Example: Large multicentre trials report both p-values and CIs because the CI additionally tells you about the precision (width) and clinical magnitude of the effect — a statistically significant p-value with a tiny effect size may not be clinically meaningful, which the CI reveals but the p-value alone does not.
Why It Matters: Exam questions frequently test whether a CI crossing 1 (for RR/OR) or 0 (for a mean difference) implies non-significance — recognizing this instantly saves time.
Common Misunderstanding: A p-value of 0.03 does not mean "there's a 97% chance the result is real" or "3% chance the null hypothesis is true." It only describes how surprising the data would be under the null hypothesis — it says nothing about the probability that the null hypothesis itself is true.
Visual Learning: Classifying Study Designs
Key Terms
| Term | Definition |
|---|---|
| Incidence | New cases of disease per population at risk, over a time period |
| Prevalence | Total existing cases (new + old) at a point or over a period |
| Relative Risk (RR) | Ratio of incidence in exposed vs. unexposed; used in cohort/RCT |
| Odds Ratio (OR) | Ratio of odds of exposure in cases vs. controls; used in case-control |
| Attributable Risk | Excess incidence in exposed group directly due to exposure |
| Sensitivity | Ability of a test to correctly identify true positives |
| Specificity | Ability of a test to correctly identify true negatives |
| PPV / NPV | Probability that a positive/negative test result is a true result |
| Confounding | A third variable that distorts the true association between exposure and outcome |
| Bias | Systematic error in study design/conduct that distorts results (selection bias, recall bias, information bias) |
| p-value | Probability of the observed data (or more extreme) under the null hypothesis |
| Confidence Interval | Range likely to contain the true population parameter, with a stated confidence level |
| Number Needed to Treat (NNT) | Number of patients who must be treated to prevent one additional bad outcome |
Common Mistakes
-
Misconception: "Relative risk can be calculated from a case-control study." Why it's wrong: In a case-control study, the number of cases and controls is fixed by the investigator's sampling design, not by disease occurrence in the source population, so true incidence — and therefore RR — cannot be derived. Correct: Only the odds ratio is valid in case-control studies; it approximates RR when the disease is rare (rare disease assumption).
-
Misconception: "A screening test with 99% specificity will almost always be right when it's positive." Why it's wrong: This ignores prevalence. In a low-prevalence population, false positives from the large healthy majority can outnumber true positives from the small diseased minority, dragging PPV down substantially even with excellent specificity. Correct: PPV must be calculated using prevalence, sensitivity, and specificity together (via Bayes' theorem) — specificity alone doesn't tell you PPV.
-
Misconception: "A p-value of 0.04 means there's a 96% probability that the treatment works." Why it's wrong: The p-value is calculated assuming the null hypothesis is true; it describes the probability of the data given the null, not the probability of a hypothesis given the data. Correct: p = 0.04 means: if there truly were no effect, data this extreme would occur only 4% of the time by chance. It says nothing directly about the probability that the treatment is effective.
Comparison and Connections
| Feature | Cross-Sectional | Case-Control | Cohort | RCT |
|---|---|---|---|---|
| Starting point | Whole population, one time point | Disease status (cases vs. controls) | Exposure status | Random assignment by investigator |
| Direction | None (simultaneous) | Backward (retrospective) | Forward (usually) | Forward |
| Measure of association | Prevalence | Odds Ratio | Relative Risk | Relative Risk / Risk Reduction |
| Best for | Health planning, prevalence surveys | Rare diseases | Rare exposures, multiple outcomes | Proving causation |
| Confounding control | Weak | Weak (matching helps) | Moderate | Strong (randomization) |
| Speed/cost | Fast, cheap | Fast, cheap | Slow, expensive | Slow, most expensive |
Practice Questions
Recall
-
Define incidence and prevalence, and give the approximate formula relating them. Answer guidance: Incidence = new cases/population at risk/time; Prevalence = existing cases/population at a point or period; Prevalence ≈ Incidence × average disease duration.
-
What measure of association is reported in a case-control study and why? Answer guidance: Odds ratio — because incidence and therefore relative risk cannot be calculated when case and control numbers are fixed by study design.
Understanding
-
Explain why sensitivity and specificity remain constant across populations but PPV and NPV do not. Answer guidance: Sensitivity/specificity are properties of the test relative to disease status (rows of the table), independent of how common disease is. PPV/NPV depend on how many diseased vs. healthy people are actually tested (prevalence), which shifts the ratio of true to false positives/negatives.
-
Why is randomization the defining strength of an RCT compared to a cohort study? Answer guidance: Randomization balances both known and unknown confounders between groups on average, whereas cohort studies rely on natural exposure patterns that may be linked to other risk factors (confounding by indication, healthy worker effect, etc.).
Application
-
A new blood test for a rare cancer (prevalence 1%) has 95% sensitivity and 90% specificity. In a population of 10,000, how many people will test positive, and what fraction of those truly have cancer? Answer guidance: Diseased = 100, Healthy = 9,900. True positives = 95, False positives = 990. Total positives = 1,085. PPV = 95/1085 ≈ 8.8% — illustrating how low prevalence drags PPV down even with good sensitivity/specificity.
-
A researcher wants to study whether a rare occupational chemical exposure causes a specific cancer that takes 20 years to develop. Which study design is most appropriate and why? Answer guidance: Retrospective (historical) cohort study — start with exposed vs. unexposed workers from employment records and check current disease registries; a case-control study is possible too (since the disease is presumably also rare) but cohort is ideal when exposure records exist and the exposed population is identifiable.
Analysis
-
A cohort study reports RR = 1.8 (95% CI: 0.9–3.2) for an exposure-disease association. Is this result statistically significant? Justify your answer. Answer guidance: Not statistically significant — the 95% CI crosses 1.0, meaning "no association" is a plausible value consistent with the data, even though the point estimate suggests increased risk.
-
Compare and contrast why a case-control study and a cohort study examining the same exposure-disease pair might report different-looking numbers (OR vs. RR) even if the true association is identical. Answer guidance: Case-control OR approximates RR only under the rare disease assumption; if the disease is not rare, OR will overestimate RR (moving further from 1 in the same direction), so the two designs' reported numbers can differ even when both are measuring a true, identical association.
FAQ
1. Why can't a case-control study calculate incidence? Because the investigator, not nature, decides how many cases and how many controls to enroll — the ratio of cases to controls in the sample doesn't reflect their true ratio in the population, so you can't derive a true incidence rate from it.
2. Is a higher sensitivity always better for a screening test? Not universally — it depends on the consequence of missing a case versus over-diagnosing. For a deadly, treatable disease (e.g., HIV screening), high sensitivity is prioritized to avoid missed cases, accepting more false positives that get filtered out by a confirmatory test.
3. What's the practical difference between statistical significance and clinical significance? A large trial can find a statistically significant (p<0.05) difference that is clinically trivial (e.g., a 0.5 mmHg blood pressure drop), while a small trial might show a clinically important effect that fails to reach significance due to low power. Always look at the effect size and CI, not just the p-value.
4. Why is a cohort study considered stronger evidence than a case-control study? Because it establishes clear temporal sequence (exposure measured before outcome occurs) and directly measures incidence, reducing recall bias — case-control studies rely on participants accurately remembering past exposures, which is prone to recall bias, especially in cases who search harder for an explanation for their illness.
5. Does a 95% confidence interval mean there's a 95% chance the true value lies in that range? Not exactly, though it's commonly taught that way for practical exam purposes. Strictly, it means that if you repeated the sampling process many times, 95% of the intervals so constructed would contain the true population parameter. For exam purposes, interpreting it as "the range in which we are 95% confident the true value lies" is accepted.
Quick Revision
- Cross-sectional = prevalence, one time point, no causation.
- Case-control = starts with disease, looks backward, reports odds ratio; best for rare diseases.
- Cohort = starts with exposure, follows forward, reports relative risk; best for rare exposures.
- RCT = random allocation, controls confounding (known and unknown), strongest for causation.
- Prevalence ≈ Incidence × average duration of disease.
- RR/OR = 1 → no association; >1 → increased risk; <1 → protective.
- Sensitivity/specificity are intrinsic to the test; PPV/NPV depend on prevalence.
- SnNout: high Sensitivity, Negative result → rules Out. SpPin: high Specificity, Positive result → rules In.
- p < 0.05 → statistically significant by convention; means data this extreme is unlikely under the null, not that the hypothesis is proven.
- A CI for RR/OR that includes 1.0 (or for a mean difference that includes 0) = not statistically significant.
- Recall bias and selection bias are the classic weaknesses of case-control studies.
- NNT = 1/Absolute Risk Reduction; lower NNT means a more effective intervention.
Related Topics
Prerequisites: Basic probability and proportions; population health concepts (demography, health indicators).
Related Topics: Bias and confounding in epidemiological studies; screening and diagnostic test evaluation; disease surveillance systems; vital statistics and demographic indicators.
Next Topics: Outbreak investigation methodology; sampling techniques and sample size calculation; systematic reviews and meta-analysis; principles of health program evaluation.