Skip to main content

Experimental Design in Psychology

Learning Objectives

By the end of this topic, you should be able to:

  • Define experimental design and identify its defining feature (manipulation of an independent variable)
  • Distinguish independent, dependent, and extraneous/confounding variables in a given study
  • Explain how random assignment and control groups protect internal validity
  • Compare between-subjects, within-subjects, and factorial designs, including their trade-offs
  • Identify threats to internal and external validity in a described experiment
  • Apply experimental design concepts to critique or design a simple study

Quick Answer

Experimental design is the blueprint researchers use to test cause-and-effect claims: they deliberately manipulate an independent variable, hold other conditions constant, and measure the effect on a dependent variable. What makes an experiment different from every other method is control — through random assignment, control groups, and standardized procedures, researchers rule out alternative explanations for their results. This is why experiments are the only method that can support causal claims, and why so much of experimental design is really about anticipating and eliminating confounds before they happen, rather than explaining them away afterward.

What Makes a Design "Experimental"

A study earns the label "experiment" only when the researcher actively manipulates something and randomly assigns participants to conditions. If either piece is missing — the researcher observes rather than manipulates, or assigns groups non-randomly (e.g., comparing existing smokers to existing non-smokers) — it's a quasi-experiment or correlational study, not a true experiment, and it cannot support the same causal conclusions.

The essential building blocks:

  • Independent variable (IV) — the factor the researcher manipulates (e.g., dose of caffeine).
  • Dependent variable (DV) — the outcome measured, presumed to be affected by the IV (e.g., memory recall score).
  • Extraneous variables — any other factor that could affect the DV; if not controlled, one becomes a confound that offers an alternative explanation for the results.
  • Control group — a comparison group that does not receive the treatment, establishing a baseline against which the treatment group's results are judged.

Protecting Internal Validity

Internal validity is the question "did the IV actually cause this change in the DV, or is something else responsible?" Every design choice below exists to protect it.

Random assignment distributes participant differences (motivation, prior experience, personality) evenly across groups by chance, so that, on average, the only systematic difference between groups is the treatment itself. Without it, pre-existing group differences become a plausible rival explanation for any result.

Standardized procedures ensure every participant in a condition receives an identical experience except for the manipulated variable itself — same instructions, same room, same timing.

Blinding — in a single-blind study, participants don't know which condition they're in (preventing expectation effects); in a double-blind study, the experimenter doesn't know either (preventing experimenter bias from leaking into how they administer the study or interpret ambiguous behavior).

Manipulation checks verify the IV actually worked as intended — for example, asking participants to rate how anxious a stressor made them feel, to confirm the anxiety manipulation actually produced anxiety before checking its downstream effects.

External Validity: The Trade-Off

Every control added to boost internal validity typically costs some real-world realism. A tightly controlled lab study of aggression using a button-press task tells you little about how aggression unfolds in a bar fight. Factors that affect external validity include the sample's demographics, whether the setting is a lab or the field, and how artificial the manipulated stimulus is compared to real life. Good experimental design is a constant negotiation between these two forms of validity — not a checklist where you maximize both independently.

Types of Experimental Designs

Between-subjects design — each participant experiences only one condition (e.g., either the caffeine group or the placebo group). Random assignment across groups is the main defense against confounds. It avoids carryover effects entirely but requires more participants to reach the same statistical power as a within-subjects design.

Within-subjects (repeated measures) design — the same participants go through every condition. This dramatically reduces the impact of individual differences, since each person acts as their own control, and it needs fewer participants. Its main threat is order effects: practice, fatigue, or boredom from earlier conditions bleeding into later ones. Researchers control this using counterbalancing — systematically varying the order conditions are presented in across participants.

Factorial design — two or more independent variables are manipulated simultaneously (e.g., exercise level × meditation, each with two levels, producing a 2×2 design). Factorial designs can reveal an interaction effect — where the effect of one IV depends on the level of another — which neither IV studied alone could ever show. This is often the most exam-relevant design concept, because interaction effects are conceptually tricky and frequently tested.

Real-World Applications

Pharmaceutical trials use between-subjects, double-blind, placebo-controlled designs to test whether a drug's effect is real and not a placebo response. Educational researchers use within-subjects or factorial designs to isolate which combination of teaching methods produces the best learning outcomes. UX researchers running A/B tests on websites are, in effect, running simple between-subjects experiments — one group sees version A, another version B, and the "dependent variable" is a click-through or conversion rate. The logic taught in this chapter is the same logic underlying evidence-based practice across medicine, education, and product design.

Key Terms

TermDefinitionRelated Concept
Independent Variable (IV)The variable manipulated by the researcherDependent Variable
Dependent Variable (DV)The outcome measured to assess the IV's effectIndependent Variable
ConfoundAn uncontrolled variable that offers an alternative explanation for the resultsExtraneous Variable
Random AssignmentRandomly placing participants into conditions to equalize groupsInternal Validity
Control GroupThe comparison group that does not receive the experimental treatmentTreatment Group
Double-BlindNeither participant nor experimenter knows the assigned conditionSingle-Blind, Experimenter Bias
Manipulation CheckA measure confirming the IV was successfully implementedValidity
Between-Subjects DesignDifferent participants experience each conditionWithin-Subjects Design
Within-Subjects DesignThe same participants experience all conditionsOrder Effects, Counterbalancing
CounterbalancingSystematically varying condition order to control order effectsWithin-Subjects Design
Factorial DesignA design manipulating two or more IVs at onceInteraction Effect
Interaction EffectWhen the effect of one IV depends on the level of another IVFactorial Design

Common Mistakes

Misconception: Any study that compares two groups is an experiment. Why it's wrong: Comparing pre-existing groups (e.g., men vs. women, smokers vs. non-smokers) without random assignment is a quasi-experimental or correlational comparison, not a true experiment — you cannot rule out that the groups already differed in relevant ways before the study began. Correct understanding: A true experiment requires both manipulation of the IV and random assignment to conditions. Without random assignment, causal claims are much weaker.


Misconception: A within-subjects design is always better because it needs fewer participants. Why it's wrong: Fewer participants is an advantage, but within-subjects designs introduce order effects (fatigue, practice, boredom) that between-subjects designs don't have, and they aren't suitable when experiencing one condition would permanently change how a participant responds to another (e.g., you can't "un-teach" a skill before testing the control condition). Correct understanding: The choice depends on the research question — within-subjects works well when order effects can be controlled via counterbalancing and conditions are non-permanent; otherwise between-subjects is safer.


Misconception: An interaction effect just means both IVs had an effect. Why it's wrong: Two IVs can each have significant main effects with no interaction at all (their effects simply add up independently), or one IV might have no main effect alone yet still interact meaningfully with the other. Correct understanding: An interaction effect specifically means the effect of one IV changes depending on the level of the other IV — e.g., exercise reduces stress only when combined with meditation, but not alone. This is a distinct statistical concept from a main effect.

Comparison and Connections

DesignParticipants per ConditionMain RiskBest Used When
Between-SubjectsDifferent peopleIndividual differences between groupsConditions would permanently affect participants
Within-SubjectsSame peopleOrder/practice/fatigue effectsNeed higher power with fewer participants
FactorialSame or different, per cellComplexity of interpreting interactionsTesting how two+ factors combine

Practice Questions

Recall

  1. Define independent variable and dependent variable, and give one example of each in a memory experiment. Answer guidance: IV = manipulated factor (e.g., amount of sleep); DV = measured outcome (e.g., number of words recalled).

  2. What is the purpose of random assignment in an experiment? Answer guidance: It distributes participant differences evenly across groups by chance, protecting internal validity by making groups comparable before treatment.

Understanding

  1. Explain why a double-blind procedure controls for more sources of bias than a single-blind procedure. Answer guidance: Single-blind only removes participant expectation effects; double-blind additionally removes experimenter bias in administering the study or interpreting results.

  2. Why do within-subjects designs require counterbalancing? Answer guidance: Because experiencing conditions in a fixed order risks practice, fatigue, or carryover effects; counterbalancing varies the order across participants so these effects cancel out rather than systematically favoring one condition.

Application

  1. A researcher wants to test whether a new study technique improves exam scores but is worried students who try it first will get better at test-taking generally, boosting their second-condition scores regardless of the technique. What design problem is this, and how should it be addressed? Answer guidance: This is an order/practice effect in a within-subjects design; address it using counterbalancing (half the group tries the new technique first, half tries the old method first).

  2. A company tests a 2×2 design: workspace type (open vs. private) × task type (creative vs. analytical) on productivity. They find open offices boost creative task performance but hurt analytical task performance. What statistical concept describes this pattern? Answer guidance: This is an interaction effect — the effect of workspace type depends on the level of task type.

Analysis

  1. A study compares anxiety levels between people who meditate regularly and those who don't, without assigning anyone to meditate. Analyze why this cannot be called a true experiment and what causal claims are and aren't justified. Answer guidance: No manipulation or random assignment occurred — meditators may differ systematically (e.g., pre-existing personality, lifestyle) before the study, so only a correlational claim is justified, not a causal one.

  2. Compare the trade-offs a researcher faces choosing between a highly controlled lab experiment and a more naturalistic field experiment when studying bystander intervention. What does each design sacrifice? Answer guidance: Lab experiments maximize internal validity/control but sacrifice external validity/realism; field experiments maximize ecological realism but sacrifice control over confounds, making causal interpretation less certain.

FAQ

Why is random assignment considered more important than a large sample size? Random assignment addresses a different problem than sample size does. It equalizes groups on both known and unknown confounding variables before the study starts, which is essential for causal inference. A large sample without random assignment can still be badly confounded — size improves precision, not causal validity. Both matter, but random assignment is what makes a design a true experiment in the first place.

Can a study have high internal validity but low external validity at the same time? Yes, and this is extremely common. A tightly controlled lab study with a homogeneous sample (e.g., undergraduates) can very confidently establish that the IV caused the DV within that setting, while still being questionable in how well that finding applies to different populations, environments, or real-world stimuli. Researchers manage this by later running field or replication studies to test generalizability.

How do researchers decide how many independent variables to include in a factorial design? Practically, complexity: each added IV multiplies the number of conditions and can make results harder to interpret, especially with three-way or higher interactions. Researchers include a second IV when there's a genuine theoretical reason to expect that it might change the effect of the first IV — not simply because more variables seem more thorough.

Is a quasi-experiment less trustworthy than a true experiment? It's less definitive for causal claims specifically, because it lacks random assignment, but it's often the only ethical or practical option (e.g., studying the effects of divorce on children, since you cannot randomly assign families to divorce). Quasi-experiments are a legitimate and common tool — they simply require more caution in causal language and more effort statistically to rule out alternative explanations.

What's the difference between a confound and an extraneous variable? Every confound is an extraneous variable, but not every extraneous variable is a confound. An extraneous variable is any factor besides the IV that could affect the DV. It becomes a confound specifically when it varies systematically alongside the IV — meaning it offers a genuine alternative explanation for the result, rather than just adding random noise.

Quick Revision

  • An experiment requires both manipulation of an IV and random assignment to conditions
  • IV = what's manipulated; DV = what's measured; confound = uncontrolled variable that varies with the IV
  • Random assignment equalizes groups on both known and unknown variables, protecting internal validity
  • Double-blind designs remove bias from both participants and experimenters
  • Between-subjects: different people per condition, no carryover, needs more participants
  • Within-subjects: same people in all conditions, more power, but risks order effects — controlled via counterbalancing
  • Factorial designs manipulate 2+ IVs and can reveal interaction effects that single-IV studies can't show
  • An interaction effect means one IV's impact depends on the level of another IV
  • Internal validity and external validity trade off against each other — more control usually means less realism
  • Quasi-experiments lack random assignment and support weaker causal claims than true experiments

Prerequisites

  • Introduction to Research Methods
  • Basic understanding of variables and hypotheses

Related Topics

  • Data Collection Methods
  • Statistical Analysis
  • Research Ethics

Next Topics

  • Data Collection Methods (how DVs are actually measured in practice)
  • Statistical Analysis (t-tests, ANOVA for analyzing experimental data)
  • Research Ethics (informed consent and risk in experimental manipulation)