Skip to main content

Introduction to Research Methods in Psychology

Learning Objectives

By the end of this topic, you should be able to:

  • Define research methods and explain why psychology relies on systematic, empirical procedures rather than intuition
  • Distinguish between experimental, correlational, and descriptive approaches to studying behavior
  • Explain the difference between validity and reliability, and identify the main types of each
  • Compare sampling techniques (random, stratified, convenience) and their effect on generalizability
  • Identify the core ethical safeguards that apply across all psychological research
  • Apply research-method vocabulary to evaluate a real study description

Quick Answer

Research methods in psychology are the systematic procedures scientists use to collect data, test hypotheses, and draw defensible conclusions about behavior and mental processes. They matter because psychology makes claims about real people's lives — in therapy, education, policy, and workplaces — and those claims need to survive scrutiny. A method is only as good as its ability to answer two questions: does it actually measure what it claims to measure (validity), and would it produce the same result again under similar conditions (reliability)? Everything else in the subject — sampling, design, ethics, statistics — exists to protect those two properties.

Why Psychology Needs a Method at All

Common sense about behavior fails constantly. People believe eyewitnesses are reliable, that venting anger reduces it, that opposites attract — all popular beliefs that controlled research has overturned or heavily qualified. The reason psychology insists on formal methods is that human intuition is a poor judge of human behavior: we notice patterns that confirm what we already believe (confirmation bias), we misremember our own past states, and we generalize from small, unrepresentative slices of experience (our friends, our culture, our mood that day).

A research method is a discipline for removing that noise. It forces the researcher to state a testable question in advance, define exactly how each variable will be measured, decide who counts as a participant before collecting data, and specify what result would count as support versus disconfirmation. None of this guarantees truth — but it makes error visible and correctable, which untested opinion never does.

Types of Research Methods

Every method below trades off control for realism, or depth for breadth. Recognizing which trade-off a method makes is the actual skill being tested in exams, not memorizing the list.

Experimental method — the researcher manipulates an independent variable and measures its effect on a dependent variable, while holding other factors constant. This is the only method that can establish cause and effect, because random assignment rules out most alternative explanations. Example: giving one group caffeine and a control group a placebo, then testing memory recall.

Correlational method — the researcher measures two or more variables as they naturally occur, without manipulating anything, and calculates how strongly they move together. Useful when manipulation is impossible or unethical (you cannot randomly assign people to smoke for ten years), but it cannot establish causation. Example: measuring the relationship between hours of TV watched and BMI.

Surveys and questionnaires — structured self-report instruments that gather data from large numbers of people efficiently. They rely on participants accurately knowing and honestly reporting their own thoughts, which is their central weakness. Example: a standardized personality inventory administered online.

Case studies — an intensive, detailed examination of a single person or small group, often used when a phenomenon is rare. They generate rich hypotheses but cannot be generalized to a population. Example: H.M.'s case, which shaped our understanding of memory and the hippocampus.

Naturalistic observation — watching and recording behavior in its real-world setting without interference, maximizing ecological validity at the cost of experimental control. Example: recording children's social interactions on a playground.

Archival research — analyzing existing records (crime statistics, hospital admissions, historical documents) to answer new questions without collecting new data. Efficient, but limited to whatever was originally recorded and for whatever purpose.

Physiological and neuropsychological measures — recording biological responses (heart rate, EEG, fMRI) or performance on standardized cognitive tests to index psychological states objectively, bypassing self-report entirely.

Validity and Reliability

These two concepts are the backbone of every methods chapter, and exam questions love testing whether students can tell them apart.

Validity asks: is this study measuring what it claims to measure? A depression questionnaire that actually captures general anxiety has a validity problem, no matter how consistently it produces scores.

  • Internal validity — confidence that the independent variable, and not some confound, caused the observed change.
  • External validity — confidence that the results generalize beyond the specific sample, setting, and time of the study.
  • Construct validity — confidence that the operational definition (e.g., "number of items recalled") truly represents the abstract concept (e.g., "memory").

Reliability asks: would this produce the same result again? A bathroom scale that reads three different weights in three minutes is unreliable — regardless of whether it's accurate.

  • Test-retest reliability — consistency of scores when the same test is given to the same people at two points in time.
  • Inter-rater reliability — consistency between two or more observers coding the same behavior.

A measure can be reliable without being valid (a scale that consistently reads five pounds too heavy is reliable but not valid), but it cannot be valid without being at least somewhat reliable — an instrument that gives random results can't be accurately measuring anything.

Sampling Methods

How participants are selected determines how far the findings can travel.

  • Random sampling — every member of the population has an equal chance of selection, which is the gold standard for generalizability but is expensive and often impractical.
  • Stratified sampling — the population is divided into meaningful subgroups (age, gender, income) and participants are randomly sampled within each, ensuring the sample mirrors the population's composition.
  • Convenience sampling — participants are selected because they are easy to access (commonly, university students). Cheap and fast, but it risks a sample that doesn't represent the broader population — a limitation psychology has been criticized for, since so much research relies on "WEIRD" samples (Western, Educated, Industrialized, Rich, Democratic).

Ethical Considerations (Preview)

Every method above operates inside ethical boundaries: informed consent, confidentiality, and debriefing. These principles are covered in depth in the Research Ethics chapter, but they apply from the moment a study is designed, not just when data collection begins.

Real-World Applications

Research methods aren't just an academic hurdle — they shape decisions with real consequences. Clinical psychologists rely on validated assessment tools to diagnose accurately. Policymakers use correlational and longitudinal data to decide where to allocate mental-health funding. Companies use surveys and controlled A/B-style experiments to evaluate whether a workplace wellness program actually improves outcomes. Being able to read a study and ask "was this correlational or experimental? was the sample representative?" is a practical skill far beyond the psychology classroom.

Key Terms

TermDefinitionRelated Concept
Research MethodA systematic procedure for collecting and analyzing data to answer a research questionScientific Method
ValidityThe extent to which a study measures what it claims to measureInternal/External/Construct Validity
ReliabilityThe consistency of a measurement across repeated trialsTest-Retest, Inter-Rater Reliability
Internal ValidityConfidence that the IV, not a confound, caused the observed effectExperimental Method
External ValidityThe degree to which findings generalize beyond the study sampleSampling, Ecological Validity
Construct ValidityThe degree to which an operational definition captures the true conceptOperationalization
Random SamplingA sampling method giving every population member an equal chance of selectionGeneralizability
Stratified SamplingSampling that preserves population subgroup proportionsRepresentative Sample
Convenience SamplingSelecting participants based on ease of accessSampling Bias
Correlational MethodMeasuring the relationship between variables without manipulationCorrelation, Causation
Case StudyAn in-depth examination of a single individual or small groupExternal Validity (low)

Common Mistakes

Misconception: A large sample size automatically makes a study valid. Why it's wrong: Sample size affects statistical power and precision, but it says nothing about whether the sample was selected in a biased way or whether the measurement itself captures the right construct. A huge, poorly sampled or badly measured study can still produce misleading conclusions. Correct understanding: Validity depends on sound measurement and representative sampling; sample size mainly helps reduce random error and increases the reliability of statistical estimates.


Misconception: "Reliable" and "valid" mean roughly the same thing. Why it's wrong: Students often use them interchangeably, but a test can be highly consistent (reliable) while measuring the wrong thing entirely (not valid) — like a scale that's off by five pounds every single time. Correct understanding: Reliability is about consistency; validity is about accuracy in measuring the intended construct. Reliability is necessary but not sufficient for validity.


Misconception: Correlational research is a "weaker" or lesser version of experimental research. Why it's wrong: Correlational designs aren't inferior — they're the only ethical or practical option for many important questions (e.g., studying the effects of childhood trauma, since you cannot ethically assign trauma to a group). Correct understanding: Each method fits a different type of question. Correlational research excels at documenting real-world relationships; experiments excel at isolating cause. Choosing the right tool for the question is the mark of good research, not a hierarchy of quality.

Comparison and Connections

MethodManipulates Variables?Establishes Causation?Typical StrengthTypical Weakness
ExperimentalYesYesHigh internal validityCan lack real-world realism
CorrelationalNoNoStudies real, unmanipulable relationshipsCannot rule out third variables
SurveyNoNoReaches large samples efficientlyRelies on honest self-report
Case StudyNoNoRich, in-depth detailCannot generalize
Naturalistic ObservationNoNoHigh ecological validityLow control, observer bias risk

Practice Questions

Recall

  1. What is the difference between validity and reliability? Answer guidance: Validity = measuring what you intend to measure; reliability = getting consistent results across repeated measurements. Give one example of each.

  2. Name the three types of validity discussed in this chapter. Answer guidance: Internal validity, external validity, construct validity — with a one-line definition of each.

Understanding

  1. Why can't correlational research establish causation, even with a very strong correlation coefficient? Answer guidance: A strong correlation could reflect reverse causation or a third (confounding) variable; only controlled manipulation and random assignment can rule these out.

  2. Explain why convenience sampling threatens external validity even if the study itself is well-designed. Answer guidance: If the sample (e.g., psychology undergraduates) differs systematically from the population of interest, findings may not generalize, regardless of how rigorous the internal design was.

Application

  1. A company wants to know if a four-day work week improves employee well-being. It cannot randomly assign employees to different schedules. What method should it use, and what limitation should it report? Answer guidance: A correlational or quasi-experimental design comparing existing groups; the limitation is inability to fully rule out confounds like differing job types or self-selection into the schedule.

  2. A researcher develops a new "test anxiety" scale. On Monday and again on Friday, the same students get very different scores despite no real change in anxiety. What property of the test is in question? Answer guidance: Reliability, specifically test-retest reliability — the instrument isn't producing consistent scores over time.

Analysis

  1. Compare an experiment and a case study as ways of investigating the psychological effects of an unusual brain injury. What can each tell you that the other cannot? Answer guidance: A case study can capture unique, unrepeatable detail impossible to ethically experiment on; an experiment would offer causal control but such injuries can't be manipulated, so case studies are the appropriate — not inferior — tool here.

  2. A study finds a strong positive correlation between ice cream sales and drowning deaths. Analyze what's wrong with concluding ice cream causes drowning, and propose the likely explanation. Answer guidance: Classic third-variable problem — hot weather increases both ice cream sales and swimming (and thus drowning risk). Illustrates why correlation cannot be read as causation.

FAQ

Is one research method always "better" than the others? No. Each method is suited to a different kind of question and a different set of constraints. Experiments are the only way to establish causation, but many important questions in psychology (the long-term effects of poverty, cultural differences in parenting) simply cannot be manipulated ethically or practically. Good researchers pick the method that fits the question, and often combine methods — for example, using correlational data to generate a hypothesis and then testing it experimentally.

Why do psychologists care so much about ethics when discussing methods? Because psychology studies people, not inert materials, and flawed ethics can directly and immediately harm participants — emotionally, socially, or physically. Historical failures, like the Stanford Prison Experiment's psychological harm to participants, shaped today's ethical standards. Methods and ethics are taught together because a "successful" study that harmed its participants isn't actually successful.

What does it mean when a study says its findings "generalize"? It means the pattern observed in the sample is likely to hold true in the broader population the sample was meant to represent. Generalizability depends heavily on how the sample was selected (random vs. convenience) and how similar the sample is to the group you want to apply the findings to.

If reliability doesn't guarantee validity, why measure it at all? Because reliability is a prerequisite — you cannot trust a measurement that gives you a different answer every time you use it, even if you suspect it's roughly accurate. Establishing reliability first is standard practice before ever claiming a measure is valid; think of it as clearing the minimum bar before the real evaluation begins.

Are surveys considered a weak method because they rely on self-report? Self-report has real limitations (social desirability bias, poor self-insight), but surveys are indispensable for measuring internal states — thoughts, feelings, attitudes — that can't be directly observed. Their weakness is a trade-off, not a disqualification; researchers manage it through careful item wording, anonymity, and cross-checking with other data sources.

Quick Revision

  • Research methods exist because intuition and anecdote are unreliable guides to human behavior
  • Experimental method: manipulates variables, only method that can show causation
  • Correlational method: measures naturally occurring relationships, cannot show causation
  • Validity = measuring the right thing; reliability = measuring consistently
  • Internal validity = confidence in cause-effect within the study; external validity = generalizability beyond it
  • Construct validity = the operational definition truly captures the abstract concept
  • Reliability is necessary but not sufficient for validity
  • Random sampling gives every person an equal chance of selection and best supports generalization
  • Convenience sampling is fast and common but threatens external validity
  • Case studies and naturalistic observation trade generalizability for depth and real-world realism
  • Third-variable and reverse-causation problems are the reason correlation ≠ causation
  • All research methods operate under core ethical principles: informed consent, confidentiality, debriefing

Prerequisites

  • Introduction to Psychology
  • Basic Statistics/Mathematics Literacy

Related Topics

  • Experimental Design
  • Data Collection Methods
  • Statistical Analysis

Next Topics

  • Experimental Design (independent/dependent variables, control groups, design types)
  • Data Collection Methods (surveys, interviews, observation in depth)
  • Research Ethics (informed consent, IRBs, protection of participants)