Skip to main content

Assessment and Evaluation in Educational Psychology

Learning Objectives

By the end of this page, you should be able to:

  • Distinguish assessment from evaluation, and formative from summative assessment.
  • Explain criterion-referenced and norm-referenced evaluation and when each is appropriate.
  • Identify common formative and summative assessment tools used in classrooms.
  • Describe key ethical considerations in designing and administering assessments.
  • Apply these concepts to design or critique an assessment strategy for a realistic classroom scenario.

Quick Answer

Assessment is the process of gathering information about what students know and can do; evaluation is what happens next — interpreting that information to make a decision, such as adjusting instruction or assigning a grade. The most important distinction to master is formative versus summative assessment: formative assessment happens during learning to guide instruction (like an exit ticket), while summative assessment happens after learning to measure the final result (like a final exam). A second key distinction is criterion-referenced evaluation (comparing a student against a fixed standard) versus norm-referenced evaluation (comparing a student against peers). These distinctions matter because using the wrong type of assessment for a given purpose — say, only ever summative testing with no formative checks — leaves teachers unable to catch and correct misunderstandings while there's still time to help.

Types of Assessments

Educational psychologists and teachers gather information about learning in two broad ways, distinguished mainly by when they happen and what they're for.

Formative Assessments

Definition: Formative assessments are ongoing evaluations used during the learning process to monitor student progress and adjust instruction in response.

Explanation: The defining feature of formative assessment isn't the format — it's the purpose. A quiz can be formative if its main function is to reveal what needs re-teaching, not to produce a final grade. Formative assessment closes the loop between teaching and learning while there's still time to act on what's discovered.

Example: An exit ticket at the end of class asking students to write down one thing they understood and one thing that's still confusing.

Real-world example: A teacher notices from a think-pair-share activity that most of the class still can't correctly identify the independent variable in an experiment, and adjusts the next day's lesson to re-teach that specific concept before moving on.

Why it matters: formative assessment is what allows teaching to be responsive rather than fixed — without it, a teacher only discovers a misunderstanding once it's too late to fix before the final grade is set.

Common misunderstanding: students often think "formative" just means "low-stakes" or "ungraded." What actually makes an assessment formative is that its results are used to adjust instruction before final judgment is made — a graded quiz can still be formative if the teacher uses the results to re-teach a concept.

Summative Assessments

Definition: Summative assessments evaluate student learning at the end of a lesson, unit, or course, providing a comprehensive measure of what has been learned.

Explanation: Where formative assessment asks "how is learning going, and what should change?", summative assessment asks "how much was learned, overall?" This makes summative assessment the natural fit for final grades and high-stakes decisions, but it offers no opportunity to intervene before the outcome is recorded.

Example: A standardized end-of-course exam covering the full semester's material.

Real-world example: A final research paper that requires students to synthesize everything learned in a unit into a single, graded product.

Why it matters: summative assessments provide the accountability and certification function of education — they're how achievement gets documented for report cards, transcripts, and college admissions.

Common misunderstanding: students sometimes think summative assessment is inherently "worse" or purely punitive compared to formative assessment. Both serve legitimate, different purposes — a course built only on formative assessment would have no mechanism for certifying that learning happened, which matters for transcripts, licensure, and academic integrity.

Evaluation Techniques

Evaluation is a distinct step from assessment: it's the interpretation of assessment data to make a decision. Two major approaches determine what a score is compared against.

Criterion-Referenced Evaluation

Definition: Criterion-referenced evaluation compares a student's performance against a predetermined, fixed standard or set of criteria — not against other students.

Explanation: The question being answered is "did this student meet the standard?" rather than "how did this student rank compared to peers?" This makes it well suited to mastery-based learning, where the goal is for every student to reach a defined competency.

Example: A math teacher checks whether a student can correctly solve quadratic equations according to a fixed rubric, regardless of how classmates performed.

Real-world example: A driving test is criterion-referenced — you pass if you meet the defined safety standards, regardless of how other test-takers that day performed.

Why it matters: criterion-referenced evaluation is fairer for measuring individual mastery of a skill, since a student's success doesn't depend on how well or poorly their peers happen to perform.

Common misunderstanding: students sometimes assume "criterion-referenced" just means "graded on a rubric." A rubric can be used in either approach — what makes evaluation criterion-referenced is that the standard is fixed in advance and independent of how the group performs.

Norm-Referenced Evaluation

Definition: Norm-referenced evaluation compares an individual student's performance to that of a peer group, typically reporting results as percentiles or rankings.

Explanation: The question being answered is "how does this student compare to others?" This is useful for identifying relative strengths and weaknesses, or for selective processes where only a limited number of spots exist.

Example: A standardized reading comprehension test that reports a student's score as a percentile relative to a national sample of test-takers.

Real-world example: College admissions tests are typically norm-referenced, since they're used partly to compare applicants against one another for a limited number of spots.

Why it matters: norm-referenced evaluation is useful for large-scale comparison and selection decisions, but it doesn't directly tell you whether a student has mastered a specific skill — a student could score above average on a norm-referenced test while still not meeting an absolute competency standard.

Common misunderstanding: students often think a "high percentile" automatically means strong mastery of content. It only means the student outperformed a certain proportion of the comparison group — if the whole group performed poorly, a high percentile could still reflect a low absolute level of mastery.

Assessment Tools and Technologies

Modern classrooms increasingly rely on technology to support both formative and summative assessment. Learning Management Systems (like Canvas, Blackboard, or Moodle) support automated grading, discussion forums, and assignment submission, streamlining routine administrative work. Educational software such as Kahoot and Quizlet make low-stakes formative checks fast and engaging, often gamifying quick recall practice. Online proctoring tools support the integrity of remote summative assessments by reducing opportunities for academic dishonesty.

Why it matters: technology doesn't change the underlying formative/summative distinction — it changes how efficiently data can be gathered and acted on, which matters most for formative assessment, where speed of feedback is part of what makes it effective.

Ethical Considerations in Assessment

Because assessment results shape real decisions about students' futures, educational psychologists emphasize several ethical obligations: avoiding bias in how assessments are designed (so that questions don't unfairly favor a particular cultural or linguistic background), protecting student privacy and confidentiality of results, giving clear instructions and expectations so students aren't penalized for confusion about the format, and offering timely feedback so results can still be used to support learning.

Real-world example: A test that relies heavily on cultural references unfamiliar to English language learners may underestimate their actual content knowledge — not because they lack the skill being tested, but because the format of the question introduces an unrelated barrier. Careful assessment design tries to separate the skill being measured from unrelated obstacles.

Visual: From Assessment to Decision

Key Terms

TermDefinition
AssessmentThe process of gathering information about student learning
EvaluationThe interpretation and use of assessment data to make an educational decision
Formative assessmentOngoing assessment used during learning to monitor progress and guide instruction
Summative assessmentAssessment conducted at the end of a lesson, unit, or course to measure overall learning
Criterion-referenced evaluationEvaluation that compares performance against a predetermined, fixed standard
Norm-referenced evaluationEvaluation that compares an individual's performance to that of a peer group

Common Mistakes

#MisconceptionWhy It's WrongCorrect Understanding
1Formative assessment just means "ungraded" or "low-stakes"What defines formative assessment is that results are used to adjust instruction before a final judgment, regardless of whether it's gradedA graded quiz can still be formative if its purpose is to inform re-teaching, not just to record a score
2A high percentile on a norm-referenced test always means strong mastery of the contentA percentile only reflects standing relative to a specific comparison group, which could itself perform well or poorly overallNorm-referenced scores show relative standing, not absolute mastery — criterion-referenced evaluation is needed to confirm mastery of a specific standard
3Summative assessment is inherently worse or purely punitive compared to formative assessmentBoth serve legitimate, different purposes: formative guides learning in progress, summative certifies and documents final achievementA well-designed course uses both — formative assessment to support learning along the way, summative assessment to certify what was ultimately achieved

Comparison and Connections

FeatureFormative AssessmentSummative Assessment
TimingDuring the learning processAt the end of a lesson, unit, or course
Main purposeGuide and adjust instructionMeasure and certify overall achievement
ExampleExit ticket, think-pair-share, self-assessment rubricFinal exam, standardized test, research paper
StakesTypically lower, though not always ungradedTypically higher, often tied to grades or certification
FeatureCriterion-Referenced EvaluationNorm-Referenced Evaluation
Comparison pointA fixed, predetermined standardA peer group
Answers the questionDid the student meet the standard?How does the student rank compared to others?
Best suited forMastery-based learning, competency checksSelection processes, large-scale comparison
ExampleDriving test, mastery rubric for solving equationsStandardized percentile-based test, college admissions test

Practice Questions

Recall

  1. Define assessment and evaluation, and explain how they differ. Answer guidance: assessment is gathering information about learning; evaluation is interpreting and using that data to make a decision — assessment feeds evaluation.
  2. Give one example each of a formative and a summative assessment tool. Answer guidance: formative — exit ticket, think-pair-share, quiz used to guide re-teaching; summative — final exam, standardized test, research paper.

Understanding

  1. Explain why a graded quiz can still be considered formative assessment. Answer guidance: what makes an assessment formative is that its results are used to adjust instruction before a final judgment is made, not whether it's graded — a graded quiz used to identify what needs re-teaching still functions formatively.
  2. Why doesn't a high percentile on a norm-referenced test guarantee strong absolute mastery of content? Answer guidance: a percentile only reflects a student's standing relative to a specific comparison group; if the whole group performed at a low absolute level, a high percentile could still represent weak mastery of the actual content.

Application

  1. A teacher wants to know, mid-unit, whether students are ready to move on to a harder topic. Which type of assessment should they use, and why? Answer guidance: formative assessment (e.g., an exit ticket or quick quiz) because the purpose is to inform an instructional decision (move on or re-teach) while there's still time to act.
  2. A school needs to decide which of 200 applicants will receive 20 scholarship spots based on test performance. Should the school use criterion-referenced or norm-referenced evaluation, and why? Answer guidance: norm-referenced evaluation, because the decision requires ranking applicants against one another for a limited number of spots, rather than simply checking whether each meets a fixed standard.

Analysis

  1. Compare the ethical risks of using a culturally biased summative test for high-stakes decisions versus using the same test formatively. Answer guidance: bias is a serious concern in both cases, but the stakes differ — a biased summative test can permanently and unfairly affect a student's grade, placement, or future opportunities, while a biased formative assessment, if recognized, can still be corrected through re-teaching or a different assessment approach before final judgment is made; the argument for careful ethical assessment design is stronger when stakes are higher.
  2. Evaluate the claim that a course could rely entirely on formative assessment and skip summative assessment altogether. Answer guidance: a strong answer notes formative-only assessment provides excellent support for learning but offers no clear mechanism for certifying final achievement, which matters for transcripts, licensure, and academic integrity — most educational systems therefore need both types serving their distinct purposes.

FAQ

Is a midterm exam formative or summative? It depends on how it's used — if the midterm's primary purpose is to inform changes to the rest of the course (formative), it functions formatively; if it primarily contributes to a final grade for material already completed, it functions summatively. Many midterms actually serve both purposes.

Can the same test be evaluated using both criterion-referenced and norm-referenced methods? Yes — for example, a research paper could be graded against a fixed rubric (criterion-referenced) while also being ranked among classmates for a scholarship or award (norm-referenced), as long as both purposes are clearly communicated.

Why do ethical considerations matter so much in assessment design? Because assessment results influence real decisions about grades, placement, and opportunities — a biased or unclear assessment can unfairly disadvantage students for reasons unrelated to what's actually being measured.

Do educational technology tools like Kahoot count as "real" assessment? Yes, when used purposefully — their main value is usually formative, since they provide fast, low-stakes data teachers can use immediately to check understanding and adjust instruction.

What's the risk of relying only on norm-referenced evaluation in a classroom? It can obscure whether students have actually mastered the material, since it only shows relative standing — a class could all score "average" while genuinely lacking key skills, which is why criterion-referenced checks are important for confirming actual competency.

Quick Revision

  • Assessment = gathering data on learning; evaluation = interpreting that data to make a decision.
  • Formative assessment happens during learning to guide instruction (exit tickets, quizzes used for re-teaching).
  • Summative assessment happens at the end to measure overall achievement (final exams, standardized tests).
  • What makes an assessment formative is its use, not whether it's graded or "low-stakes."
  • Criterion-referenced evaluation compares performance to a fixed standard (mastery-based).
  • Norm-referenced evaluation compares performance to a peer group (percentile/ranking-based).
  • Norm-referenced scores show relative standing, not guaranteed absolute mastery.
  • Technology tools (LMS, Kahoot, proctoring software) speed up data collection but don't change the underlying formative/summative or criterion/norm distinctions.
  • Ethical assessment design avoids cultural/linguistic bias, protects privacy, and provides timely, clear feedback.
  • Most effective courses combine both assessment types and, where relevant, both evaluation approaches for different purposes.

Prerequisites: Classroom Management

Related Topics: Standardized testing, rubric design, bias in psychological and educational testing

Next Topics: Special Education Needs