What happens when schools stop testing depends on which assessment they remove. Ending a high-stakes exit exam can reduce the influence of one result over graduation. Ending low-stakes classroom checks can remove useful feedback and retrieval practice. Ending system-wide standardised assessment can reduce comparable information about schools and student groups.
A short recall quiz, a final school examination, and a national monitoring assessment serve different purposes. Treating them as one activity creates a false choice between constant testing and no assessment. A more useful question is what each test measures, which decisions depend on it, what pressure it creates, and how its function will be replaced.
The OECD’s student-assessment framework supports a balance between formative and summative purposes and between teacher-developed and external evidence. This approach recognises that one assessment method cannot provide all the information needed for classroom teaching, student certification, and system monitoring.
Answer Summary: Removing selected high-stakes tests may reduce test-based barriers, concentrated pressure, and incentives to narrow teaching. Removing all assessment is different: schools may lose feedback, certification evidence, and comparable system information. A balanced lower-stakes system retains formative checks, uses several demonstrations of learning, moderates teacher judgement, and preserves proportionate external monitoring.
Table of Content
- What Does Testing Mean in Schools?
- What Changes When Tests Are Removed?
- What Does the Evidence Say About High-Stakes Testing?
- What Happens Without Low-Stakes Checks?
- What Can Replace Exams or Standardised Tests?
- What Are the Trade-Offs Between Assessment Alternatives?
- Who Gains, and Who May Be Overlooked?
- What Does a Balanced Lower-Stakes System Look Like?
- Six Questions to Ask Before Removing a Test
- Claims the Evidence Does Not Support
- Bottom Line
Key Takeaways:
-
Removing one test is not the same as ending assessment.
-
High-stakes tests can influence teaching, progression, and student pressure.
-
Low-stakes quizzing can support learning when used appropriately.
-
Teacher assessment offers broader evidence but requires moderation.
-
Projects and portfolios introduce workload and comparability concerns.
-
Common system data need a replacement when standardised tests are removed.
-
No single assessment method serves every educational purpose.
What Does Testing Mean in Schools?
School testing includes several forms of assessment with different purposes, stakes, and consequences. Understanding these distinctions prevents conclusions about one type of test from being applied to every form of assessment.
| Assessment type | Main purpose and stakes | What removal changes |
|---|---|---|
| Low-stakes classroom checks | Feedback, diagnosis, and learning practice | Teachers lose immediate information unless questioning, observation, or other checks replace it. |
| School exams and grades | Reporting, promotion, and course decisions | Schools rely more on coursework, projects, portfolios, or teacher judgement. |
| Exit or certification exams | Graduation, selection, or credentials | A test-based gate is reduced, but another basis for certification is required. |
| System-wide standardised assessment | Monitoring and comparison | Comparable information weakens unless another monitoring method is introduced. |
Assessment, Testing, and Related Terms
Assessment is the process of gathering and interpreting evidence about learning. A test is one tool within that broader process.
Formative assessment is used while learning is developing so that teachers and students can decide what to do next. Summative assessment records or certifies achievement after a period of learning. Diagnostic assessment identifies a student’s starting point or particular learning needs.
A standardised assessment uses centrally determined or consistently applied tasks, administration procedures, or scoring arrangements. A high-stakes assessment has important consequences for a student, teacher, school, or institution.
These features are separate. A test can be standardised without carrying high stakes, while a locally designed assessment can still determine an important outcome. The distinction is explained further in Collegenp’s overview of standardised and nonstandardised assessments.
Low-Stakes Classroom Checks
Low-stakes checks include short quizzes, oral prompts, exit tickets, and practice problems that have little or no effect on a final grade. Some also function as retrieval practice by asking students to recall information rather than reread it.
A systematic and meta-analytic review of classroom quizzing combined evidence from 48,478 students across 222 independent studies. It reported a medium average benefit from quizzing, with an overall effect size of (g = 0.499). The finding concerns classroom testing used as a learning activity; it does not show that high-stakes examinations improve learning.
This distinction matters because a brief ungraded recall activity has a different purpose from an examination that determines graduation. Schools can reduce high-stakes consequences while retaining useful retrieval practice in daily lessons.
School Exams and Grades
School examinations and grades summarise attainment for reporting, promotion, placement, and communication with families or later institutions. Removing an examination does not remove these decisions. It shifts the evidence toward coursework, teacher judgement, common tasks, portfolios, or performance demonstrations.
This shift may provide a broader picture of performance, but it can also increase variation between teachers, classes, or schools. A replacement system therefore needs explicit standards, shared criteria, and review procedures.
High-Stakes Exit or Certification Exams
Exit and certification examinations determine or contribute to decisions such as graduation, credential award, selection, or progression. Removing such an examination reduces the chance that one testing event alone blocks a student.
The education system must still define how students demonstrate the required standard. Possible alternatives include completed coursework, common performance tasks, moderated grades, competency demonstrations, or a combination of evidence.
System-Wide Standardised Assessment
System-wide assessments produce common information for monitoring education systems, schools, subjects, or student groups. They may be used to identify trends, compare outcomes, or examine whether differences between groups are narrowing or widening.
External monitoring does not always require testing every student. A system may use representative samples, inspections, moderated common tasks, or a broader set of indicators. These approaches reduce reliance on universal testing but do not produce identical information.
What Changes When Tests Are Removed?
The first effects usually appear in teaching priorities, student pressure, feedback routines, progression decisions, and the information available to education authorities.
Teaching Time and Curriculum Priorities
High-stakes tests can influence what schools teach because educators have incentives to focus on assessed subjects, formats, and scoring rules.
A qualitative metasynthesis of 49 studies found recurring changes in curriculum content, the form of knowledge taught, and classroom pedagogy under high-stakes testing. The patterns were not uniform: some studies reported narrowing and more test-centred instruction, while a minority found expansion or more integrated teaching under particular test structures. The findings are therefore context-dependent rather than universal.
Reducing high-stakes consequences may create more room for discussion, practical work, extended writing, arts, or inquiry. It does not automatically produce a broader curriculum. Schools still need clear learning standards, appropriate resources, teacher preparation, and evidence that students are progressing.
Readers examining this issue in more detail can refer to Collegenp’s discussion of high-stakes testing and student outcomes.
Feedback and Early Detection
Removing low-stakes checks can reduce teachers’ access to timely evidence about student understanding. Misconceptions may remain hidden until a later assignment, project, or course exposes them.
Schools can gather formative evidence through:
-
targeted questions during lessons;
-
student explanations and demonstrations;
-
annotated drafts and work samples;
-
observation notes and conferences;
-
brief diagnostic activities;
-
peer and self-assessment supported by clear criteria.
These methods support assessment as a basis for better school learning only when teachers interpret the evidence and adjust instruction. Collecting information without changing teaching provides little practical value.
Pressure and Motivation
Removing one high-stakes examination may reduce the pressure attached to that event. It does not establish that students will experience less pressure overall.
Pressure can shift from a final examination to repeated coursework, project deadlines, oral presentations, teacher grading, or selection procedures. Some students may prefer several assessed tasks, while others may find continuous assessment more demanding than one final examination.
Assessment design should therefore consider:
-
how frequently students are assessed;
-
how predictable the requirements are;
-
how much each result affects progression;
-
whether conditions are accessible;
-
whether students receive useful feedback;
-
whether support is available before major decisions.
The aim should not be to assume that one format is comfortable for every student. It should be to avoid making a narrow or imperfect measure carry more consequence than it can justify.
Graduation and Progression
Removing an exit examination changes the rules for receiving a credential. It may allow students who have completed required coursework to graduate without passing one additional test.
A 2026 U.S. study of exit-exam removal examined state-level graduation requirements between the 2008–2009 and 2019–2020 school years. Its difference-in-differences analysis estimated that removing exit-exam requirements increased school-level four-year adjusted cohort graduation rates by about three percentage points. The estimated effect was larger for students classified at the time as limited English proficient.
The measured outcome was graduation, not subject knowledge, later educational performance, employment, or credential quality. The study therefore supports the conclusion that exit examinations can act as graduation barriers in the studied U.S. context. It does not show that eliminating every form of school testing improves learning.
Comparability and Accountability
Common assessments create a shared signal, although that signal is incomplete and can be misinterpreted. When it disappears, comparisons across classrooms, schools, regions, or student groups become more difficult.
A system that removes common testing needs another way to check consistency and identify learning gaps. Possible replacements include:
-
sample-based assessments;
-
moderated common tasks;
-
school inspections;
-
reviews of student work;
-
course-completion information;
-
attendance and engagement indicators;
-
surveys of students and teachers.
These measures answer different questions. Attendance data cannot replace evidence of subject knowledge, and a common performance task cannot fully describe school climate. A credible monitoring system must be clear about what each indicator can and cannot show.

What Does the Evidence Say About High-Stakes Testing?
Research does not provide one global answer. Studies examine different jurisdictions, test types, populations, and outcomes, so their findings should not be treated as though they measure the same result.
| Outcome | Evidence | Supported conclusion and limit |
|---|---|---|
| Classroom learning | Meta-analysis of 222 classroom studies | Low-stakes quizzing can support learning; this does not validate high-stakes testing. |
| Curriculum and pedagogy | Qualitative metasynthesis of 49 studies | High-stakes testing can influence curriculum and teaching; effects vary by test structure and context. |
| Graduation | U.S. state-policy study | Removing exit-exam requirements increased measured graduation rates; it did not measure learning. |
| Teacher judgement | Official review of teacher and test-based assessment | Teacher and external assessment have different strengths and sources of error. |
| System monitoring | OECD assessment framework | Systems need a combination of internal and external evidence rather than one universal measure. |
Curriculum Effects Depend on Test Design
Evidence about curriculum mainly concerns systems where tests carry substantial consequences. When school ratings, sanctions, graduation, or staff evaluation depend heavily on results, educators have stronger reasons to focus on what the test rewards.
The central policy issue is therefore not standardisation alone. It is the combination of limited measurement, strong consequences, and insufficient alternative evidence.
A common assessment aligned with a broad curriculum may clarify expectations. A narrow assessment tied to major consequences can produce incentives to concentrate on tested material. These are different policy designs and should not be discussed as though they have identical effects.
Graduation Is Not the Same as Learning
A higher graduation rate may mean that fewer students are blocked by one test after completing their coursework. It does not by itself establish that academic standards rose, fell, or remained unchanged.
When an exit examination is removed, policymakers still need to explain:
-
which course requirements remain;
-
what evidence demonstrates attainment;
-
how teacher judgements are moderated;
-
whether students can appeal decisions;
-
how standards remain understandable across schools.
A policy can reduce unnecessary gatekeeping while retaining explicit expectations. It can also weaken confidence in a credential if the replacement evidence is unclear or inconsistent.
What Happens Without Low-Stakes Checks?
Removing low-stakes checks is different from removing graduation or accountability examinations. It may reduce visible testing, but it can also remove a learning method and an early-warning system.
Testing Can Support Learning
Retrieval practice requires students to recall information from memory. Short quizzes, oral questions, practice prompts, or self-tests can therefore strengthen learning as well as measure it.
The evidence does not mean that more quizzes are always better. Results can depend on:
-
the quality of the questions;
-
the timing of practice;
-
whether students receive feedback;
-
the content being learned;
-
the age and prior knowledge of students;
-
whether the task carries grades or penalties.
A low-stakes check should help students and teachers decide what needs further attention. Once it becomes a frequent source of punishment or ranking, its educational function changes.
Feedback Must Still Come From Somewhere
Schools can operate without formal examinations, but teaching still requires evidence about what students understand.
Observation, questioning, portfolios, projects, oral tasks, and teacher conferences can provide that evidence. Each method also has limitations. Observation may vary between teachers, projects may conceal unequal outside help, and self-assessment may be inaccurate.
Combining several forms of evidence reduces dependence on the weaknesses of any one method.
What Can Replace Exams or Standardised Tests?
No single alternative performs every assessment function. The replacement should match the decision being made, whether that decision concerns classroom feedback, progress reporting, certification, selection, or system monitoring.
| Method | Main strength and use | Limitation and safeguard |
|---|---|---|
| Teacher assessment with common rubrics | Captures repeated performance across different tasks | Requires moderation, exemplars, training, and review routes. |
| Projects and portfolios | Show extended work, application, and revision | Require authorship checks, workload limits, and shared scoring criteria. |
| Oral assessment and demonstrations | Reveal reasoning, communication, and practical performance | Require accessible formats and consistent scoring procedures. |
| Common performance tasks | Produce richer evidence under shared expectations | Require careful design and cross-school moderation. |
| Sample-based external assessment | Tracks system trends while testing fewer students | Cannot provide individual results for every student or school. |
Teacher Assessment and Common Rubrics
Teacher assessment can draw on months of work and a wider range of activities than a short examination. It may include practical tasks, extended writing, improvement over time, and performance that is difficult to capture under timed conditions.
Its main limitation is consistency. Teachers may use different evidence, interpret standards differently, or be influenced by expectations unrelated to the intended learning outcome.
An Ofqual review of teacher and test-based assessment warned against assuming that every difference between teacher judgement and test results proves teacher error. The methods may use different evidence and contain different sources of bias. The review nevertheless identified risks to the dependability of decisions based entirely on teacher assessment.
Moderation can include:
-
comparing samples of student work;
-
discussing scoring decisions;
-
using shared rubrics and examples;
-
reviewing selected judgements;
-
training teachers to apply standards;
-
providing an appeal process.
Projects and Portfolios
Projects and portfolios can show planning, application, revision, and sustained work. They may capture knowledge and skills that are difficult to demonstrate in a short written examination.
They also create access and reliability concerns. Students may have unequal time, equipment, technology, quiet study space, or outside assistance.
Schools can reduce these risks through:
-
staged submissions;
-
supervised components;
-
oral checks;
-
shared scoring criteria;
-
records of revisions;
-
limits on work completed outside school.
Projects are useful when extended application is part of the intended learning. They are less suitable when a decision requires a quick, independently completed demonstration of foundational knowledge.
Oral Assessment and Demonstrations
Oral tasks and demonstrations can assess explanation, communication, live reasoning, or practical performance. They may be appropriate when the intended learning outcome cannot be represented adequately through written answers alone.
For consequential decisions, students need clear criteria, comparable prompts, suitable accessibility adjustments, and a review process. Recorded samples, moderation, or a second assessor may improve consistency.
Sample-Based External Assessment
Sample-based assessment can preserve system-level trend information without requiring every student to take every external test.
This approach can reduce testing exposure and separate system monitoring from individual consequences. It cannot diagnose every student or provide a complete evaluation of every school, so local formative evidence and other quality checks remain necessary.
What Are the Trade-Offs Between Assessment Alternatives?
Assessment methods should be judged by the evidence they produce and the decisions they inform, not by whether they are described as traditional or alternative.
Validity concerns whether an assessment supports the intended interpretation of performance. Reliability concerns consistency across tasks, scorers, or occasions. Comparability concerns whether results can be interpreted across students, schools, or years. Moderation is the process of reviewing and aligning assessment judgements.
| Criterion | Question to ask | Typical trade-off |
|---|---|---|
| Validity | Does the task represent the intended learning? | A narrow test may omit complex skills; a broad project may include unrelated advantages. |
| Reliability | Would another scorer or occasion produce a similar result? | Common examinations support consistency; local judgement requires moderation. |
| Authenticity | Does the task resemble meaningful use of knowledge? | Richer tasks may be more difficult to standardise. |
| Comparability | Can results be interpreted across settings? | Common tasks aid comparison; local tasks reflect context more closely. |
| Bias and access | Who may be advantaged by the format or conditions? | Examinations and coursework create different barriers. |
| Workload | What time and resources are required? | Richer evidence may increase teaching and moderation demands. |
| Stakes | What consequence follows from one result? | High stakes magnify the effect of measurement error. |
No assessment method performs equally well against every criterion. A mixed system uses several forms of evidence so that the limitations of one measure are partly checked by another.
Who Gains, and Who May Be Overlooked?
Reducing high-stakes testing may benefit students who have completed required coursework but are blocked by one examination. It may also give teachers more room to recognise performance across several tasks.
Other students may lose a common external opportunity to demonstrate achievement. Teacher judgement can vary, while projects can favour students with greater access to time, technology, quiet study space, or outside support.
External examinations also have validity and accessibility limitations. Replacing them does not create a neutral system; it introduces a different set of advantages and risks.
Equity therefore depends on more than removing a test. Schools need to examine:
-
who receives support for projects and coursework;
-
whether assessment adjustments are available;
-
whether standards are applied consistently;
-
whether students can challenge decisions;
-
whether group-level learning gaps remain visible;
-
whether the replacement increases teacher workload unevenly.
A system may reduce one form of inequality while creating another. Equity safeguards must be designed into the replacement rather than assumed.
What Does a Balanced Lower-Stakes System Look Like?
A lower-stakes assessment system does not need to eliminate assessment. It can reduce the influence of one result while preserving evidence for teaching, reporting, certification, and monitoring.
Five elements provide a practical structure:
-
Clear learning standards describing what students are expected to know and do.
-
Regular formative evidence that carries limited consequences and informs teaching.
-
Several measures for important progression or certification decisions.
-
Moderation, review, and appeal processes for professional judgement.
-
Proportionate external monitoring through samples, common tasks, or independent checks.
| Level | Evidence that should remain | Main safeguard |
|---|---|---|
| Classroom | Questions, observations, drafts, retrieval tasks, and diagnostics | Use evidence promptly and provide accessible feedback. |
| School | Coursework, projects, common tasks, and teacher judgement | Apply shared criteria, moderation, and review routes. |
| Credential | Several measures linked to explicit standards | Explain how evidence is combined and avoid reliance on one result. |
| System | Sample assessments, inspections, moderated work, and wider indicators | Preserve comparability, subgroup visibility, and transparent use. |
The balance should differ by age, subject, credential, and purpose. A primary classroom needs different evidence from an upper-secondary certification system. A practical subject may require demonstration, while system monitoring needs stable information across time.
Six Questions to Ask Before Removing a Test
Before removing a test, a school or policymaker should ask:
-
What purpose does the test currently serve?
-
Which decisions depend on its result?
-
What pressure, distortion, or barrier is linked to its design?
-
What evidence will replace the information it provides?
-
How will the replacement be moderated and reviewed?
-
How will learning gaps and unequal outcomes remain visible?
These questions distinguish a deliberate change in assessment method from an accidental loss of feedback, certification, or accountability.
A responsible transition should also specify:
-
when the new system begins;
-
which students and courses it covers;
-
how teachers will be prepared;
-
how workload will be managed;
-
how results will be communicated;
-
when the new system will be reviewed.
Claims the Evidence Does Not Support
The available evidence does not support the following universal claims:
-
Students learn more whenever tests are removed.
-
Removing examinations reduces stress for every student.
-
Higher graduation rates prove stronger academic learning.
-
Teacher assessment is free from bias or inconsistency.
-
Projects and portfolios can serve every assessment purpose.
-
Low-stakes quizzes are harmless in every design.
-
External tests are the only way to monitor equity.
-
One assessment system suits every jurisdiction and age group.
Each claim merges different assessment types, populations, outcomes, and contexts. Responsible conclusions keep the method, purpose, stakes, evidence, population, and limitation connected.
Bottom Line
Removing high-stakes tests may reduce gatekeeping, concentrated pressure, and incentives to narrow teaching. Removing low-stakes checks as well can weaken feedback and retrieval practice. Ending system-wide assessment can reduce comparable information unless another monitoring method replaces it.
The evidence supports a replacement-system approach rather than a choice between constant testing and no assessment. Schools can combine fewer high-stakes consequences with formative feedback, multiple demonstrations of learning, moderated teacher judgement, and limited external checks.
The fairness and quality of the replacement system matter as much as the test being removed.
Education