What Is Regression Testing? Techniques + a Reusable Checklist

Regression testing explained: risk-based selection, prioritization that works, why flaky tests lie, a reusable checklist, and how to keep suites fast.

Regression testing answers one question after every change: did anything that used to work still work? New features get the attention, demos, and test plans — but the bugs that burn teams are almost always in code nobody touched. The ISTQB defines regression testing as "a type of change-related testing to detect whether defects have been introduced or uncovered in unchanged areas of the software" (see the ISTQB Glossary entry). The phrase "unchanged areas" is the part teams get wrong: you are not retesting the change, you are testing everything the change could have broken around it.

Why regressions happen (and why re-running everything is not the answer)

A regression sneaks in through the seams between modules. Typical mechanisms:

  • Shared code changes. A utility function gets a "safe" tweak; three callers downstream depended on the old edge-case behavior.
  • Refactors with green unit tests. Unit mocks encode the old assumptions, so a broken integration still passes every unit suite.
  • Config, dependency, and data drift. The code is identical but the library version, migration, or feature flag underneath it moved.
  • Fixed bugs that were never folded back into the suite. The bug returns because no test ever existed that would have caught it.

The instinctive answer — re-run the entire suite on every commit — fails economically long before it fails technically. A 4,000-case suite at 30 seconds per case is 33 hours of wall-clock time per pass. By the time the pass finishes, the code has moved again. Practical regression testing is therefore a selection problem: run the right subset, fast, on every change.

Choosing what to run: risk-based selection

Rank every candidate case on two axes — the likelihood the change broke it, and the impact if it silently broke. Then apply a simple selection order:

  • All cases touching changed modules — including their neighbors at integration boundaries, not just the changed functions.
  • All cases for bug fixes shipped in this release. A fixed bug without a regression case is a timed reinstall: fold every fix into the suite permanently.
  • High-traffic, revenue-adjacent flows (login, checkout, signup, search) even when nothing near them changed — these are the failures users find first.
  • Historically flaky or defect-dense areas. Modules that produced bugs last quarter statistically produce them again.

Everything else runs on a rotation — nightly, weekly, or before release only. That rotation is what keeps daily passes inside a coffee break instead of a day.

Ordering the subset: prioritization, not shuffling

Once the subset is chosen, order decides what you learn. Regression passes get cut short in practice — a hotfix lands, a release window closes, CI minutes run out — so the cases that run first should be the ones that detect the most faults first. Researchers call this the rate of fault detection and score it with APFD (average percentage of faults detected), 0–100, for how early a suite reveals the faults it can catch.

Elbaum, Malishevsky and Rothermel (IEEE Transactions on Software Engineering, 2002) compared 16 prioritization techniques across eight C programs and thousands of replications. Two findings matter in practice:

  • Prioritization works, and granularity buys a little. Every technique studied improved fault-detection rate over unordered suites. Statement-level ordering narrowly beat function-level ordering — in the all-programs experiment the best statement-level technique averaged APFD 80.73 against 77.45 for its function-level equivalent — so a cheap, coarse ordering captures most of the gain.
  • No ordering wins everywhere. The relative effectiveness of the techniques varied significantly across target programs, and whether the gains turn into savings depends on each team's cost factors — so re-measure your own suite every few months instead of copying a technique off a blog.

A practical ordering you can apply without instrumenting anything: changed modules first, then cases that failed in the last three releases or live in your most defect-dense areas, then the revenue-adjacent flows, then everything else by last-run date. If the pass is abandoned halfway, the first two buckets have already run. Read the study (IEEE TSE 28(2), 2002).

When a red build isn't a regression: flaky tests

The fastest way to destroy a regression suite's authority is to let it cry wolf. A flaky case — one that passes and fails without any change to the code under test — teaches the team to re-run the pipeline until it goes green, at which point real regressions ride through on the next retry. The evidence on how common this is:

EvidenceWhat the study foundSource
61 Java open-source projects on Travis CIJust under 13% of 935 failing builds that later passed again had failed because of flaky tests, not real defectsParry et al., ACM TOSEM 31(1), 2021
Order-dependent cases94 of 96 sampled order-dependent tests produced a false alarm — the suite failed with no bug present in the codeZhang et al., ISSTA 2014 (surveyed in Parry et al.)
Reproducibility in isolationOf 107 flaky tests found by re-running whole suites, only 50 could be reproduced as flaky when run aloneLam et al. (surveyed in Parry et al.)
Developer survey59% of developers said they deal with flaky tests on a monthly, weekly, or daily basisParry et al., ACM TOSEM 31(1), 2021
Where flakiness hidesOrder-dependent tests account for up to 16% of flaky-test bug reports and 9% of flaky-test repairsParry et al., ACM TOSEM 31(1), 2021

Three rules keep the signal clean without slowing the release. First, quarantine, don't retry silently: tag the case, exclude it from the merge gate, and give it an owner and a deadline — a quarantine list that only grows is the same failure with extra steps. Second, treat order-dependence as a defect in the suite, not the product: isolation-only debugging is unreliable, so the fix is usually restoring the test's own setup and teardown. Third, make sure the case that would have caught a fixed bug exists at all — if a regression walked past your suite once, the missing case is the deliverable. The Test Case Generator builds the boundary and negative cases a hurried fix usually skips, and the RTM Generator maps requirements to cases, so a changed requirement tells you which cases to re-run. Read the flaky-test survey (ACM TOSEM 31(1), 2021).

Smoke vs sanity vs regression: the layer cake

The three terms blur together because they are layers of the same pyramid, distinguished by breadth and timing rather than technique:

LayerQuestion it answersBreadthWhen it runs
SmokeIs the build even testable?A handful of critical-path cases (app loads, login works, core API responds)First minutes after every build/deploy — a broken smoke aborts further testing
SanityDoes the changed area behave?Narrow, unscripted checks around the specific changeAfter a minor fix, instead of a full pass
RegressionDid the change break anything else?The selected (or full) existing suitePer merge to main, nightly, and before every release

A useful rule of thumb: smoke is a gate, sanity is a spot-check, regression is an insurance policy. If your smoke suite fails, skip the regression pass — you will only collect noise.

A regression checklist you can reuse

Copy this into your release process. It is ordered — a failure at the top changes what you do at the bottom:

  1. Smoke first. 5–10 critical-path cases. If these fail, stop and fix the build; a regression pass on a broken build measures nothing.
  2. Changed-module cases. Every case whose steps touch a file, endpoint, or screen in the diff — plus one integration case per boundary.
  3. Fix-verification cases. Every bug fixed in this release has a case that reproduces the original bug, and it passes.
  4. Core-flow cases. Login/checkout/search/upload, regardless of the diff.
  5. Data-boundary cases. Empty inputs, maximum lengths, special characters — the class of bugs that survives "it worked in dev".
  6. Cross-browser/device sample. At minimum one pass on the second most popular browser your analytics show, not just the one developers use.
  7. Visual spot-check. Layout breaks (a missing CSS file, a broken bundle) survive functional green lights; look at one page per template.
  8. Log results per case, including passes. "Regression passed" as a single checkbox hides which cases were skipped.

Working regression checks into the broader process

Regression cases are written once and executed hundreds of times, so their quality bar is higher than one-off test cases: exact test data, no dependence on execution order, and self-contained teardown. The structure follows the same anatomy as any good case — if your team is still standardizing case writing, start with our guide on how to write test cases, then convert the review comments you get into permanent suite rules. When a regression pass finds a defect, the report needs the same reproducibility discipline — our bug report template covers the severity/symptom separation that keeps triage fast.

Metrics close the loop. Track two numbers per pass: the percentage of the planned suite actually executed (execution rate) and the defect removal efficiency of the release. If execution rate keeps dropping below ~80%, the suite is too big for its window — prune or parallelize instead of silently skipping. The QA Metrics Calculator computes both, plus defect density per module, which tells you where the next regression is most likely to come from.

Where standards fit

If your organization needs a formal reference — for audits, certifications, or process documentation — the international standard for software testing is the ISO/IEC/IEEE 29119 series, whose first part defines the core concepts and vocabulary (ISO/IEC/IEEE 29119-1:2022). It formalizes exactly the layered structure above: test levels, test types, and regression as change-related testing. You do not need the standard to test well, but citing it ends "is this the right process?" arguments faster than any blog post — including this one.

Related reading