What Is Regression Testing? Techniques + a Reusable Checklist
Regression testing explained: risk-based selection, prioritization that works, why flaky tests lie, a reusable checklist, and how to keep suites fast.
Regression testing answers one question after every change: did anything that used to work still work? New features get the attention, demos, and test plans — but the bugs that burn teams are almost always in code nobody touched. The ISTQB defines regression testing as "a type of change-related testing to detect whether defects have been introduced or uncovered in unchanged areas of the software" (see the ISTQB Glossary entry). The phrase "unchanged areas" is the part teams get wrong: you are not retesting the change, you are testing everything the change could have broken around it.
Why regressions happen (and why re-running everything is not the answer)
A regression sneaks in through the seams between modules. Typical mechanisms:
- Shared code changes. A utility function gets a "safe" tweak; three callers downstream depended on the old edge-case behavior.
- Refactors with green unit tests. Unit mocks encode the old assumptions, so a broken integration still passes every unit suite.
- Config, dependency, and data drift. The code is identical but the library version, migration, or feature flag underneath it moved.
- Fixed bugs that were never folded back into the suite. The bug returns because no test ever existed that would have caught it.
The instinctive answer — re-run the entire suite on every commit — fails economically long before it fails technically. A 4,000-case suite at 30 seconds per case is 33 hours of wall-clock time per pass. By the time the pass finishes, the code has moved again. Practical regression testing is therefore a selection problem: run the right subset, fast, on every change.
Choosing what to run: risk-based selection
Rank every candidate case on two axes — the likelihood the change broke it, and the impact if it silently broke. Then apply a simple selection order:
- All cases touching changed modules — including their neighbors at integration boundaries, not just the changed functions.
- All cases for bug fixes shipped in this release. A fixed bug without a regression case is a timed reinstall: fold every fix into the suite permanently.
- High-traffic, revenue-adjacent flows (login, checkout, signup, search) even when nothing near them changed — these are the failures users find first.
- Historically flaky or defect-dense areas. Modules that produced bugs last quarter statistically produce them again.
Everything else runs on a rotation — nightly, weekly, or before release only. That rotation is what keeps daily passes inside a coffee break instead of a day.
Ordering the subset: prioritization, not shuffling
Once the subset is chosen, order decides what you learn. Regression passes get cut short in practice — a hotfix lands, a release window closes, CI minutes run out — so the cases that run first should be the ones that detect the most faults first. Researchers call this the rate of fault detection and score it with APFD (average percentage of faults detected), 0–100, for how early a suite reveals the faults it can catch.
Elbaum, Malishevsky and Rothermel (IEEE Transactions on Software Engineering, 2002) compared 16 prioritization techniques across eight C programs and thousands of replications. Two findings matter in practice:
- Prioritization works, and granularity buys a little. Every technique studied improved fault-detection rate over unordered suites. Statement-level ordering narrowly beat function-level ordering — in the all-programs experiment the best statement-level technique averaged APFD 80.73 against 77.45 for its function-level equivalent — so a cheap, coarse ordering captures most of the gain.
- No ordering wins everywhere. The relative effectiveness of the techniques varied significantly across target programs, and whether the gains turn into savings depends on each team's cost factors — so re-measure your own suite every few months instead of copying a technique off a blog.
A practical ordering you can apply without instrumenting anything: changed modules first, then cases that failed in the last three releases or live in your most defect-dense areas, then the revenue-adjacent flows, then everything else by last-run date. If the pass is abandoned halfway, the first two buckets have already run. Read the study (IEEE TSE 28(2), 2002).
When a red build isn't a regression: flaky tests
The fastest way to destroy a regression suite's authority is to let it cry wolf. A flaky case — one that passes and fails without any change to the code under test — teaches the team to re-run the pipeline until it goes green, at which point real regressions ride through on the next retry. The evidence on how common this is:
| Evidence | What the study found | Source |
|---|---|---|
| 61 Java open-source projects on Travis CI | Just under 13% of 935 failing builds that later passed again had failed because of flaky tests, not real defects | Parry et al., ACM TOSEM 31(1), 2021 |
| Order-dependent cases | 94 of 96 sampled order-dependent tests produced a false alarm — the suite failed with no bug present in the code | Zhang et al., ISSTA 2014 (surveyed in Parry et al.) |
| Reproducibility in isolation | Of 107 flaky tests found by re-running whole suites, only 50 could be reproduced as flaky when run alone | Lam et al. (surveyed in Parry et al.) |
| Developer survey | 59% of developers said they deal with flaky tests on a monthly, weekly, or daily basis | Parry et al., ACM TOSEM 31(1), 2021 |
| Where flakiness hides | Order-dependent tests account for up to 16% of flaky-test bug reports and 9% of flaky-test repairs | Parry et al., ACM TOSEM 31(1), 2021 |
Three rules keep the signal clean without slowing the release. First, quarantine, don't retry silently: tag the case, exclude it from the merge gate, and give it an owner and a deadline — a quarantine list that only grows is the same failure with extra steps. Second, treat order-dependence as a defect in the suite, not the product: isolation-only debugging is unreliable, so the fix is usually restoring the test's own setup and teardown. Third, make sure the case that would have caught a fixed bug exists at all — if a regression walked past your suite once, the missing case is the deliverable. The Test Case Generator builds the boundary and negative cases a hurried fix usually skips, and the RTM Generator maps requirements to cases, so a changed requirement tells you which cases to re-run. Read the flaky-test survey (ACM TOSEM 31(1), 2021).
Smoke vs sanity vs regression: the layer cake
The three terms blur together because they are layers of the same pyramid, distinguished by breadth and timing rather than technique:
| Layer | Question it answers | Breadth | When it runs |
|---|---|---|---|
| Smoke | Is the build even testable? | A handful of critical-path cases (app loads, login works, core API responds) | First minutes after every build/deploy — a broken smoke aborts further testing |
| Sanity | Does the changed area behave? | Narrow, unscripted checks around the specific change | After a minor fix, instead of a full pass |
| Regression | Did the change break anything else? | The selected (or full) existing suite | Per merge to main, nightly, and before every release |
A useful rule of thumb: smoke is a gate, sanity is a spot-check, regression is an insurance policy. If your smoke suite fails, skip the regression pass — you will only collect noise.
A regression checklist you can reuse
Copy this into your release process. It is ordered — a failure at the top changes what you do at the bottom:
- Smoke first. 5–10 critical-path cases. If these fail, stop and fix the build; a regression pass on a broken build measures nothing.
- Changed-module cases. Every case whose steps touch a file, endpoint, or screen in the diff — plus one integration case per boundary.
- Fix-verification cases. Every bug fixed in this release has a case that reproduces the original bug, and it passes.
- Core-flow cases. Login/checkout/search/upload, regardless of the diff.
- Data-boundary cases. Empty inputs, maximum lengths, special characters — the class of bugs that survives "it worked in dev".
- Cross-browser/device sample. At minimum one pass on the second most popular browser your analytics show, not just the one developers use.
- Visual spot-check. Layout breaks (a missing CSS file, a broken bundle) survive functional green lights; look at one page per template.
- Log results per case, including passes. "Regression passed" as a single checkbox hides which cases were skipped.
Working regression checks into the broader process
Regression cases are written once and executed hundreds of times, so their quality bar is higher than one-off test cases: exact test data, no dependence on execution order, and self-contained teardown. The structure follows the same anatomy as any good case — if your team is still standardizing case writing, start with our guide on how to write test cases, then convert the review comments you get into permanent suite rules. When a regression pass finds a defect, the report needs the same reproducibility discipline — our bug report template covers the severity/symptom separation that keeps triage fast.
Metrics close the loop. Track two numbers per pass: the percentage of the planned suite actually executed (execution rate) and the defect removal efficiency of the release. If execution rate keeps dropping below ~80%, the suite is too big for its window — prune or parallelize instead of silently skipping. The QA Metrics Calculator computes both, plus defect density per module, which tells you where the next regression is most likely to come from.
Where standards fit
If your organization needs a formal reference — for audits, certifications, or process documentation — the international standard for software testing is the ISO/IEC/IEEE 29119 series, whose first part defines the core concepts and vocabulary (ISO/IEC/IEEE 29119-1:2022). It formalizes exactly the layered structure above: test levels, test types, and regression as change-related testing. You do not need the standard to test well, but citing it ends "is this the right process?" arguments faster than any blog post — including this one.
Related reading
- How to Write Test Cases (Examples + Template) — the suite you re-run is only as good as the individual cases in it.
- Bug Report Template (+ Examples That Get Fixed Faster) — what to file when the regression pass catches a real break.
- How to Test JSON APIs Without Writing Code — smoke-check API endpoints interactively before wiring them into the suite.
- Test Estimation Calculator — price the regression passes before someone asks when the build is safe; three-point estimates give you a range instead of a guess.