Smoke Testing vs Sanity Testing: the Difference, Finally Clear

Smoke vs sanity testing: what each one is, when to run which, a side-by-side comparison table, and real examples from web and API projects.

Smoke and sanity testing get conflated because both are short, both are shallow, and both run before anyone does "real" testing. The confusion is mostly harmless — until a build fails and a team argues for two hours about which pass "counts". The distinction is simple once you anchor it to purpose: smoke testing asks "is this build stable enough to test at all?" while sanity testing asks "did this specific change do what it claimed?" Smoke is a gate on the build; sanity is a spot-check on the change.

Smoke testing: the build gate

The name comes from hardware: power on a board and watch for smoke before doing anything elaborate. In software, a smoke suite is a fixed, tiny set of critical-path checks — typically 5 to 15 cases — that must pass before anyone invests time in deeper testing:

  • The app loads and renders its main screen (no blank page, no bundle error).
  • A user can sign in with a known test account.
  • The primary flow of the product completes (add to cart, submit the form, send the message).
  • The main API endpoint responds with a valid payload.

Two properties matter more than the exact list. First, smoke suites are fixed — same cases every build, so "green smoke" is a comparable signal across builds, not a fresh judgment call. Second, smoke is pass/fail for the build itself: a smoke failure means nobody tests this build, full stop. Running a full regression pass on top of a failed smoke just manufactures bug reports about a build everyone already knew was broken.

Sanity testing: the change spot-check

Sanity testing is narrower and ad hoc. When a developer fixes one bug or ships a small change, sanity checks confirm the change works and its immediate neighbors did not break — before scheduling anything more thorough. It is usually unscripted, done by the tester who knows the area, in minutes:

  • The login-timeout bug fix: verify the original repro now passes, plus one check that a normal login still works.
  • A bumped dependency: one pass through the screens that library touches.
  • A config change: the one endpoint it affects returns sane data.

Sanity fails fast on purpose: if the spot-check looks wrong, the change goes straight back to the developer without a formal bug lifecycle. If it looks right, the change becomes eligible for the broader regression pass.

Side by side

DimensionSmoke testingSanity testing
GoalIs the build testable at all?Is this specific change sensible?
ScopeCritical path across the whole appOnly the changed module/feature and its neighbors
Scripted?Fixed, scripted suite, same every buildUsually unscripted, improvised per change
TriggerEvery new build or deployEvery minor fix or small change
Who runs itAutomation or QA, right after buildOften the developer or the area's tester
On failureBuild rejected; deeper testing abortedChange bounced back without a formal bug cycle
Typical duration5–30 minutes, automated where possibleMinutes, by hand
AnalogyGate at the front doorPeek through the keyhole

A note on regional vocabulary: some teams — following older textbooks — use "sanity" to mean the shallow build check and "smoke" for something else entirely. When you join a team that argues about the words, skip the dictionary fight and map the purposes instead: which check gates builds, which spot-checks changes. The ISO/IEC/IEEE 29119 testing vocabulary (ISO/IEC/IEEE 29119-1:2022) sidesteps both labels and classifies retesting by purpose and level — a useful tiebreaker in the naming debate.

Where each fits in the pipeline

The three-tier rhythm looks like this in practice:

  1. Build completes → smoke suite runs. Green? Testing proceeds. Red? Build is rejected and the regression pass never starts.
  2. Small change lands → sanity check. Developer or area tester verifies the fix and its immediate surroundings in minutes.
  3. Merge to main / nightly / pre-release → regression pass. The selected subset (or full suite) confirms nothing else moved. Smoke is the gate, sanity is the filter, regression is the insurance.

Teams that skip the smoke gate pay for it in the most expensive way possible: hours of regression results on builds that never had a working login. Teams that skip sanity funnel every typo fix into the full QA queue and drown it. The two short passes exist precisely to protect the long one.

Smoke suites and regression suites share cases — keep them in sync

Your smoke cases are a subset of your regression suite, extracted by priority. That means they obey the same quality bar: exact test data, observable expected results, stable IDs. If your cases are not yet at that standard, fix the foundation first — our guide on how to write test cases walks the anatomy with examples. And because smoke runs on every build, keep the checks honest with real data: the Fake Data Generator produces consistent test personas and payloads, so a smoke login never breaks because someone's fixture expired.

When a smoke or sanity check does catch a defect, it usually graduates into the suite as a permanent case — and if the finding needs a report, our bug report template keeps it actionable instead of dying in a chat message.

How many cases belong in a smoke suite?

The honest answer is "as few as will still catch a broken build" — but there is a workable sizing rule. Start from the critical path, then add one case per integration boundary the build crosses:

  • One per entry point: the primary screen or route a real user lands on.
  • One per authentication path: sign-in, and a token/session refresh if the app uses one.
  • One per primary write: the app's main "create something" action, verified in the data store — not just in the UI.
  • One per external dependency: payment, email, auth provider, third-party API — the calls that fail for reasons your code cannot fix.
  • One per build artefact: the bundle actually loads, the service worker or background worker boots, the migration applies cleanly.

If that lands above roughly 15 cases, the suite has drifted into regression territory: move the extras to the regression pass and keep the gate fast enough that nobody is tempted to skip it. A 40-minute smoke suite gets bypassed on a Friday release; a 6-minute one does not.

What a broken build actually costs

The case for a cheap gate in front of expensive testing is not just intuition. In The Economic Impacts of Inadequate Infrastructure for Software Testing (NIST Planning Report 02-3, prepared by RTI for NIST, May 2002 — full report), the estimated national annual cost of inadequate software-testing infrastructure was $22.2 billion to $59.5 billion. Four findings in that study map directly onto smoke and sanity practice:

Finding (NIST Planning Report 02-3)Reported figureWhat it means for a smoke gate
Estimated annual U.S. cost of an inadequate testing infrastructure $22.2B–$59.5B (Table ES-4: $59.5B) Testing effort is scarce; spending it on unstable builds is the worst possible allocation.
Share of those costs borne by software users rather than developers Over half (~60% users / ~40% developers) A defect that escapes to a user costs far more than one caught in the build gate — exactly what a smoke failure prevents.
Potential cost reduction from feasible infrastructure improvements $22.2B annually Cheap, repeatable gates (fixed smoke suites, scripted sanity checks) are the improvement most teams can actually implement.
Maintenance reduction surveyed organizations attributed to eliminating errors and bugs ~14.4% of maintenance spend Prevention pays on the maintenance line, years after the release that skipped the gate.

Read those together and the pipeline rhythm from the previous section stops looking like bureaucracy: the smoke gate is the cheapest place in the process to find out that a build cannot be tested, and the sanity spot-check is the cheapest way to keep a one-line fix from consuming a regression slot. Size the effort for the suite you are protecting with the Test Estimation Calculator rather than by gut feel, and log what the gate caught — a smoke suite that has never failed in six months is either excellent or asleep.

Related reading