Research guide
Visual Regression Testing: Complete Practical Guide
Visual regression testing catches the layout, spacing, and styling regressions that functional E2E assertions miss. This guide covers how it works, baselines and thresholds, masks, CI integration, and when to use it.
What visual regression catches
Functional assertions prove elements exist; visual testing proves the product still looks right. It catches layout shifts, spacing, color, typography, overlap, and responsive-breakpoint regressions.
Those are the failures that reach users as “something just looks off” bug reports that functional tests never flagged.
How it works
A run captures a screenshot, compares it to a stored baseline, produces a diff, and passes or fails against a threshold. Baselines are versioned and approved, so the comparison is always against a known-good state.
The pipeline is the same shape as functional testing, with the assertion replaced by image comparison.
Baselines and thresholds
A baseline is the approved reference image, and a threshold decides how much difference is a failure. A low threshold flags every font-rendering quirk; a high one lets real regressions through.
Thresholds should be set per view and tuned by the noise you actually see, not by a single global default.
Masks and dynamic regions
Dynamic content, such as dates, ads, auth widgets, and animations, will diff on every run and drown out real changes. Mask those regions so the comparison focuses on what should be stable.
A mask list is part of the test definition, kept current as the page changes.
Selecting what to capture
Capture viewport, element, or full-page screenshots for the critical pages and components, across the breakpoints you support. Element captures are more stable than full pages.
Start with the pages users see most and the components that change most often.
Running it in CI and reviewing diffs
Run visual tests in CI, send failures to an approval flow, and use AI classification to separate real changes from noise before a human reviews. The review is the point: an unapproved baseline change is exactly what you are trying to catch.
CueTest includes a visual regression runner that captures capabilities and keeps versioned baselines for exactly this workflow.
Key takeaways
- Visual tests prove the product still looks right, not just that elements exist.
- Versioned baselines and per-view thresholds keep the suite useful instead of noisy.
- Mask dynamic regions or every run will diff like a failure.