Research guide

E2E Test Failed: Bug, Flake, Environment or Test Problem?

Every E2E failure belongs to one of four buckets: bug, flake, environment, or test problem. This guide gives a triage framework that tells you which, and what evidence to gather before escalating.

Four buckets, one question

Every failure is one of four things: a bug, where the app does not match spec; a flake, where the test is unstable; an environment problem, where CI or test infrastructure failed; or a test problem, where the assertion or setup is wrong.

Each bucket has a different owner and a different fix, so classifying first prevents the team from fixing the wrong thing.

Gather evidence before you classify

Collect what changed, the failing step, the screenshot and trace, whether the test passed on a rerun, and whether it passes locally.

The deploy log and the trace answer most classification questions before a human has to read the test.

Signals that say flake

A test that passes on rerun, fails on timing-sensitive assertions, breaks only in CI, or passes when the worker count is lowered is almost certainly a flake.

Flakes need ownership, not just a rerun. An unowned flake repeats and accumulates until the suite is ignored.

Signals that say environment

When the whole suite or unrelated tests time out at once, suspect the environment: exhausted resources, shared test-data collisions, a third-party outage, or a misconfigured CI container.

Environment failures look nothing like the failing test, so look at the pattern across the run, not the single failure.

Signals that say test problem

A brittle assertion, a wrong selector, missing setup, or a test that asserts behavior the product intentionally changed is a test problem.

These are the easiest to fix and the easiest to rerun past. If the test contradicts current product behavior, update the test and note it for the next review.

Signals that say bug

A failure that reproduces deterministically, matches a behavior change that just shipped, and fails with the same evidence every time is a bug.

This is the case to escalate with evidence rather than rerun. Rerunning a real bug hides a release risk behind a green button.

Turn classification into policy

Write down when a rerun is safe, flake and environment, and when it is not, bug and test problem, plus who owns each bucket. The policy turns triage from instinct into process.

Pair this with the flaky-debugging workflow and the fails-only-in-CI guide for the two hardest buckets.

Key takeaways

  • Classify before you rerun: bug, flake, environment, or test problem.
  • Gather evidence (what changed, trace, retry history) before deciding.
  • Rerunning a bug hides a release risk; rerunning a flake is fine if it’s owned.

Sources

Related CueTest resources