Research guide

How to Debug an E2E Test That Only Fails in CI

A test that passes locally but fails in CI is usually not a random flake. It exposes an environment difference. This guide covers the root causes and a workflow to identify the difference.

Playwright example

Wait for application state, not elapsed time

Timing-dependent check

await page.getByRole('button', {
  name: 'Place order',
}).click()

await page.waitForTimeout(2_000)
await expect(page.getByText('Order confirmed'))
  .toBeVisible()

State-based check

const order = page.waitForResponse((response) =>
  response.url().endsWith('/api/orders') &&
  response.status() === 201,
)

await page.getByRole('button', {
  name: 'Place order',
}).click()

await order
await expect(page.getByText('Order confirmed'))
  .toBeVisible()

Treat the failure as an environment difference

CI differs from a laptop in concrete ways: headless mode, parallel workers sharing resources, a sandboxed network, a different base URL, stricter time limits, and a fresh database.

A test that fails only in CI shows that one of those differences matters. Once you identify it, correct the specific timing, state, network, or configuration assumption.

Check the common CI differences

Headless-only behavior, such as animations, WebGL, and fonts that render differently; timing under load when many workers start at once; shared test data colliding across parallel workers; missing environment variables; and network calls blocked by the CI sandbox.

The trace and console logs in CI usually name the culprit: a timeout, a missing element, or a request that never returned.

Reproduce the CI environment locally

Run the test headed and headless, restrict the worker count to match CI, lower or raise concurrency, set the same environment variables, and use the same browser launch flags CI uses.

Reproducing the difference locally converts a mystery into a variable you can change and observe.

Inspect the artifacts from the failed job

Read the trace, screenshot, video, and console output from the CI run and compare them with local. A timeout here, a missing element there, or a console error that never appears locally points straight at the difference.

Use the artifacts to locate the first point where local and CI behavior diverge.

Fix the difference instead of hiding it

Wait on state with web-first assertions instead of fixed sleeps, scope test data per worker, set a deterministic base URL, and add health checks for third-party services CI depends on.

Resist blanket retries as the default. A retry can hide the environmental bug the test is meant to catch.

Keep CI configuration visible

Document the environment variables and browser flags CI relies on, and make config a single source of truth shared by local and CI.

For the smoke set that should survive the CI environment, see the deploy-time smoke guide.

Use a fixed diagnostic order

First rerun the failed test alone with the same browser, worker count, environment variables, locale, timezone, and container image. Then run it with the full shard to detect shared-state or resource contention.

Compare the first failed action across the trace, browser console, and network log. Change one variable at a time and keep the failure packet from the original job so each rerun has a stable reference.

Capture enough evidence before the next failure

Retain the Playwright trace on the first retry, plus the screenshot, browser console, failed requests, commit SHA, browser version, and worker index. Record the test-data identifier when parallel workers create server-side state.

A screenshot shows the final frame. The trace shows which locator resolved, which request stalled, and whether navigation replaced the page between the action and assertion.

When the “flake” is really a changed journey

Some fails-only-in-CI cases are not environment differences at all. The scripted path drifted from the real journey, and CI is simply the first place the drift is exercised at full speed with real timing. If a journey keeps failing across retries and workers, treat it as a possible product change rather than a flake, and verify the expected result against the live interface.

An agentic journey test can settle the argument. Instead of a fixed locator script, you describe the user outcome and let a live browser run adapt and re-verify it. When the coded path and the agentic path disagree, the agentic run shows whether a customer can still complete the flow in the current UI, turning a CI dispute into release evidence.

Key takeaways

  • Fails-only-in-CI is an environment difference, so reproduce the CI environment locally.
  • Check headless mode, worker count, shared data, and env vars before touching the test.
  • Fix the difference with explicit waits and isolated data, not blanket retries.

Sources

Related CueTest resources