Skip to content

e2e: all runs cancel at the 20m timeout — 'playwright install-deps' hangs on apt #246

Description

@aspiers

Symptom

Every e2e run is currently cancelling at the 20-minute job timeout without executing a single test. This blocks the run check on all PRs.

Three runs, three identical hangs:

Run Ref SHA Outcome
31580298348 #234 branch 9d0d0fb cancelled @ 20m
31580298348 (rerun) #234 branch 9d0d0fb cancelled @ 20m
31583701010 main, workflow_dispatch vs pr-base bf775ee cancelled @ 20m

The third is the important one: unmodified main, no PR involved. This is infrastructure, not any branch's code.

Where it hangs

Install Playwright system dependencies (cache hit) (.github/workflows/e2e-tests.yml:151-153). Log tail is byte-identical across all three:

09:37:50  Get:3 http://archive.ubuntu.com/ubuntu noble-updates InRelease [126 kB]
09:37:50  Get:4 http://archive.ubuntu.com/ubuntu noble-backports InRelease [126 kB]
09:57:11  ##[error]The operation was canceled.

~19.5 minutes of no output, then the job timeout. Teardown confirms what was stuck:

Terminate orphan process: pid (4108) (npm exec playwright install-deps chromium)

The job never reaches the Railway deployment lookup, so no test runs.

Contributing factor: the runner cannot reach mirror.arizona.edu at all — every request to it logs Ign:. Both mirror.arizona.edu and archive.ubuntu.com respond normally from outside CI (HTTP 200, ~1.5s and ~0.5s), so this is specific to the Blacksmith runner's network.

This started recently: the preceding 28 e2e runs all succeeded, including 86068b5 on the same #234 branch ~10h before the first failure.

Why the workflow makes this maximally painful

- name: Install Playwright system dependencies (cache hit)
  if: steps.playwright-cache.outputs.cache-hit == 'true'
  run: npx playwright install-deps chromium
  1. No step-level timeout. A stalled apt consumes the job's entire 20-minute budget, so a network problem presents as an opaque cancelled rather than a legible failure.
  2. No retry.
  3. The name is misleading. "(cache hit)" refers to the browser cache; the step still does a full apt-get update + install over the network every run. The cache restores fine (Cache restored from key: playwright-Linux-6410a6aa...) and buys no protection here.

Suggested fix

Roughly, in increasing order of ambition:

  • Add timeout-minutes: 5 to the step so a mirror stall fails fast and legibly instead of eating the job.
  • Wrap in a retry (e.g. nick-fields/retry) so a single bad mirror does not red the run.
  • Consider whether install-deps is needed at all on a runner image that may already carry the chromium system libs — if so, the step could be dropped or made non-fatal (|| true) with a smoke check that chromium actually launches.
  • Optionally pin apt to a known-good mirror, or strip the unreachable mirror.arizona.edu entry, for the duration of the step.

I have not implemented any of these — filing for a decision first, since it touches shared CI rather than any one PR.

Impact

run is a merge-blocking check, so until this is fixed no PR can go green. #234 is otherwise clean (16 checks passing, 0 unresolved threads, Coveralls +0.04%) and blocked only by this and the required approval.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions