Symptom
Every e2e run is currently cancelling at the 20-minute job timeout without executing a single test. This blocks the run check on all PRs.
Three runs, three identical hangs:
The third is the important one: unmodified main, no PR involved. This is infrastructure, not any branch's code.
Where it hangs
Install Playwright system dependencies (cache hit) (.github/workflows/e2e-tests.yml:151-153). Log tail is byte-identical across all three:
09:37:50 Get:3 http://archive.ubuntu.com/ubuntu noble-updates InRelease [126 kB]
09:37:50 Get:4 http://archive.ubuntu.com/ubuntu noble-backports InRelease [126 kB]
09:57:11 ##[error]The operation was canceled.
~19.5 minutes of no output, then the job timeout. Teardown confirms what was stuck:
Terminate orphan process: pid (4108) (npm exec playwright install-deps chromium)
The job never reaches the Railway deployment lookup, so no test runs.
Contributing factor: the runner cannot reach mirror.arizona.edu at all — every request to it logs Ign:. Both mirror.arizona.edu and archive.ubuntu.com respond normally from outside CI (HTTP 200, ~1.5s and ~0.5s), so this is specific to the Blacksmith runner's network.
This started recently: the preceding 28 e2e runs all succeeded, including 86068b5 on the same #234 branch ~10h before the first failure.
Why the workflow makes this maximally painful
- name: Install Playwright system dependencies (cache hit)
if: steps.playwright-cache.outputs.cache-hit == 'true'
run: npx playwright install-deps chromium
- No step-level timeout. A stalled apt consumes the job's entire 20-minute budget, so a network problem presents as an opaque
cancelled rather than a legible failure.
- No retry.
- The name is misleading. "(cache hit)" refers to the browser cache; the step still does a full
apt-get update + install over the network every run. The cache restores fine (Cache restored from key: playwright-Linux-6410a6aa...) and buys no protection here.
Suggested fix
Roughly, in increasing order of ambition:
- Add
timeout-minutes: 5 to the step so a mirror stall fails fast and legibly instead of eating the job.
- Wrap in a retry (e.g.
nick-fields/retry) so a single bad mirror does not red the run.
- Consider whether
install-deps is needed at all on a runner image that may already carry the chromium system libs — if so, the step could be dropped or made non-fatal (|| true) with a smoke check that chromium actually launches.
- Optionally pin apt to a known-good mirror, or strip the unreachable
mirror.arizona.edu entry, for the duration of the step.
I have not implemented any of these — filing for a decision first, since it touches shared CI rather than any one PR.
Impact
run is a merge-blocking check, so until this is fixed no PR can go green. #234 is otherwise clean (16 checks passing, 0 unresolved threads, Coveralls +0.04%) and blocked only by this and the required approval.
Symptom
Every e2e run is currently cancelling at the 20-minute job timeout without executing a single test. This blocks the
runcheck on all PRs.Three runs, three identical hangs:
workflow_dispatchvspr-baseThe third is the important one: unmodified
main, no PR involved. This is infrastructure, not any branch's code.Where it hangs
Install Playwright system dependencies (cache hit)(.github/workflows/e2e-tests.yml:151-153). Log tail is byte-identical across all three:~19.5 minutes of no output, then the job timeout. Teardown confirms what was stuck:
The job never reaches the Railway deployment lookup, so no test runs.
Contributing factor: the runner cannot reach
mirror.arizona.eduat all — every request to it logsIgn:. Bothmirror.arizona.eduandarchive.ubuntu.comrespond normally from outside CI (HTTP 200, ~1.5s and ~0.5s), so this is specific to the Blacksmith runner's network.This started recently: the preceding 28 e2e runs all succeeded, including 86068b5 on the same #234 branch ~10h before the first failure.
Why the workflow makes this maximally painful
cancelledrather than a legible failure.apt-get update+ install over the network every run. The cache restores fine (Cache restored from key: playwright-Linux-6410a6aa...) and buys no protection here.Suggested fix
Roughly, in increasing order of ambition:
timeout-minutes: 5to the step so a mirror stall fails fast and legibly instead of eating the job.nick-fields/retry) so a single bad mirror does not red the run.install-depsis needed at all on a runner image that may already carry the chromium system libs — if so, the step could be dropped or made non-fatal (|| true) with a smoke check that chromium actually launches.mirror.arizona.eduentry, for the duration of the step.I have not implemented any of these — filing for a decision first, since it touches shared CI rather than any one PR.
Impact
runis a merge-blocking check, so until this is fixed no PR can go green. #234 is otherwise clean (16 checks passing, 0 unresolved threads, Coveralls +0.04%) and blocked only by this and the required approval.