Let e2e runs share a host, instead of colliding on fixed ports #118

Merged
cruelacid merged 2 commits from worktree-e2e-parallel into main 2026-09-23 16:20:31 +01:00
Owner

Supersedes #117. Part of #77 and #80. Also phase 5 of the agent-queue plan ("parallel-safe e2e").

Why: the first afternoon with the e2e gate on pull requests, overlapping runs failed with EADDRINUSE 127.0.0.1:21236 (508/509, then #115 and #116). Nothing had ever kept two runs apart. The job-level concurrency is ignored by Forgejo, and runner containers share the host network namespace.

What: each run isolates itself instead of queueing:

  • Port block per run (packages/e2e/port-block.ts). The run claims 32 ports in 20000–32767 by holding the block's last port open for the whole run. Suite ports are the base plus a fixed offset (test/ports.ts), so a lone run stays on 21234–21260. A pinned NECTENDA_E2E_PORT_BASE that is taken fails loudly.
  • Per-checkout TMPDIR, /tmp/nectenda-e2e-<hash> (run-scope.ts), short enough for Chromium's socket path limit.
  • clean.sh is scoped to that dir and this checkout's chromedriver. ./clean.sh all is the old machine-wide sweep. run-detached.sh logs in the checkout.
  • The mirror lock is atomic and refuses a second suite in one checkout. Release is by pid.
  • CI: Xvfb runs with -nolisten local, and the run prints its display and the abstract-socket count. The dead job-level group and e2e.yml's group are removed. Listener diagnostics cover the whole range.

Verified locally:

  • Two full suites ran in two worktrees at once, both green 16/16, on blocks 32640 and 32064.
  • While both ran, each clean.sh check counted only its own processes (8 and 12 test processes; 2 and 3 chromedrivers).
  • A pinned held block stops at claim.
  • The mirror lock refuses a second suite and names the first.
  • Mutation-checked through the real global setup: an unheld sentinel fails 2 tests, and an unpublished base fails 1.

Still to verify here: this run overlapping another PR's run on the runner, both green. Each should log a different block, and the abstract X11 count should be 0.

Changelog

NONE

🤖 Generated with Claude Code

https://claude.ai/code/session_01StURdiv33xnMfE2XRyg8Lt

Supersedes #117. Part of #77 and #80. Also phase 5 of the agent-queue plan ("parallel-safe e2e"). **Why:** the first afternoon with the e2e gate on pull requests, overlapping runs failed with `EADDRINUSE 127.0.0.1:21236` (508/509, then #115 and #116). Nothing had ever kept two runs apart. The job-level `concurrency` is ignored by Forgejo, and runner containers share the host network namespace. **What:** each run isolates itself instead of queueing: - **Port block per run** (`packages/e2e/port-block.ts`). The run claims 32 ports in 20000–32767 by holding the block's last port open for the whole run. Suite ports are the base plus a fixed offset (`test/ports.ts`), so a lone run stays on 21234–21260. A pinned `NECTENDA_E2E_PORT_BASE` that is taken fails loudly. - **Per-checkout `TMPDIR`**, `/tmp/nectenda-e2e-<hash>` (`run-scope.ts`), short enough for Chromium's socket path limit. - **`clean.sh` is scoped** to that dir and this checkout's chromedriver. `./clean.sh all` is the old machine-wide sweep. `run-detached.sh` logs in the checkout. - **The mirror lock is atomic** and refuses a second suite in one checkout. Release is by pid. - **CI:** Xvfb runs with `-nolisten local`, and the run prints its display and the abstract-socket count. The dead job-level group and `e2e.yml`'s group are removed. Listener diagnostics cover the whole range. **Verified locally:** - Two full suites ran in two worktrees at once, both green 16/16, on blocks 32640 and 32064. - While both ran, each `clean.sh check` counted only its own processes (8 and 12 test processes; 2 and 3 chromedrivers). - A pinned held block stops at claim. - The mirror lock refuses a second suite and names the first. - Mutation-checked through the real global setup: an unheld sentinel fails 2 tests, and an unpublished base fails 1. **Still to verify here:** this run overlapping another PR's run on the runner, both green. Each should log a different block, and the abstract X11 count should be 0. ## Changelog NONE 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01StURdiv33xnMfE2XRyg8Lt
Let e2e runs share a host, instead of colliding on fixed ports
Some checks failed
e2e / multi (push) Failing after 55s
Release note / release-note (pull_request) Successful in 11s
CI / e2e (pull_request) Has been cancelled
CI / promote (pull_request) Has been cancelled
CI / build (pull_request) Has been cancelled
71ec161597
The e2e gate moved onto pull requests, and on its first afternoon two
overlapping runs failed with EADDRINUSE on 127.0.0.1:21236. Nothing had ever
kept them apart: the e2e job's concurrency block was job-level, which
Forgejo ignores, and the runner's containers share the host's network
namespace. It had not mattered while e2e ran only on main.

Each run now isolates itself rather than queueing behind the others:

- Ports: the global setup claims a 32-port block in 20000-32767 by holding
  its last port open for the whole run (port-block.ts). Another run finds
  that sentinel taken and moves on; a run pinned with
  NECTENDA_E2E_PORT_BASE fails loudly instead. Suite ports are base plus a
  fixed offset (test/ports.ts), so a lone run keeps 21234-21260.
- Temp dir: TMPDIR is /tmp/nectenda-e2e-<hash of the checkout>, short
  enough for Chromium's singleton socket, so every Obsidian instance's
  user-data dir names its checkout.
- clean.sh matches only that dir and this checkout's chromedriver, where it
  used to kill every test instance on the machine; `all` keeps the old
  sweep. run-detached.sh logs in the checkout, not /tmp.
- The mirror lock is taken atomically and refuses a second live suite in
  the same checkout; release by pid cannot drop another run's lock.
- CI: Xvfb runs with -nolisten local, so a job cannot reach another's
  display through the shared abstract socket, and the run prints what it
  sees. The no-op job-level group and e2e.yml's group are gone.

Checked: two full suites in two worktrees at once on the Mac, both 16/16
green on blocks 32640 and 32064, and each clean.sh counted only its own
processes. A pinned held block stops the run at claim. ports.test.ts runs
through the real setup; an unheld sentinel fails 2 tests and an unpublished
base fails 1.

Supersedes #117. Part of #77, #80.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01StURdiv33xnMfE2XRyg8Lt
Let the port-range diagnostic find nothing without failing the job
All checks were successful
Release note / release-note (pull_request) Successful in 11s
CI / build (pull_request) Successful in 4m19s
CI / e2e (pull_request) Successful in 4m42s
CI / promote (pull_request) Has been skipped
e2e / multi (push) Successful in 4m31s
CI / build (push) Successful in 4m57s
CI / e2e (push) Successful in 4m9s
CI / promote (push) Successful in 32s
fb8bee07e3
shell: bash runs with pipefail, and the new listener listing ended in a
grep that exits 1 on no matches — the usual case — so e2e.yml run 517
failed before the suite started. awk does the filtering now, and the
pipeline tolerates an empty result.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01StURdiv33xnMfE2XRyg8Lt
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
Nectenda/nectenda!118
No description provided.