Serialise CI at the workflow level, where Forgejo actually applies it #117

Closed
cruelacid wants to merge 1 commit from worktree-e2e-serial into main
Owner

Part of #77, and fixes a consequence of #80. Merge this first: every other PR's e2e can collide until it lands.

What broke: main's e2e for the #113 merge (run 508) and #114's (run 509) both failed with EADDRINUSE 127.0.0.1:21236. They ran at the same time: 508 from 14:33:36 to 14:37:05, 509 from 14:34:02 to 14:37:29. The suites use fixed ports on a network=host runner.

Why: the e2e job's concurrency: group: e2e was job-level, and Forgejo's Actions reference says "Multiple jobs within a workflow are not affected by the concurrency setting". It had never serialised anything. That didn't matter while e2e ran only on pushes to main, and #80 put it on every PR.

Fix: a workflow-level concurrency: { group: e2e, cancel-in-progress: false } at the top of ci.yml, the group e2e.yml already uses. Runs queue rather than cancel, so no PR's run cancels another's and a later push can't cancel main mid-promote. The cost is that whole CI runs serialise, at about 10 minutes each.

It also corrects the README entry I wrote for #113's failed run 504. That run overlapped 489 by 48 s, so a collision is a candidate cause, not the known hang.

:stable did not move: 508's promote never ran, so production is still on 796d617.

Changelog

NONE

🤖 Generated with Claude Code

https://claude.ai/code/session_01StURdiv33xnMfE2XRyg8Lt

Part of #77, and fixes a consequence of #80. **Merge this first**: every other PR's e2e can collide until it lands. **What broke:** `main`'s e2e for the #113 merge (run 508) and #114's (run 509) both failed with `EADDRINUSE 127.0.0.1:21236`. They ran at the same time: 508 from 14:33:36 to 14:37:05, 509 from 14:34:02 to 14:37:29. The suites use fixed ports on a `network=host` runner. **Why:** the e2e job's `concurrency: group: e2e` was job-level, and Forgejo's Actions reference says "Multiple jobs within a workflow are not affected by the concurrency setting". It had never serialised anything. That didn't matter while e2e ran only on pushes to `main`, and #80 put it on every PR. **Fix:** a workflow-level `concurrency: { group: e2e, cancel-in-progress: false }` at the top of `ci.yml`, the group `e2e.yml` already uses. Runs queue rather than cancel, so no PR's run cancels another's and a later push can't cancel `main` mid-promote. The cost is that whole CI runs serialise, at about 10 minutes each. It also corrects the README entry I wrote for #113's failed run 504. That run overlapped 489 by 48 s, so a collision is a candidate cause, not the known hang. `:stable` did not move: 508's promote never ran, so production is still on `796d617`. ## Changelog NONE 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01StURdiv33xnMfE2XRyg8Lt
Serialise CI at the workflow level, where Forgejo actually applies it
All checks were successful
Release note / release-note (pull_request) Successful in 11s
CI / build (pull_request) Successful in 4m10s
CI / e2e (pull_request) Successful in 4m36s
CI / promote (pull_request) Has been skipped
17cde8ee8e
The e2e job's concurrency block never did anything: Forgejo applies
concurrency to whole workflows only ("Multiple jobs within a workflow are
not affected by the concurrency setting"). It went unnoticed while e2e ran
only on pushes to main. Once it ran on pull requests, three of them and a
merge ran it at once, and runs 508 and 509 failed with EADDRINUSE on
127.0.0.1:21236 — the fixed ports on a host-networked runner.

The group moves to the top of ci.yml, shared with e2e.yml's nightly and
queued rather than cancelled. Whole runs now wait for each other, build
included; promote needs e2e green in its own run, which rules out splitting
e2e into a separate workflow to save those minutes.

The README's note on run 504 is corrected: that run overlapped 489 by 48 s,
so a collision is a candidate cause and it should not be counted as the
unexplained hang.

Part of #77, #80.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01StURdiv33xnMfE2XRyg8Lt
Author
Owner

Closing unmerged, by decision. Its run passed (515), but serialising every CI run is the wrong fix. It would make ten-minute runs queue behind each other precisely when the agent work queue wants several pull requests in flight.

Instead, e2e becomes safe to run concurrently: a port block per run held by a sentinel listener, a checkout-scoped `clean.sh` and TMPDIR, and Xvfb without the abstract socket. That is phase 5 of the agent-queue plan, done now. This PR's diagnosis stands: job-level `concurrency` is ignored by Forgejo. Its README correction to the run-504 entry moves into the replacement PR.

Closing unmerged, by decision. Its run passed (515), but serialising every CI run is the wrong fix. It would make ten-minute runs queue behind each other precisely when the agent work queue wants several pull requests in flight. Instead, e2e becomes safe to run concurrently: a port block per run held by a sentinel listener, a checkout-scoped \`clean.sh\` and TMPDIR, and Xvfb without the abstract socket. That is phase 5 of the agent-queue plan, done now. This PR's diagnosis stands: job-level \`concurrency\` is ignored by Forgejo. Its README correction to the run-504 entry moves into the replacement PR.
cruelacid closed this pull request 2026-09-23 15:53:56 +01:00
All checks were successful
Release note / release-note (pull_request) Successful in 11s
Required
Details
CI / build (pull_request) Successful in 4m10s
Required
Details
CI / e2e (pull_request) Successful in 4m36s
Required
Details
CI / promote (pull_request) Has been skipped

Pull request closed

Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
Nectenda/nectenda!117
No description provided.