Measure CI capacity and clconsole load, and alert when the host is pressed #126

Merged
cruelacid merged 1 commit from worktree-ci-monitoring into main 2026-09-23 17:30:29 +01:00
Owner

Part of #77. It follows up on raising the runner to capacity 5.

  • scripts/ci-capacity.mjs [--hours N] is a read-only report from action_run_job. Per job it gives queue wait (mean and p90) and run time, and run time against how many jobs were already running when it started. It counts still-running jobs, which a mutation showed the first draft did not.
  • deploy/clconsole/ci-host-monitor.{sh,service,timer} is a user timer, run every minute. It logs load, CPU and memory pressure (PSI), available memory and running CI jobs to ~/ci-host/metrics.tsv (14 days). It pushes up/down to an Uptime Kuma push monitor: down at >=25% CPU pressure over 5 min, >=10% memory pressure over 1 min, or unreadable pressure. It only logs until KUMA_PUSH_URL is set.
  • deploy/clconsole/README.md covers install, and how to read the two together to decide on more or fewer runners.

Already installed on clconsole from this branch. First reading: load 6.22, CPU pressure 6.9%, memory pressure 0, 26.7 GB available, 2 jobs.

Not yet alerting: the Kuma push monitor needs creating, with credentials this session doesn't have.

Tests: 7 for the monitor against a fake `/proc` and the deployer's fake curl, and 5 for the report. Four rules in each were mutation-checked.

Changelog

NONE

🤖 Generated with Claude Code

https://claude.ai/code/session_01StURdiv33xnMfE2XRyg8Lt

Part of #77. It follows up on raising the runner to capacity 5. - **`scripts/ci-capacity.mjs [--hours N]`** is a read-only report from `action_run_job`. Per job it gives queue wait (mean and p90) and run time, and run time against how many jobs were already running when it started. It counts still-running jobs, which a mutation showed the first draft did not. - **`deploy/clconsole/ci-host-monitor.{sh,service,timer}`** is a user timer, run every minute. It logs load, CPU and memory pressure (PSI), available memory and running CI jobs to `~/ci-host/metrics.tsv` (14 days). It pushes **up/down** to an Uptime Kuma push monitor: down at >=25% CPU pressure over 5 min, >=10% memory pressure over 1 min, or unreadable pressure. It only logs until `KUMA_PUSH_URL` is set. - **`deploy/clconsole/README.md`** covers install, and how to read the two together to decide on more or fewer runners. **Already installed on clconsole** from this branch. First reading: load 6.22, CPU pressure 6.9%, memory pressure 0, 26.7 GB available, 2 jobs. **Not yet alerting:** the Kuma push monitor needs creating, with credentials this session doesn't have. Tests: 7 for the monitor against a fake \`/proc\` and the deployer's fake curl, and 5 for the report. Four rules in each were mutation-checked. ## Changelog NONE 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01StURdiv33xnMfE2XRyg8Lt
Measure CI capacity and clconsole load, and alert when the host is pressed
Some checks failed
Release note / release-note (pull_request) Successful in 14s
CI / promote (pull_request) Has been cancelled
CI / e2e (pull_request) Has been cancelled
CI / build (pull_request) Has been cancelled
7fc6fde100
Runner capacity went 2 -> 3 -> 5 on 23 September on numbers read by hand.
Two tools keep reading them:

scripts/ci-capacity.mjs reports, per job, queue wait and run time, and run
time against how many jobs were already running — a run time that climbs
with concurrency means oversubscription whatever the wait says. It counts a
job still running (stopped = 0) as running, which the first draft did not,
and a mutation found.

deploy/clconsole/ci-host-monitor.sh runs every minute from a user timer on
clconsole, logs load, CPU and memory pressure (PSI), available memory and
running CI jobs, and pushes up or down to an Uptime Kuma push monitor —
down at 25% CPU pressure over five minutes or 10% memory over one, and
down when pressure cannot be read at all. clconsole also runs Windows and a
game server, so contention is noticed when it happens rather than as a
flaky test later.

Tested against a fake /proc and the deployer's fake curl; four rules in the
monitor and four in the report were each broken to confirm a test fails.

Part of #77.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01StURdiv33xnMfE2XRyg8Lt
cruelacid force-pushed worktree-ci-monitoring from 7fc6fde100
Some checks failed
Release note / release-note (pull_request) Successful in 14s
CI / promote (pull_request) Has been cancelled
CI / e2e (pull_request) Has been cancelled
CI / build (pull_request) Has been cancelled
to ffb3ea1417
All checks were successful
e2e / multi (push) Successful in 4m25s
Release note / release-note (pull_request) Successful in 13s
CI / build (pull_request) Successful in 5m15s
CI / e2e (pull_request) Successful in 4m6s
CI / promote (pull_request) Has been skipped
CI / e2e (push) Successful in 4m2s
CI / build (push) Successful in 4m37s
CI / promote (push) Successful in 31s
2026-09-23 17:21:16 +01:00
Compare
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
Nectenda/nectenda!126
No description provided.