Move the ops box's Uptime Kuma to 2.5.5, rehearsed on a copy of its data #135

Merged
cruelacid merged 3 commits from worktree-kuma-2 into main 2026-09-23 21:31:19 +01:00
Owner

Task: NEC-88

Pins the ops box's Uptime Kuma to 2.5.5 by digest, replacing :1. This PR only changes the pin and the runbook. The box itself is upgraded afterwards, by the procedure in deploy/UPGRADES.md (snapshot, banner, named-service pull).

Rehearsed off the box on an online .backup of the live kuma.db:

  • Migration took 47 s locally, 45 s of it the aggregate-table step.
  • All 17 monitors, both push tokens, both maintenances and their links, 34 notification links, and the status page with status.nectenda.com came across unchanged. The only difference is a new, empty column on monitor_group.
  • 2.x keeps only "important" raw heartbeats after aggregating: 161,625 rows became 15,307. The uptime history is kept in the aggregates.
  • kuma-compat.mjs passed 12/12 on the migrated copy and 12/12 on 1.23.17. The version branch differs visibly between them: conditions and incidents[] on 2.x, versus no conditions and a single incident on 1.x.
  • One unexplained miss: the first login against the fresh 2.5.5 container got no acknowledgement within 8 s. A retry worked and it didn't recur.

Changelog

NONE

🤖 Generated with Claude Code

https://claude.ai/code/session_014jXw5B2oryJCk6KxdkjZJR

Task: NEC-88 Pins the ops box's Uptime Kuma to `2.5.5` by digest, replacing `:1`. This PR only changes the pin and the runbook. The box itself is upgraded afterwards, by the procedure in `deploy/UPGRADES.md` (snapshot, banner, named-service pull). **Rehearsed off the box** on an online `.backup` of the live `kuma.db`: - Migration took 47 s locally, 45 s of it the aggregate-table step. - All 17 monitors, both push tokens, both maintenances and their links, 34 notification links, and the status page with `status.nectenda.com` came across unchanged. The only difference is a new, empty column on `monitor_group`. - 2.x keeps only "important" raw heartbeats after aggregating: 161,625 rows became 15,307. The uptime history is kept in the aggregates. - `kuma-compat.mjs` passed 12/12 on the migrated copy and 12/12 on 1.23.17. The version branch differs visibly between them: `conditions` and `incidents[]` on 2.x, versus no `conditions` and a single `incident` on 1.x. - One unexplained miss: the first `login` against the fresh 2.5.5 container got no acknowledgement within 8 s. A retry worked and it didn't recur. ## Changelog NONE 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_014jXw5B2oryJCk6KxdkjZJR
Move the ops box's Uptime Kuma to 2.5.5, rehearsed on a copy of its data
Some checks failed
Release note / release-note (pull_request) Successful in 13s
CI / build (pull_request) Successful in 4m20s
CI / e2e (pull_request) Failing after 8m41s
CI / promote (pull_request) Has been skipped
79aba91a35
The pin was `:1`, which can never express the move to 2.x. It is now
2.5.5 by digest. Before changing it, an online `.backup` of the live
kuma.db was opened by 2.5.5 off the box: the migration took 47 s, every
monitor, push token, maintenance link and the status page's custom
domain came across unchanged, and kuma-compat.mjs passed 12/12 on the
migrated copy and on 1.23.17, with the version branch visibly different.

UPGRADES.md records the result, including that 2.x deletes most raw
heartbeats after aggregating them, and fixes the rehearsal recipe: take
the copy with .backup rather than through snapshot.sh (which retires the
oldest snapshot), and switch notifications off in the copy so it cannot
send real DOWN alerts.

Task: NEC-88

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014jXw5B2oryJCk6KxdkjZJR
Record the repro-runaway hang recurring on a run that overlapped nothing
All checks were successful
Release note / release-note (pull_request) Successful in 13s
CI / build (pull_request) Successful in 4m21s
CI / e2e (pull_request) Successful in 4m50s
CI / promote (pull_request) Has been skipped
27ff09c658
Run 611 on this branch hung the same test the same way, with no other e2e
job running and nothing under packages/ changed. The server log is silent
for the five minutes, which says the vault stopped talking rather than
looped. Tracked as NEC-89.

Task: NEC-88

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014jXw5B2oryJCk6KxdkjZJR
cruelacid force-pushed worktree-kuma-2 from 27ff09c658
All checks were successful
Release note / release-note (pull_request) Successful in 13s
CI / build (pull_request) Successful in 4m21s
CI / e2e (pull_request) Successful in 4m50s
CI / promote (pull_request) Has been skipped
to 5dbbef9c8f
Some checks failed
Release note / release-note (pull_request) Successful in 13s
CI / build (pull_request) Successful in 4m19s
CI / e2e (pull_request) Failing after 8m40s
CI / promote (pull_request) Has been skipped
2026-09-23 20:36:49 +01:00
Compare
Record the repro-runaway hang's third occurrence, on run 615
All checks were successful
Release note / release-note (pull_request) Successful in 13s
CI / build (pull_request) Successful in 4m29s
CI / e2e (pull_request) Successful in 4m48s
CI / promote (pull_request) Has been skipped
CI / build (push) Successful in 4m38s
CI / e2e (push) Successful in 4m49s
CI / promote (push) Successful in 30s
6bb4465311
Two hangs in three runs on this branch, none overlapping another suite.

Task: NEC-88

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014jXw5B2oryJCk6KxdkjZJR
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
Nectenda/nectenda!135
No description provided.