Delete old images from the deployer, and stop reading a full disk as a bad build #158

Merged
nectenda-agent merged 1 commit from worktree-image-housekeeping into main 2026-09-24 21:38:50 +01:00
Collaborator

Task: NEC-101 — https://projectron.nerchure.com/tasks/101

The identity host went down on 24 September 2026. Its disk filled with about a hundred old builds, because the deployer never deletes an image. /api/ready failed on the disk floor. The deployer read that as a bad build and rolled back, and the rollback failed the same way. It then set AUTO_DEPLOY=off and posted a status-page incident. Recovered by hand; this stops it recurring.

  • Prune. Keeps the running build, its rollback target (from history.log) and :stable. Builds are matched by revision label, because :stable is pulled by digest and has no tag. Runs before each pull, after a successful deploy and on every tick that does not deploy, so an upgraded host clears its backlog on the first tick. Only $IMAGE, never rmi -f, never fatal. Also deployer.sh prune [--dry-run].
  • Disk-bound deploys. A build that reports its own sha but fails readiness with "reason":"disk" prunes and retries. If still full it records disk-bound and pushes down, and does not roll back or turn itself off.
  • Early warning. The heartbeat adds disk low: N MB free while free space is under twice the floor. status gains disk_free, disk_floor and local_builds.

Tests: 15 new ones in deploy/test/deployer.test.ts, with fake docker images/rmi/df and readiness failing on disk. Each rule was mutation-checked; a substring-keep mutation survived at first and the test was tightened. The real docker images/image inspect output was checked read-only against eu1: duplicated digest refs, 15 pre-label images, df -Pk parsing.

Not yet verified live. After merge, install-deployer.sh carries this to the hosts (it is not automatic). Run deployer.sh prune --dry-run on eu1 first: it still holds ~100 old builds, left on purpose as test data.

Follow-up, as a separate PR: a CI registry build cache, plus moving GIT_SHA to the end of the runtime stages so builds share layers.

Changelog

NONE

🤖 Generated with Claude Code

https://claude.ai/code/session_0187PLDdwErbQsm6aZi7KXuH

Task: NEC-101 — https://projectron.nerchure.com/tasks/101 The identity host went down on 24 September 2026. Its disk filled with about a hundred old builds, because the deployer never deletes an image. `/api/ready` failed on the disk floor. The deployer read that as a bad build and rolled back, and the rollback failed the same way. It then set `AUTO_DEPLOY=off` and posted a status-page incident. Recovered by hand; this stops it recurring. - **Prune.** Keeps the running build, its rollback target (from `history.log`) and `:stable`. Builds are matched by revision label, because `:stable` is pulled by digest and has no tag. Runs before each pull, after a successful deploy and on every tick that does not deploy, so an upgraded host clears its backlog on the first tick. Only `$IMAGE`, never `rmi -f`, never fatal. Also `deployer.sh prune [--dry-run]`. - **Disk-bound deploys.** A build that reports its own sha but fails readiness with `"reason":"disk"` prunes and retries. If still full it records `disk-bound` and pushes `down`, and does not roll back or turn itself off. - **Early warning.** The heartbeat adds `disk low: N MB free` while free space is under twice the floor. `status` gains `disk_free`, `disk_floor` and `local_builds`. Tests: 15 new ones in `deploy/test/deployer.test.ts`, with fake `docker images`/`rmi`/`df` and readiness failing on disk. Each rule was mutation-checked; a substring-keep mutation survived at first and the test was tightened. The real `docker images`/`image inspect` output was checked read-only against eu1: duplicated digest refs, 15 pre-label images, `df -Pk` parsing. Not yet verified live. After merge, `install-deployer.sh` carries this to the hosts (it is not automatic). Run `deployer.sh prune --dry-run` on eu1 first: it still holds ~100 old builds, left on purpose as test data. Follow-up, as a separate PR: a CI registry build cache, plus moving `GIT_SHA` to the end of the runtime stages so builds share layers. ## Changelog NONE 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_0187PLDdwErbQsm6aZi7KXuH
Delete old images from the deployer, and stop reading a full disk as a bad build
All checks were successful
Release note / release-note (pull_request) Successful in 12s
CI / e2e (pull_request) Successful in 5m10s
CI / build (pull_request) Successful in 5m28s
CI / promote (pull_request) Has been skipped
e34919a549
On 24 September 2026 the identity host filled its 38 GB disk with about a
hundred old builds, because nothing ever deleted an image. Readiness failed
on its disk floor. The deployer rolled back, the rollback met the same floor,
and the failure path switched auto-deploy off and posted a public incident.

The deployer now keeps three builds of its own image: the running one, its
rollback target and :stable, matched by revision label because :stable
arrives by digest with no tag. It removes the rest before each pull, after a
successful deploy and on every tick that does not deploy. A readiness failure
whose reason is the disk prunes and asks again, and never rolls back or turns
itself off. The heartbeat says when free space is under twice the floor.

Task: NEC-101

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0187PLDdwErbQsm6aZi7KXuH
cruelacid force-pushed worktree-image-housekeeping from e34919a549
All checks were successful
Release note / release-note (pull_request) Successful in 12s
CI / e2e (pull_request) Successful in 5m10s
CI / build (pull_request) Successful in 5m28s
CI / promote (pull_request) Has been skipped
to 8974ec37e9
Some checks failed
CI / build (pull_request) Failing after 1m23s
Release note / release-note (pull_request) Successful in 11s
CI / promote (pull_request) Has been cancelled
CI / e2e (pull_request) Has been cancelled
2026-09-24 21:30:31 +01:00
Compare
cruelacid force-pushed worktree-image-housekeeping from 8974ec37e9
Some checks failed
CI / build (pull_request) Failing after 1m23s
Release note / release-note (pull_request) Successful in 11s
CI / promote (pull_request) Has been cancelled
CI / e2e (pull_request) Has been cancelled
to 15c3d34593
All checks were successful
Release note / release-note (pull_request) Successful in 13s
CI / build (pull_request) Successful in 5m22s
CI / e2e (pull_request) Successful in 5m31s
CI / promote (pull_request) Has been skipped
CI / e2e (push) Successful in 5m6s
CI / build (push) Successful in 5m52s
CI / promote (push) Successful in 34s
2026-09-24 21:33:04 +01:00
Compare
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
Nectenda/nectenda!158
No description provided.