docker-maintenance/ adds a systemd timer + prune.sh for the root LV that holds Docker's data-root and has filled to 100% before, risking ENOSPC corruption. The script sticks to the safe reclaim set (builder cache, dangling images, stopped containers) and deliberately avoids `-a` and volume pruning, which can destroy live data when run unattended. kanboard/backup.sh and healthcheck.sh bring Kanboard in line with the other services. seafile/ and vaultwarden/ backup scripts get fixes carried from the stability audit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
5.1 KiB
Docker Maintenance — Growth Guard (kb#184)
Background
2026-07-24: root LV (/, holds Docker's data-root) filled to 1.5G free / 100% used —
risk of ENOSPC corruption across every service on Agap. Emergency reclaim (builder
prune + dangling images + stopped containers) took it to 25G free / 90% used.
2026-07-26 (this task): re-measured before any action — still 24G free / 90% used
(the emergency fix hadn't regrown, but hadn't been made recurring either). Ran
prune.sh live: root went to 40G free / 83% used. Breakdown before/after:
| Before | After | |
|---|---|---|
| Images (total/reclaimable) | 88GB / 62.41GB (70%) | 87.2GB / 62.41GB (71%) |
| Build Cache (total/reclaimable) | 18.74GB / 14.49GB | 4.25GB / 0B |
| Root free | 24G (90% used) | 40G (83% used) |
The 62.41GB of image "reclaimable" space barely moved — that's almost entirely
tagged-but-unused images (docker image prune -a territory), not dangling ones.
See the decision below.
What prune.sh does (safe scope only)
docker builder prune -f # build cache — always safe, fully rebuildable
docker container prune -f # stopped/exited containers only
docker image prune -f # DANGLING images only (untagged) — no -a
Logs before/after docker system df to prune.log, and pushes the post-prune
build-cache-reclaimable size (bytes) to Zabbix host AgapHost (hostid 10776),
trapper item docker.buildcache.reclaimable.bytes (itemid 70624, type 2 —
confirmed to accept history.push on this Zabbix instance, unlike the type-2
"Calculated" items other tasks found dead; verified live 2026-07-26).
Test without changing anything:
bash /home/alvis/agap_git/docker-maintenance/prune.sh --dry-run
Explicitly NOT automated (human decision required)
docker image prune -a— would reclaim ~62GB of tagged-but-currently-unused images (e.g.cognee/cognee-mcp:1.2.2at 17.9GB,cogneevariants, oldghcr.io/open-webui/open-webui,lscr.io/linuxserver/calibre/zoteroat 3.5GB each — many belong to services that are stopped/replaced but the image may still be wanted for a quick restart). Decision recorded, not made: someone with knowledge of which of these services are truly retired should either (a) rundocker image prune -amanually after reviewing the image list (docker system df -v), or (b) curate a keep-list and prune around it. Not scheduled — this script will never run-a.docker volume prune— never automated; can destroy live data for a volume that's temporarily unmounted. Not touched by this script or its guard.
Zabbix
- Item:
docker.buildcache.reclaimable.byteson hostAgapHost(hostid 10776, itemid 70624), type Trapper, units B. Pushed once perprune.shrun (post-prune value — after a scheduled run this should almost always read near 0). - Trigger (triggerid 32970): fires if
last()of that item exceeds 20GB (21474836480 bytes) — i.e. the build cache grew back past the safe threshold between scheduled runs, meaning the timer isn't running or needs to run more often. - Zabbix runs on lizacer (
http://192.168.1.4:81), not Agap.
Scheduled run — NOT INSTALLED (handoff required)
This directory contains the tested script only. Installing the schedule is a human step (per Agap automation policy: no unattended agent installs a cron/ timer that acts on a live target). To install:
Cron (matches the existing agap_git/kanboard pattern — user crontab, no sudo):
crontab -e
# add:
0 4 * * * /home/alvis/agap_git/docker-maintenance/prune.sh >> /home/alvis/agap_git/docker-maintenance/prune.log 2>&1
Runs daily at 04:00. Adjust frequency if build-cache growth between runs proves faster than expected (watch the Zabbix trigger above).
Alternative — systemd user timer, unit files provided in this directory
(docker-prune.service, docker-prune.timer), not installed:
mkdir -p ~/.config/systemd/user
cp /home/alvis/agap_git/docker-maintenance/docker-prune.service ~/.config/systemd/user/
cp /home/alvis/agap_git/docker-maintenance/docker-prune.timer ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now docker-prune.timer
Either mechanism is fine; cron matches existing Agap convention (kanboard healthcheck/backup) so it's the recommended default.
Acceptance status (kb#184)
- Immediate reclaim (2026-07-24, prior task) — done.
- Recurring prune script — written, tested (dry-run + live run confirmed it reclaims builder cache + dangling images safely).
- Schedule installed — script + cron line + systemd unit files are ready;
installing them is a human step (see above). Handoff: run the
crontab -ecommand above, or install the systemd timer. - Build-cache Zabbix trigger — created (triggerid 32970), verified
history.pushlands (item read back 14490000000 during test, then 0 after the live prune run). docker image prune -adecision — recorded above as an open human decision, not automated.- df / free space — 40G free (83% used) as of 2026-07-26, up from 24G/90%.