# Docker Maintenance — Growth Guard (kb#184) ## Background 2026-07-24: root LV (`/`, holds Docker's data-root) filled to 1.5G free / 100% used — risk of ENOSPC corruption across every service on Agap. Emergency reclaim (builder prune + dangling images + stopped containers) took it to 25G free / 90% used. 2026-07-26 (this task): re-measured before any action — still 24G free / 90% used (the emergency fix hadn't regrown, but hadn't been made recurring either). Ran `prune.sh` live: root went to **40G free / 83% used**. Breakdown before/after: | | Before | After | |---|---|---| | Images (total/reclaimable) | 88GB / 62.41GB (70%) | 87.2GB / 62.41GB (71%) | | Build Cache (total/reclaimable) | 18.74GB / 14.49GB | 4.25GB / 0B | | Root free | 24G (90% used) | 40G (83% used) | The 62.41GB of image "reclaimable" space barely moved — that's almost entirely **tagged-but-unused** images (`docker image prune -a` territory), not dangling ones. See the decision below. ## What `prune.sh` does (safe scope only) ``` docker builder prune -f # build cache — always safe, fully rebuildable docker container prune -f # stopped/exited containers only docker image prune -f # DANGLING images only (untagged) — no -a ``` Logs before/after `docker system df` to `prune.log`, and pushes the post-prune build-cache-reclaimable size (bytes) to Zabbix host `AgapHost` (hostid 10776), trapper item `docker.buildcache.reclaimable.bytes` (itemid 70624, **type 2 — confirmed to accept `history.push` on this Zabbix instance, unlike the type-2 "Calculated" items other tasks found dead; verified live 2026-07-26**). Test without changing anything: ```bash bash /home/alvis/agap_git/docker-maintenance/prune.sh --dry-run ``` ## Explicitly NOT automated (human decision required) - **`docker image prune -a`** — would reclaim ~62GB of tagged-but-currently-unused images (e.g. `cognee/cognee-mcp:1.2.2` at 17.9GB, `cognee` variants, old `ghcr.io/open-webui/open-webui`, `lscr.io/linuxserver/calibre`/`zotero` at 3.5GB each — many belong to services that are stopped/replaced but the image may still be wanted for a quick restart). **Decision recorded, not made**: someone with knowledge of which of these services are truly retired should either (a) run `docker image prune -a` manually after reviewing the image list (`docker system df -v`), or (b) curate a keep-list and prune around it. Not scheduled — this script will never run `-a`. - **`docker volume prune`** — never automated; can destroy live data for a volume that's temporarily unmounted. Not touched by this script or its guard. ## Zabbix - Item: `docker.buildcache.reclaimable.bytes` on host `AgapHost` (hostid 10776, itemid 70624), type Trapper, units B. Pushed once per `prune.sh` run (post-prune value — after a scheduled run this should almost always read near 0). - Trigger (triggerid 32970): fires if `last()` of that item exceeds 20GB (21474836480 bytes) — i.e. the build cache grew back past the safe threshold between scheduled runs, meaning the timer isn't running or needs to run more often. - Zabbix runs on **lizacer** (`http://192.168.1.4:81`), not Agap. ## Scheduled run — NOT INSTALLED (handoff required) This directory contains the tested script only. Installing the schedule is a human step (per Agap automation policy: no unattended agent installs a cron/ timer that acts on a live target). To install: **Cron** (matches the existing `agap_git/kanboard` pattern — user crontab, no sudo): ```bash crontab -e # add: 0 4 * * * /home/alvis/agap_git/docker-maintenance/prune.sh >> /home/alvis/agap_git/docker-maintenance/prune.log 2>&1 ``` Runs daily at 04:00. Adjust frequency if build-cache growth between runs proves faster than expected (watch the Zabbix trigger above). Alternative — **systemd user timer**, unit files provided in this directory (`docker-prune.service`, `docker-prune.timer`), not installed: ```bash mkdir -p ~/.config/systemd/user cp /home/alvis/agap_git/docker-maintenance/docker-prune.service ~/.config/systemd/user/ cp /home/alvis/agap_git/docker-maintenance/docker-prune.timer ~/.config/systemd/user/ systemctl --user daemon-reload systemctl --user enable --now docker-prune.timer ``` Either mechanism is fine; cron matches existing Agap convention (kanboard healthcheck/backup) so it's the recommended default. ## Acceptance status (kb#184) - [x] Immediate reclaim (2026-07-24, prior task) — done. - [x] Recurring prune script — written, tested (dry-run + live run confirmed it reclaims builder cache + dangling images safely). - [ ] **Schedule installed** — script + cron line + systemd unit files are ready; installing them is a human step (see above). **Handoff: run the `crontab -e` command above, or install the systemd timer.** - [x] Build-cache Zabbix trigger — created (triggerid 32970), verified `history.push` lands (item read back 14490000000 during test, then 0 after the live prune run). - [x] `docker image prune -a` decision — recorded above as an open human decision, not automated. - [x] df / free space — 40G free (83% used) as of 2026-07-26, up from 24G/90%.