Day-2 ops: backups, monitoring, recovery procedures, and the roll-forward / roll-back plans.
Three files written daily to
/srv/nc-files/backups/ on homework03 (NFS, real path
/slab/container_storage/office/backups/ on desslok):
| File | Contents | Typical size | Recovery use |
|---|---|---|---|
office-YYYYMMDD-pgdump.sql.gz |
PostgreSQL full dump via pg_dumpall from the AIO database container. All ~155 Nextcloud tables. |
~600 KB (empty) → grows with users/files | Restore the database after a Nextcloud corruption or migration to new hardware. |
office-YYYYMMDD-aio-config.tar.gz |
The mastercontainer's configuration.json (office suite choice, domain, datadir, passwords) + database-dump bind target. |
~6 KB | Reconstruct the AIO install state without going through the setup wizard again. |
office-YYYYMMDD-ncdata.tar.gz |
Tar of /srv/nc-files/nextcloud/ (user-uploaded files) — excludes backups/ to avoid recursion. |
Empty (~100 B) until users upload files, then grows | Restore user files after data loss. |
Triggered by a systemd timer on hector, daily at
03:30 UTC (with up to 15 min random delay). The unit SSHes into
homework03 (no password prompt — keys only) and runs the script
with sudo.
ssh tigo@hector
systemctl list-timers office-backup*
# Expect: NEXT shown for the next 03:30 UTC ± 15 min
# Manual trigger for testing
sudo -n systemctl start office-backup.service
sleep 30
systemctl status office-backup.service | head -5
# Expect: Active: inactive (dead) → success
14 days. The script prunes via find ... -mtime +14 -delete
after the daily write. Same-day reruns overwrite (date-only stamp)
— intentional; we don't want to keep multiple copies per day.
configuration.json + the NFS data,
but you'll lose any state stored in those volumes (e.g. the
mastercontainer's domain-validation certificates cache). In
practice these regenerate on first boot.
/srv/nc-files/nextcloud/ while the filesystem is
actively being written to by the nextcloud container. The tar
will see a consistent enough snapshot for crash-consistent
recovery; for true point-in-time recovery, you'd want to
quiesce Nextcloud (set maintenance mode) for the duration of
the tar, which we haven't done.
Gatus runs on monitor (10.0.0.75), port 10010.
Suggested checks for the Nextcloud stack:
| Endpoint | What | Severity |
|---|---|---|
https://office.rmf44.xyz/login | Public ingress (Caddy → Apache → PHP-FPM) | P1 outage |
https://100.79.142.164:8443 | AIO admin UI (mastercontainer direct) | P2 if down |
docker stats — nextcloud-aio-* | Container health | P2 if any restart loop |
NFS — /srv/nc-files on homework03 | Mount up + writable | P1 (data loss risk) |
The backup pipeline's last-run status is readable via
systemctl status office-backup.service on hector;
adding a Gatus check on this is straightforward via SSH exec.
Full restore assumes a clean homework03 + intact NFS on desslok.
ssh homework03
cd /usr/local/containers/nextcloudaio
sudo -n docker compose down
LATEST=$(ls -t /srv/nc-files/backups/office-*-aio-config.tar.gz | head -1)
tar -C /usr/local/containers/nextcloudaio -xzf "$LATEST"
nextcloud/):
LATEST=$(ls -t /srv/nc-files/backups/office-*-ncdata.tar.gz | head -1)
# Tar contains /nextcloud/ at root
tar -C /srv/nc-files -xzf "$LATEST"
# Start only the database container first
sudo -n docker compose up -d nextcloud-aio-mastercontainer
sleep 30
# Wait for the database container to come up via mastercontainer
sudo -n docker exec nextcloud-aio-database pg_isready -U nextcloud
LATEST=$(ls -t /srv/nc-files/backups/office-*-pgdump.sql.gz | head -1)
zcat "$LATEST" | sudo -n docker exec -i nextcloud-aio-database psql -U nextcloud -d nextcloud_database
sudo -n docker compose restart
sleep 60
curl -skI https://office.rmf44.xyz/login
# Expect: HTTP/2 200
No need for a full restore — just untar one file:
ssh desslok
LATEST=$(ls -t /slab/container_storage/office/backups/office-*-ncdata.tar.gz | head -1)
tar -C / -xzf "$LATEST" nextcloud/admin/files/path/to/file
# Adjust for the user + path
If the admin password is lost:
ssh homework03
# Reset via OCC
sudo -n docker exec -u www-data nextcloud-aio-nextcloud \
php /var/www/html/occ user:resetpassword admin --password-from-env
# Reads password from NEXTCLOUD_ADMIN_PASSWORD env var
# (default: same as setup wizard)
AIO manages its own updates: when a new all-in-one
image is published, mastercontainer pulls the new image and
triggers a rolling update of all side containers.
To manually trigger an update:
ssh homework03
cd /usr/local/containers/nextcloudaio
sudo -n docker compose pull
sudo -n docker compose up -d
# Wait 5-10 min for all side containers to roll
Before a major update, take a manual backup:
sudo -n systemctl start office-backup.service on
hector, then verify the files exist on desslok before pulling
new images.
The OnlyOffice container was retired on 2026-08-10. To bring it
back, you'd need the saved tarball at
/home/tigo/onlyoffice-stack-backup-20260810.tar.gz
on hawker. Rollback time estimate: ~30 minutes (restore compose,
start containers, restore Caddy vhost, smoke test).
Recommendation: keep that tarball for at least one more month, then archive to cold storage. If the new AIO stack proves stable, drop the tarball after that.
The current 15 GB homework03 is tight on RAM. If we add Talk or Fulltextsearch later, the host won't fit. To roll forward to a bigger host:
/usr/local/containers/nextcloudaio/ over (or
rebuild from the saved aio-config.tar.gz).docs/skill/nextcloud-aio-deploy (Hermes skill) — abbreviated deploy workflow