Nextcloud Office  ›  Operations

Operations

Day-2 ops: backups, monitoring, recovery procedures, and the roll-forward / roll-back plans.

Backup pipeline

Four files written daily to /srv/nc-files/backups/ on homework03 (NFS, real path /slab/container_storage/office/backups/ on desslok):

FileContentsTypical sizeRecovery use
office-YYYYMMDD-pgdump.sql.gz PostgreSQL full dump via pg_dumpall from the AIO database container. All ~155 Nextcloud tables. ~600 KB (empty) → grows with users/files Restore the database after a Nextcloud corruption or migration to new hardware.
office-YYYYMMDD-aio-config.tar.gz The mastercontainer's configuration.json (office suite choice, domain, datadir, passwords) + database-dump bind target. ~6 KB Reconstruct the AIO install state without going through the setup wizard again.
office-YYYYMMDD-aio-nextcloud-app.tar.gz Tar of /usr/local/containers/nextcloudaio/nextcloud-aio-nextcloud/ — the AIO-managed local app volume (Nextcloud app code, installed apps, config/). ~200-500 MB depending on installed apps Survives a fresh AIO install: restore this AND the config tarball to skip the entire setup wizard and preserve installed apps.
office-YYYYMMDD-ncdata.tar.gz Tar of /srv/nc-files/ root (user files: admin/, race/, appdata_*/, etc.) — excludes backups/ to avoid recursion. Empty (~100 B) until users upload files, then grows Restore user files after data loss.

Daily cron schedule

Triggered by a systemd timer on hector, daily at 03:30 UTC (with up to 15 min random delay). The unit SSHes into homework03 (no password prompt — keys only) and runs the script with sudo.

ssh tigo@hector
systemctl list-timers office-backup*
# Expect: NEXT shown for the next 03:30 UTC ± 15 min

# Manual trigger for testing
sudo -n systemctl start office-backup.service
sleep 30
systemctl status office-backup.service | head -5
# Expect: Active: inactive (dead) → success

Retention policy

14 days. The script prunes via find ... -mtime +14 -delete after the daily write. Same-day reruns overwrite (date-only stamp) — intentional; we don't want to keep multiple copies per day.

What this doesn't cover

Monitoring & alerting

Gatus endpoints to watch

Gatus runs on monitor (10.0.0.75), port 10010. Suggested checks for the Nextcloud stack:

EndpointWhatSeverity
https://office.rmf44.xyz/loginPublic ingress (Caddy → Apache → PHP-FPM)P1 outage
https://100.79.142.164:8443AIO admin UI (mastercontainer direct)P2 if down
docker stats — nextcloud-aio-*Container healthP2 if any restart loop
NFS — /srv/nc-files on homework03Mount up + writableP1 (data loss risk)

The backup pipeline's last-run status is readable via systemctl status office-backup.service on hector; adding a Gatus check on this is straightforward via SSH exec.

Recovery procedures

Restore from a daily backup

Full restore assumes a clean homework03 + intact NFS on desslok.

  1. Stop the AIO stack:
    ssh homework03
    cd /usr/local/containers/nextcloudaio
    sudo -n docker compose down
  2. Restore the AIO config (replaces configuration.json):
    LATEST=$(ls -t /srv/nc-files/backups/office-*-aio-config.tar.gz | head -1)
    tar -C /usr/local/containers/nextcloudaio -xzf "$LATEST"
  3. Restore the local nextcloud app volume (AIO-managed code + installed apps):
    LATEST=$(ls -t /srv/nc-files/backups/office-*-aio-nextcloud-app.tar.gz | head -1)
    tar -C /usr/local/containers/nextcloudaio -xzf "$LATEST"
  4. Restore user files (NFS root, no intermediate nextcloud/):
    LATEST=$(ls -t /srv/nc-files/backups/office-*-ncdata.tar.gz | head -1)
    # Tar contains files at root (admin/, race/, appdata_*/, ...)
    tar -C /srv/nc-files -xzf "$LATEST"
  5. Restore the database (drop + reload):
    # Start only the database container first
    sudo -n docker compose up -d nextcloud-aio-mastercontainer
    sleep 30
    # Wait for the database container to come up via mastercontainer
    sudo -n docker exec nextcloud-aio-database pg_isready -U nextcloud
    LATEST=$(ls -t /srv/nc-files/backups/office-*-pgdump.sql.gz | head -1)
    zcat "$LATEST" | sudo -n docker exec -i nextcloud-aio-database psql -U nextcloud -d nextcloud_database
  6. Restart the AIO stack:
    sudo -n docker compose restart
    sleep 60
    curl -skI https://office.rmf44.xyz/login
    # Expect: HTTP/2 200

Restore a single file

No need for a full restore — just untar one file:

ssh desslok
LATEST=$(ls -t /slab/container_storage/office/backups/office-*-ncdata.tar.gz | head -1)
# Tar contains files at root; restore one user's file:
tar -C / -xzf "$LATEST" race/files/path/to/file
# Adjust for the user + path; user dirs are at NFS root (no nextcloud/ prefix)

Re-initialize the admin user

If the admin password is lost:

ssh homework03
# Reset via OCC
sudo -n docker exec -u www-data nextcloud-aio-nextcloud \
  php /var/www/html/occ user:resetpassword admin --password-from-env
# Reads password from NEXTCLOUD_ADMIN_PASSWORD env var
# (default: same as setup wizard)

Updates & upgrades

AIO manages its own updates: when a new all-in-one image is published, mastercontainer pulls the new image and triggers a rolling update of all side containers.

To manually trigger an update:

ssh homework03
cd /usr/local/containers/nextcloudaio
sudo -n docker compose pull
sudo -n docker compose up -d
# Wait 5-10 min for all side containers to roll

Before a major update, take a manual backup: sudo -n systemctl start office-backup.service on hector, then verify the files exist on desslok before pulling new images.

Rollback (revert to OnlyOffice)

The OnlyOffice container was retired on 2026-08-10. To bring it back, you'd need the saved tarball at /home/tigo/onlyoffice-stack-backup-20260810.tar.gz on hawker. Rollback time estimate: ~30 minutes (restore compose, start containers, restore Caddy vhost, smoke test).

Recommendation: keep that tarball for at least one more month, then archive to cold storage. If the new AIO stack proves stable, drop the tarball after that.

Roll-forward (move to dedicated AIO host)

The current 15 GB homework03 is tight on RAM. If we add Talk or Fulltextsearch later, the host won't fit. To roll forward to a bigger host:

  1. Stop AIO on homework03 (preserve data on desslok via NFS).
  2. Provision a bigger host (recommend: 32 GB RAM, NVMe).
  3. Mount the same NFS export at the same path.
  4. Copy /usr/local/containers/nextcloudaio/ over (or rebuild from the saved aio-config.tar.gz).
  5. Update Caddy upstream IP on hawker.
  6. Run a manual backup immediately to confirm the new host can write to the same NFS.

Append-only references