- topology.svg, container-tree.svg, data-flow.svg: show /srv/nc-files bound directly into nextcloud container at /mnt/ncdata (no intermediate /mnt/nc-data/nextcloud-data layer). - architecture.html: explain rewire, update storage table, remove pre-rewire 'why two layers' section. - procedure.html: NEXTCLOUD_DATADIR=/srv/nc-files, no mastercontainer bind, add §2.5 Rewire 2026-08-11 with recovery procedure, update backup script to also tar local nextcloud app volume, fix 4.3 verification (302 -> /login, status.php, file visibility). - operations.html: 4 backup files now (added aio-nextcloud-app.tar.gz), restore procedures updated, single-file restore uses new tar layout. - troubleshooting.html: add 3 new sections — appdata-missing, nextcloud-not-spawning, collabora-discovery-warning. - request-flow.svg: remove nextcloud/ subdir from NFS write path. - index.html, README.md: update paths and metadata.
271 lines
10 KiB
HTML
271 lines
10 KiB
HTML
<!DOCTYPE html>
|
||
<html lang="en">
|
||
<head>
|
||
<meta charset="UTF-8">
|
||
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||
<title>Operations — Nextcloud Office</title>
|
||
<link rel="stylesheet" href="assets/style.css">
|
||
</head>
|
||
<body>
|
||
<div id="wrapper">
|
||
|
||
<div class="crumbs"><a href="index.html">Nextcloud Office</a> › Operations</div>
|
||
|
||
<ul class="nav">
|
||
<li><a href="index.html">Overview</a></li>
|
||
<li><a href="architecture.html">Architecture</a></li>
|
||
<li><a href="procedure.html">Procedure</a></li>
|
||
<li><a href="troubleshooting.html">Troubleshooting</a></li>
|
||
<li><a href="operations.html" class="active">Operations</a></li>
|
||
</ul>
|
||
|
||
<h1>Operations</h1>
|
||
<p>
|
||
Day-2 ops: backups, monitoring, recovery procedures, and the
|
||
roll-forward / roll-back plans.
|
||
</p>
|
||
|
||
<h2>Backup pipeline</h2>
|
||
|
||
<p>
|
||
Four files written daily to
|
||
<code>/srv/nc-files/backups/</code> on homework03 (NFS, real path
|
||
<code>/slab/container_storage/office/backups/</code> on desslok):
|
||
</p>
|
||
|
||
<table>
|
||
<tr><th>File</th><th>Contents</th><th>Typical size</th><th>Recovery use</th></tr>
|
||
<tr>
|
||
<td><code>office-YYYYMMDD-pgdump.sql.gz</code></td>
|
||
<td>PostgreSQL full dump via <code>pg_dumpall</code> from the AIO database container. All ~155 Nextcloud tables.</td>
|
||
<td>~600 KB (empty) → grows with users/files</td>
|
||
<td>Restore the database after a Nextcloud corruption or migration to new hardware.</td>
|
||
</tr>
|
||
<tr>
|
||
<td><code>office-YYYYMMDD-aio-config.tar.gz</code></td>
|
||
<td>The mastercontainer's <code>configuration.json</code> (office suite choice, domain, datadir, passwords) + database-dump bind target.</td>
|
||
<td>~6 KB</td>
|
||
<td>Reconstruct the AIO install state without going through the setup wizard again.</td>
|
||
</tr>
|
||
<tr>
|
||
<td><code>office-YYYYMMDD-aio-nextcloud-app.tar.gz</code></td>
|
||
<td>Tar of <code>/usr/local/containers/nextcloudaio/nextcloud-aio-nextcloud/</code> — the AIO-managed local app volume (Nextcloud app code, installed apps, <code>config/</code>).</td>
|
||
<td>~200-500 MB depending on installed apps</td>
|
||
<td>Survives a fresh AIO install: restore this AND the config tarball to skip the entire setup wizard and preserve installed apps.</td>
|
||
</tr>
|
||
<tr>
|
||
<td><code>office-YYYYMMDD-ncdata.tar.gz</code></td>
|
||
<td>Tar of <code>/srv/nc-files/</code> root (user files: <code>admin/</code>, <code>race/</code>, <code>appdata_*/</code>, etc.) — excludes <code>backups/</code> to avoid recursion.</td>
|
||
<td>Empty (~100 B) until users upload files, then grows</td>
|
||
<td>Restore user files after data loss.</td>
|
||
</tr>
|
||
</table>
|
||
|
||
<h3>Daily cron schedule</h3>
|
||
<p>
|
||
Triggered by a systemd timer on <code>hector</code>, daily at
|
||
03:30 UTC (with up to 15 min random delay). The unit SSHes into
|
||
homework03 (no password prompt — keys only) and runs the script
|
||
with <code>sudo</code>.
|
||
</p>
|
||
|
||
<pre><code>ssh tigo@hector
|
||
systemctl list-timers office-backup*
|
||
# Expect: NEXT shown for the next 03:30 UTC ± 15 min
|
||
|
||
# Manual trigger for testing
|
||
sudo -n systemctl start office-backup.service
|
||
sleep 30
|
||
systemctl status office-backup.service | head -5
|
||
# Expect: Active: inactive (dead) → success</code></pre>
|
||
|
||
<h3>Retention policy</h3>
|
||
<p>
|
||
14 days. The script prunes via <code>find ... -mtime +14 -delete</code>
|
||
after the daily write. Same-day reruns overwrite (date-only stamp)
|
||
— intentional; we don't want to keep multiple copies per day.
|
||
</p>
|
||
|
||
<h3>What this doesn't cover</h3>
|
||
<ul>
|
||
<li>
|
||
<strong>Container runtime state</strong> — AIO's named volumes
|
||
on local ext4 are NOT backed up by this pipeline. If homework03
|
||
loses its disk, the AIO setup wizard will rebuild containers
|
||
from the saved <code>configuration.json</code> + the NFS data,
|
||
but you'll lose any state stored in those volumes (e.g. the
|
||
mastercontainer's domain-validation certificates cache). In
|
||
practice these regenerate on first boot.
|
||
</li>
|
||
<li>
|
||
<strong>NFS quiescence</strong> — the tar reads
|
||
<code>/srv/nc-files/</code> (the NFS root) while the
|
||
filesystem is actively being written to by the nextcloud
|
||
container. The tar will see a consistent enough snapshot for
|
||
crash-consistent recovery; for true point-in-time recovery,
|
||
you'd want to quiesce Nextcloud (set maintenance mode) for the
|
||
duration of the tar, which we haven't done.
|
||
</li>
|
||
</ul>
|
||
|
||
<h2>Monitoring & alerting</h2>
|
||
|
||
<h3>Gatus endpoints to watch</h3>
|
||
<p>
|
||
Gatus runs on <code>monitor (10.0.0.75)</code>, port 10010.
|
||
Suggested checks for the Nextcloud stack:
|
||
</p>
|
||
<table>
|
||
<tr><th>Endpoint</th><th>What</th><th>Severity</th></tr>
|
||
<tr><td><code>https://office.rmf44.xyz/login</code></td><td>Public ingress (Caddy → Apache → PHP-FPM)</td><td>P1 outage</td></tr>
|
||
<tr><td><code>https://100.79.142.164:8443</code></td><td>AIO admin UI (mastercontainer direct)</td><td>P2 if down</td></tr>
|
||
<tr><td>docker stats — <code>nextcloud-aio-*</code></td><td>Container health</td><td>P2 if any restart loop</td></tr>
|
||
<tr><td>NFS — <code>/srv/nc-files</code> on homework03</td><td>Mount up + writable</td><td>P1 (data loss risk)</td></tr>
|
||
</table>
|
||
|
||
<p>
|
||
The backup pipeline's last-run status is readable via
|
||
<code>systemctl status office-backup.service</code> on hector;
|
||
adding a Gatus check on this is straightforward via SSH exec.
|
||
</p>
|
||
|
||
<h2>Recovery procedures</h2>
|
||
|
||
<h3>Restore from a daily backup</h3>
|
||
<p>
|
||
Full restore assumes a clean homework03 + intact NFS on desslok.
|
||
</p>
|
||
<ol>
|
||
<li>
|
||
Stop the AIO stack:
|
||
<pre><code>ssh homework03
|
||
cd /usr/local/containers/nextcloudaio
|
||
sudo -n docker compose down</code></pre>
|
||
</li>
|
||
<li>
|
||
Restore the AIO config (replaces configuration.json):
|
||
<pre><code>LATEST=$(ls -t /srv/nc-files/backups/office-*-aio-config.tar.gz | head -1)
|
||
tar -C /usr/local/containers/nextcloudaio -xzf "$LATEST"</code></pre>
|
||
</li>
|
||
<li>
|
||
Restore the local nextcloud app volume (AIO-managed code +
|
||
installed apps):
|
||
<pre><code>LATEST=$(ls -t /srv/nc-files/backups/office-*-aio-nextcloud-app.tar.gz | head -1)
|
||
tar -C /usr/local/containers/nextcloudaio -xzf "$LATEST"</code></pre>
|
||
</li>
|
||
<li>
|
||
Restore user files (NFS root, no intermediate <code>nextcloud/</code>):
|
||
<pre><code>LATEST=$(ls -t /srv/nc-files/backups/office-*-ncdata.tar.gz | head -1)
|
||
# Tar contains files at root (admin/, race/, appdata_*/, ...)
|
||
tar -C /srv/nc-files -xzf "$LATEST"</code></pre>
|
||
</li>
|
||
<li>
|
||
Restore the database (drop + reload):
|
||
<pre><code># Start only the database container first
|
||
sudo -n docker compose up -d nextcloud-aio-mastercontainer
|
||
sleep 30
|
||
# Wait for the database container to come up via mastercontainer
|
||
sudo -n docker exec nextcloud-aio-database pg_isready -U nextcloud
|
||
LATEST=$(ls -t /srv/nc-files/backups/office-*-pgdump.sql.gz | head -1)
|
||
zcat "$LATEST" | sudo -n docker exec -i nextcloud-aio-database psql -U nextcloud -d nextcloud_database</code></pre>
|
||
</li>
|
||
<li>
|
||
Restart the AIO stack:
|
||
<pre><code>sudo -n docker compose restart
|
||
sleep 60
|
||
curl -skI https://office.rmf44.xyz/login
|
||
# Expect: HTTP/2 200</code></pre>
|
||
</li>
|
||
</ol>
|
||
|
||
<h3>Restore a single file</h3>
|
||
<p>
|
||
No need for a full restore — just untar one file:
|
||
</p>
|
||
<pre><code>ssh desslok
|
||
LATEST=$(ls -t /slab/container_storage/office/backups/office-*-ncdata.tar.gz | head -1)
|
||
# Tar contains files at root; restore one user's file:
|
||
tar -C / -xzf "$LATEST" race/files/path/to/file
|
||
# Adjust for the user + path; user dirs are at NFS root (no nextcloud/ prefix)</code></pre>
|
||
|
||
<h3>Re-initialize the admin user</h3>
|
||
<p>
|
||
If the admin password is lost:
|
||
</p>
|
||
<pre><code>ssh homework03
|
||
# Reset via OCC
|
||
sudo -n docker exec -u www-data nextcloud-aio-nextcloud \
|
||
php /var/www/html/occ user:resetpassword admin --password-from-env
|
||
# Reads password from NEXTCLOUD_ADMIN_PASSWORD env var
|
||
# (default: same as setup wizard)</code></pre>
|
||
|
||
<h2>Updates & upgrades</h2>
|
||
|
||
<p>
|
||
AIO manages its own updates: when a new <code>all-in-one</code>
|
||
image is published, mastercontainer pulls the new image and
|
||
triggers a rolling update of all side containers.
|
||
</p>
|
||
|
||
<p>
|
||
To manually trigger an update:
|
||
</p>
|
||
<pre><code>ssh homework03
|
||
cd /usr/local/containers/nextcloudaio
|
||
sudo -n docker compose pull
|
||
sudo -n docker compose up -d
|
||
# Wait 5-10 min for all side containers to roll</code></pre>
|
||
|
||
<p>
|
||
<strong>Before a major update</strong>, take a manual backup:
|
||
<code>sudo -n systemctl start office-backup.service</code> on
|
||
hector, then verify the files exist on desslok before pulling
|
||
new images.
|
||
</p>
|
||
|
||
<h2>Rollback (revert to OnlyOffice)</h2>
|
||
|
||
<p>
|
||
The OnlyOffice container was retired on 2026-08-10. To bring it
|
||
back, you'd need the saved tarball at
|
||
<code>/home/tigo/onlyoffice-stack-backup-20260810.tar.gz</code>
|
||
on hawker. Rollback time estimate: ~30 minutes (restore compose,
|
||
start containers, restore Caddy vhost, smoke test).
|
||
</p>
|
||
|
||
<p>
|
||
<strong>Recommendation:</strong> keep that tarball for at least
|
||
one more month, then archive to cold storage. If the new AIO
|
||
stack proves stable, drop the tarball after that.
|
||
</p>
|
||
|
||
<h2>Roll-forward (move to dedicated AIO host)</h2>
|
||
|
||
<p>
|
||
The current 15 GB homework03 is tight on RAM. If we add Talk or
|
||
Fulltextsearch later, the host won't fit. To roll forward to a
|
||
bigger host:
|
||
</p>
|
||
<ol>
|
||
<li>Stop AIO on homework03 (preserve data on desslok via NFS).</li>
|
||
<li>Provision a bigger host (recommend: 32 GB RAM, NVMe).</li>
|
||
<li>Mount the same NFS export at the same path.</li>
|
||
<li>Copy <code>/usr/local/containers/nextcloudaio/</code> over (or
|
||
rebuild from the saved <code>aio-config.tar.gz</code>).</li>
|
||
<li>Update Caddy upstream IP on hawker.</li>
|
||
<li>Run a manual backup immediately to confirm the new host can
|
||
write to the same NFS.</li>
|
||
</ol>
|
||
|
||
<h2>Append-only references</h2>
|
||
|
||
<ul>
|
||
<li><a href="https://github.com/nextcloud/all-in-one">Nextcloud AIO docs</a> — official compose + variable reference</li>
|
||
<li><a href="https://docs.nextcloud.com/server/latest/admin_manual/">Nextcloud admin manual</a> — OCC, app installation, user mgmt</li>
|
||
<li><a href="https://www.collaboraoffice.com/code/">Collabora CODE</a> — WOPI integration details</li>
|
||
<li><code>docs/skill/nextcloud-aio-deploy</code> (Hermes skill) — abbreviated deploy workflow</li>
|
||
</ul>
|
||
|
||
</div>
|
||
</body>
|
||
</html> |