Files
office/operations.html
T
race d172771aa1 Initial docs: Nextcloud AIO office suite (Collabora + Whiteboard)
- Full deployment reference for office.rmf44.xyz
- Architecture, procedure, troubleshooting, operations pages
- 4 SVG diagrams (topology, container-tree, data-flow, request-flow)
- Mirrors gite_replacement template structure
- Verified via 70/70 ad-hoc checks on 2026-08-10
2026-08-10 15:43:46 -05:00

258 lines
9.5 KiB
HTML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Operations — Nextcloud Office</title>
<link rel="stylesheet" href="assets/style.css">
</head>
<body>
<div id="wrapper">
<div class="crumbs"><a href="index.html">Nextcloud Office</a> &nbsp;›&nbsp; Operations</div>
<ul class="nav">
<li><a href="index.html">Overview</a></li>
<li><a href="architecture.html">Architecture</a></li>
<li><a href="procedure.html">Procedure</a></li>
<li><a href="troubleshooting.html">Troubleshooting</a></li>
<li><a href="operations.html" class="active">Operations</a></li>
</ul>
<h1>Operations</h1>
<p>
Day-2 ops: backups, monitoring, recovery procedures, and the
roll-forward / roll-back plans.
</p>
<h2>Backup pipeline</h2>
<p>
Three files written daily to
<code>/srv/nc-files/backups/</code> on homework03 (NFS, real path
<code>/slab/container_storage/office/backups/</code> on desslok):
</p>
<table>
<tr><th>File</th><th>Contents</th><th>Typical size</th><th>Recovery use</th></tr>
<tr>
<td><code>office-YYYYMMDD-pgdump.sql.gz</code></td>
<td>PostgreSQL full dump via <code>pg_dumpall</code> from the AIO database container. All ~155 Nextcloud tables.</td>
<td>~600 KB (empty) → grows with users/files</td>
<td>Restore the database after a Nextcloud corruption or migration to new hardware.</td>
</tr>
<tr>
<td><code>office-YYYYMMDD-aio-config.tar.gz</code></td>
<td>The mastercontainer's <code>configuration.json</code> (office suite choice, domain, datadir, passwords) + database-dump bind target.</td>
<td>~6 KB</td>
<td>Reconstruct the AIO install state without going through the setup wizard again.</td>
</tr>
<tr>
<td><code>office-YYYYMMDD-ncdata.tar.gz</code></td>
<td>Tar of <code>/srv/nc-files/nextcloud/</code> (user-uploaded files) — excludes <code>backups/</code> to avoid recursion.</td>
<td>Empty (~100 B) until users upload files, then grows</td>
<td>Restore user files after data loss.</td>
</tr>
</table>
<h3>Daily cron schedule</h3>
<p>
Triggered by a systemd timer on <code>hector</code>, daily at
03:30 UTC (with up to 15 min random delay). The unit SSHes into
homework03 (no password prompt — keys only) and runs the script
with <code>sudo</code>.
</p>
<pre><code>ssh tigo@hector
systemctl list-timers office-backup*
# Expect: NEXT shown for the next 03:30 UTC ± 15 min
# Manual trigger for testing
sudo -n systemctl start office-backup.service
sleep 30
systemctl status office-backup.service | head -5
# Expect: Active: inactive (dead) → success</code></pre>
<h3>Retention policy</h3>
<p>
14 days. The script prunes via <code>find ... -mtime +14 -delete</code>
after the daily write. Same-day reruns overwrite (date-only stamp)
— intentional; we don't want to keep multiple copies per day.
</p>
<h3>What this doesn't cover</h3>
<ul>
<li>
<strong>Container runtime state</strong> — AIO's named volumes
on local ext4 are NOT backed up by this pipeline. If homework03
loses its disk, the AIO setup wizard will rebuild containers
from the saved <code>configuration.json</code> + the NFS data,
but you'll lose any state stored in those volumes (e.g. the
mastercontainer's domain-validation certificates cache). In
practice these regenerate on first boot.
</li>
<li>
<strong>NFS quiescence</strong> — the tar reads
<code>/srv/nc-files/nextcloud/</code> while the filesystem is
actively being written to by the nextcloud container. The tar
will see a consistent enough snapshot for crash-consistent
recovery; for true point-in-time recovery, you'd want to
quiesce Nextcloud (set maintenance mode) for the duration of
the tar, which we haven't done.
</li>
</ul>
<h2>Monitoring &amp; alerting</h2>
<h3>Gatus endpoints to watch</h3>
<p>
Gatus runs on <code>monitor (10.0.0.75)</code>, port 10010.
Suggested checks for the Nextcloud stack:
</p>
<table>
<tr><th>Endpoint</th><th>What</th><th>Severity</th></tr>
<tr><td><code>https://office.rmf44.xyz/login</code></td><td>Public ingress (Caddy → Apache → PHP-FPM)</td><td>P1 outage</td></tr>
<tr><td><code>https://100.79.142.164:8443</code></td><td>AIO admin UI (mastercontainer direct)</td><td>P2 if down</td></tr>
<tr><td>docker stats — <code>nextcloud-aio-*</code></td><td>Container health</td><td>P2 if any restart loop</td></tr>
<tr><td>NFS — <code>/srv/nc-files</code> on homework03</td><td>Mount up + writable</td><td>P1 (data loss risk)</td></tr>
</table>
<p>
The backup pipeline's last-run status is readable via
<code>systemctl status office-backup.service</code> on hector;
adding a Gatus check on this is straightforward via SSH exec.
</p>
<h2>Recovery procedures</h2>
<h3>Restore from a daily backup</h3>
<p>
Full restore assumes a clean homework03 + intact NFS on desslok.
</p>
<ol>
<li>
Stop the AIO stack:
<pre><code>ssh homework03
cd /usr/local/containers/nextcloudaio
sudo -n docker compose down</code></pre>
</li>
<li>
Restore the AIO config (replaces configuration.json):
<pre><code>LATEST=$(ls -t /srv/nc-files/backups/office-*-aio-config.tar.gz | head -1)
tar -C /usr/local/containers/nextcloudaio -xzf "$LATEST"</code></pre>
</li>
<li>
Restore user files (overwrites the NFS share's <code>nextcloud/</code>):
<pre><code>LATEST=$(ls -t /srv/nc-files/backups/office-*-ncdata.tar.gz | head -1)
# Tar contains /nextcloud/ at root
tar -C /srv/nc-files -xzf "$LATEST"</code></pre>
</li>
<li>
Restore the database (drop + reload):
<pre><code># Start only the database container first
sudo -n docker compose up -d nextcloud-aio-mastercontainer
sleep 30
# Wait for the database container to come up via mastercontainer
sudo -n docker exec nextcloud-aio-database pg_isready -U nextcloud
LATEST=$(ls -t /srv/nc-files/backups/office-*-pgdump.sql.gz | head -1)
zcat "$LATEST" | sudo -n docker exec -i nextcloud-aio-database psql -U nextcloud -d nextcloud_database</code></pre>
</li>
<li>
Restart the AIO stack:
<pre><code>sudo -n docker compose restart
sleep 60
curl -skI https://office.rmf44.xyz/login
# Expect: HTTP/2 200</code></pre>
</li>
</ol>
<h3>Restore a single file</h3>
<p>
No need for a full restore — just untar one file:
</p>
<pre><code>ssh desslok
LATEST=$(ls -t /slab/container_storage/office/backups/office-*-ncdata.tar.gz | head -1)
tar -C / -xzf "$LATEST" nextcloud/admin/files/path/to/file
# Adjust for the user + path</code></pre>
<h3>Re-initialize the admin user</h3>
<p>
If the admin password is lost:
</p>
<pre><code>ssh homework03
# Reset via OCC
sudo -n docker exec -u www-data nextcloud-aio-nextcloud \
php /var/www/html/occ user:resetpassword admin --password-from-env
# Reads password from NEXTCLOUD_ADMIN_PASSWORD env var
# (default: same as setup wizard)</code></pre>
<h2>Updates &amp; upgrades</h2>
<p>
AIO manages its own updates: when a new <code>all-in-one</code>
image is published, mastercontainer pulls the new image and
triggers a rolling update of all side containers.
</p>
<p>
To manually trigger an update:
</p>
<pre><code>ssh homework03
cd /usr/local/containers/nextcloudaio
sudo -n docker compose pull
sudo -n docker compose up -d
# Wait 5-10 min for all side containers to roll</code></pre>
<p>
<strong>Before a major update</strong>, take a manual backup:
<code>sudo -n systemctl start office-backup.service</code> on
hector, then verify the files exist on desslok before pulling
new images.
</p>
<h2>Rollback (revert to OnlyOffice)</h2>
<p>
The OnlyOffice container was retired on 2026-08-10. To bring it
back, you'd need the saved tarball at
<code>/home/tigo/onlyoffice-stack-backup-20260810.tar.gz</code>
on hawker. Rollback time estimate: ~30 minutes (restore compose,
start containers, restore Caddy vhost, smoke test).
</p>
<p>
<strong>Recommendation:</strong> keep that tarball for at least
one more month, then archive to cold storage. If the new AIO
stack proves stable, drop the tarball after that.
</p>
<h2>Roll-forward (move to dedicated AIO host)</h2>
<p>
The current 15 GB homework03 is tight on RAM. If we add Talk or
Fulltextsearch later, the host won't fit. To roll forward to a
bigger host:
</p>
<ol>
<li>Stop AIO on homework03 (preserve data on desslok via NFS).</li>
<li>Provision a bigger host (recommend: 32 GB RAM, NVMe).</li>
<li>Mount the same NFS export at the same path.</li>
<li>Copy <code>/usr/local/containers/nextcloudaio/</code> over (or
rebuild from the saved <code>aio-config.tar.gz</code>).</li>
<li>Update Caddy upstream IP on hawker.</li>
<li>Run a manual backup immediately to confirm the new host can
write to the same NFS.</li>
</ol>
<h2>Append-only references</h2>
<ul>
<li><a href="https://github.com/nextcloud/all-in-one">Nextcloud AIO docs</a> — official compose + variable reference</li>
<li><a href="https://docs.nextcloud.com/server/latest/admin_manual/">Nextcloud admin manual</a> — OCC, app installation, user mgmt</li>
<li><a href="https://www.collaboraoffice.com/code/">Collabora CODE</a> — WOPI integration details</li>
<li><code>docs/skill/nextcloud-aio-deploy</code> (Hermes skill) — abbreviated deploy workflow</li>
</ul>
</div>
</body>
</html>