rewire 2026-08-11: single-mount NFS architecture

- topology.svg, container-tree.svg, data-flow.svg: show /srv/nc-files
  bound directly into nextcloud container at /mnt/ncdata (no
  intermediate /mnt/nc-data/nextcloud-data layer).
- architecture.html: explain rewire, update storage table, remove
  pre-rewire 'why two layers' section.
- procedure.html: NEXTCLOUD_DATADIR=/srv/nc-files, no mastercontainer
  bind, add §2.5 Rewire 2026-08-11 with recovery procedure,
  update backup script to also tar local nextcloud app volume,
  fix 4.3 verification (302 -> /login, status.php, file visibility).
- operations.html: 4 backup files now (added aio-nextcloud-app.tar.gz),
  restore procedures updated, single-file restore uses new tar layout.
- troubleshooting.html: add 3 new sections — appdata-missing,
  nextcloud-not-spawning, collabora-discovery-warning.
- request-flow.svg: remove nextcloud/ subdir from NFS write path.
- index.html, README.md: update paths and metadata.
This commit is contained in:
2026-08-11 12:09:26 -05:00
parent d172771aa1
commit b4777a5e4d
10 changed files with 423 additions and 121 deletions
+168 -3
View File
@@ -21,9 +21,10 @@
<h1>Troubleshooting</h1>
<p>
Every pitfall hit during the 2026-08-10 deployment, with root cause
and resolution. Order is roughly chronological — these are what
blocked progress at each stage.
Every pitfall hit during the 2026-08-10 deployment plus the
2026-08-11 rewire, with root cause and resolution. Order is
roughly chronological — these are what blocked progress at each
stage.
</p>
<div class="toc">
@@ -38,6 +39,9 @@
<li><a href="#onlyoffice-rejected">"OnlyOffice" rejected by AIO (must use Collabora or office flag)</a></li>
<li><a href="#nextcloud-login-flow">curl login returns 303 with empty user — CSRF cookie dance</a></li>
<li><a href="#backup-script-ownership">Backup script won't run as tigo — root-owned 755 instead</a></li>
<li><a href="#appdata-missing">"Appdata directory is not present" — wrong datadir (rewire 2026-08-11)</a></li>
<li><a href="#nextcloud-not-spawning">nextcloud container not appearing after config change</a></li>
<li><a href="#collabora-discovery-warning">Collabora logs "Could not create path .../richdocuments/remoteData/discovery"</a></li>
</ul>
</div>
@@ -351,6 +355,167 @@ sudo -n bash -n /usr/local/bin/office-backup.sh # syntax check</code></pre>
ExecStart=/usr/bin/ssh -o BatchMode=yes -o ConnectTimeout=30 \
homework03 sudo /usr/local/bin/office-backup.sh</code></pre>
<h2 id="appdata-missing">"Appdata directory is not present"</h2>
<div class="callout danger">
<p><strong>Symptom:</strong> <code>docker logs nextcloud-aio-nextcloud</code>
shows <code>Cannot write into directory "/mnt/ncdata/appdata_*"</code>
or <code>Appdata directory is not present!</code>. The login page
may load but logins loop or fail. Sometimes the nextcloud
container restart-loops.</p>
</div>
<h3>Root cause</h3>
<p>
AIO uses the value of <code>NEXTCLOUD_DATADIR</code> (or the
<code>nextcloud_datadir</code> field in
<code>configuration.json</code>) as the <strong>host bind-mount
source</strong> that goes into the nextcloud container's
<code>/mnt/ncdata</code>. If that path is not the actual NFS
mount (or is a different host path than the NFS mount), AIO
happily bind-mounts whatever the path resolves to — including an
empty local directory — and Nextcloud boots against an empty
datadir that has no <code>appdata_*</code> directory in it.
</p>
<p>
The data is still on NFS. It's just not being mounted. This
took office.rmf44.xyz offline on 2026-08-11 for ~15 minutes.
</p>
<h3>Fix (the 2026-08-11 rewire)</h3>
<ol>
<li>
Pick <strong>one</strong> host path that is the NFS mount.
We picked <code>/srv/nc-files</code>.
</li>
<li>
Set both <code>NEXTCLOUD_DATADIR</code> in compose AND
<code>nextcloud_datadir</code> in
<code>configuration.json</code> to that same path.
</li>
<li>
Remove any intermediate bind in the mastercontainer's
compose <code>volumes:</code> section — AIO doesn't need it
and it adds a layer that can drift out of sync.
</li>
<li>
Force the nextcloud container to respawn with the new bind:
<code>docker rm -f nextcloud-aio-nextcloud</code>, then
trigger <code>/api/docker/start</code> via the admin UI (with
Apache stopped first). See
<a href="#nextcloud-not-spawning">nextcloud container not appearing</a>
below for the spawn dance.
</li>
</ol>
<p>
The full sequence (mount, compose, config, respawn) is documented
in <a href="procedure.html#rewire-2026-08-11">procedure §2.5
Rewire 2026-08-11</a>.
</p>
<h2 id="nextcloud-not-spawning">nextcloud container not appearing after config change</h2>
<div class="callout warn">
<p><strong>Symptom:</strong> you updated
<code>configuration.json</code> (changed
<code>nextcloud_datadir</code>, enabled an extra container, etc.)
and restarted the mastercontainer, but the
<code>nextcloud-aio-nextcloud</code> container is missing or
running with the old config.</p>
</div>
<h3>Root cause</h3>
<p>
<strong>AIO does NOT auto-spawn the nextcloud container on
mastercontainer restart.</strong> The mastercontainer only
orchestrates the lifecycle of containers it spawns. If the
nextcloud container was already running, the mastercontainer
just observes it. If you <code>docker rm -f</code> an old
broken container and restart the mastercontainer, the
mastercontainer has no awareness that you want a new one — you
must explicitly trigger <code>/api/docker/start</code>.
</p>
<p>
Additionally, <code>isLoginAllowed()</code> in
<code>DockerActionManager.php</code> returns <code>false</code>
when Apache is starting or running and its port is open. So
<code>/api/docker/start</code> only fires when Apache is stopped.
</p>
<h3>Fix — the spawn dance</h3>
<p>
Use <code>/tmp/aio-flow.sh</code> on homework03 — it runs the
four steps atomically:
</p>
<pre><code>ssh homework03
sudo /tmp/aio-flow.sh
# 1. POST /api/docker/stop (stops Apache — login allowed)
# 2. GET /login + POST /api/auth/login (saves cookies + CSRF)
# 3. POST /api/docker/start (triggers spawn)
# 4. POST /api/docker/stop /start (re-runs for nextcloud container, restart Apache)</code></pre>
<p>
Or manually via curl:
</p>
<pre><code>JAR=/tmp/cookies.txt
BASE=https://office.rmf44.xyz:8080
# 1. Stop Apache
curl -skb $JAR -X POST "$BASE/api/docker/stop"
# 2. Login
curl -skc $JAR -o /tmp/login.html "$BASE/login"
TOKEN=$(grep -oE 'data-requesttoken="[^"]+"' /tmp/login.html | head -1 | sed 's/data-requesttoken="//;s/"$//')
curl -skb $JAR -c $JAR \
-H "requesttoken: $TOKEN" \
-d "password=$ADMIN_PASSWORD" \
-X POST "$BASE/api/auth/login"
# 3. Trigger spawn
curl -skb $JAR -c $JAR \
-H "requesttoken: $TOKEN" \
-X POST "$BASE/api/docker/start"
# 4. Restart Apache
curl -skb $JAR -X POST "$BASE/api/docker/start"</code></pre>
<h2 id="collabora-discovery-warning">Collabora logs "Could not create path .../richdocuments/remoteData/discovery"</h2>
<div class="callout info">
<p><strong>Symptom:</strong> after starting Collabora, the
nextcloud container logs lines like:</p>
<pre><code>Could not create path /mnt/ncdata/appdata_*/richdocuments/remoteData/discovery
Failed to fetch discovery endpoint</code></pre>
</div>
<h3>Root cause</h3>
<p>
The nextcloud container runs as <code>www-data</code>, but the
parent <code>appdata_*/</code> directory on NFS was created by an
earlier process (possibly the AIO entrypoint as root during first
init) and the per-app <code>richdocuments/</code> subdir doesn't
exist yet. The nextcloud container can't create it because it
doesn't own the parent.
</p>
<h3>Resolution</h3>
<p>
<strong>Wait it out.</strong> Collabora generates the discovery
JSON on first use (when a user opens a Word/Excel file in the
Collabora iframe). The warning is logged but the discovery is
cached on first successful WOPI round-trip. If Collabora is
actually broken (the iframe stays blank), check perms:
</p>
<pre><code>ssh homework03
sudo docker exec nextcloud-aio-nextcloud \
ls -la /mnt/ncdata/appdata_*/richdocuments/remoteData/
# If missing, force creation as www-data:
sudo docker exec -u www-data nextcloud-aio-nextcloud \
mkdir -p /mnt/ncdata/appdata_*/richdocuments/remoteData/</code></pre>
<p>
We have not seen this fail on a live Collabora open since the
2026-08-11 rewire — the warning appears once at startup and
then Collabora works normally.
</p>
</div>
</body>
</html>