| 1 | # Service connection checks |
| 2 | |
| 3 | On 2026-09-26, the VM's pgAdmin database contained one configured server, `Clover Postgres`, at the current Nomad address and port. A query from the pgAdmin container using its rendered environment authenticated to the `postgres` database. The `pg.studio.test` health route passed. Zenith's pgAdmin also had one saved server; Snow Globe creates its connection from the Postgres requirement on boot. |
| 4 | |
| 5 | Redis Insight's `/api/databases/0/info` initially connected to the VM Redis service. After stopping Redis allocation `7ca6e34a`, Nomad placed a new one and changed its port from `26848` to `26913`. Redis Insight's watched `nomadService` template restarted its task once; its saved connection then used `26913` and `/api/databases/0/info` connected again. A temporary Redis key survived the reallocation and was deleted afterward. Both `redis.studio.test` and `pg.studio.test` passed their HTTP checks. [Nomad's template `change_mode` defaults to `restart`](https://developer.hashicorp.com/nomad/docs/job-specification/template), and [Redis Insight applies changed preconfigured connections after a restart](https://redis.io/docs/latest/operate/redisinsight/configuration/). |
| 6 | |
| 7 | After the x86 VM crashed and rebooted on 2026-09-26, Keycloak needed several minutes to start under emulation. Shale's earlier `wait-keycloak-app` prestart task remained marked complete, so Shale restarted 19 times while its OIDC discovery endpoint returned 404. Forward Auth also retried. Both recovered when Keycloak became healthy. [Nomad does not rerun a successful non-sidecar prestart task after a task restart](https://developer.hashicorp.com/nomad/docs/job-specification/lifecycle). |
| 8 | |
| 9 | Evil.inc Forgejo and HedgeDoc remained running but unhealthy after the same reboot. Forgejo had tried PostgreSQL before it was ready and then stopped making progress. Targeted `nomad job restart -all-tasks` calls recovered both; the full 19-route Snow Globe check then passed. Snow Globe now renders a one-minute retry delay for dependent services and [Nomad `check_restart`](https://developer.hashicorp.com/nomad/docs/job-specification/check_restart) for unhealthy services. Disposable one- and two-task Nomad jobs proved that a failing group-level check restarts every running task in its group. Eight generated jobs, including Keycloak, Forgejo, and HedgeDoc, passed `nomad job validate`. |
| 10 | |
| 11 | Release `7ec7b25654702110` staged Shale from its ZFS clone and promoted to the VM after backup `20260926T134201Z-22d930` captured 22 service datasets. The promotion passed all 19 HTTPS routes. Live Nomad inspection confirmed Shale's one-minute retry delay and five-minute health-check grace, plus Forgejo's 20-minute grace from its declared healthy deadline. |
| 12 | |
| 13 | After a clean VM reboot, PostgreSQL passed first, Keycloak became healthy at 14:04 UTC, and Shale recovered on its next one-minute retry. Forgejo remained running but unhealthy after its early database connection failed. At 14:15 UTC, Nomad's group-level `check_restart` restarted both Forgejo and Anubis without intervention; both checks and Forgejo's HTTPS route passed by 14:17 UTC. The full 19-route check passed after reboot. From the Mac, Shale, Keycloak discovery, and the Shale preview returned 200 through `.studio.test` DNS and trusted TLS. The 20-minute grace made Forgejo's recovery too slow, so its service now declares a two-minute health-restart grace, separate from its deployment deadline. |
| 14 | |
| 15 | Keycloak preview `keycloak-preview-dbe9aba6` ran with the production `start` command against a cloned database. Its OIDC discovery advertised the preview hostname. Release `47342c5c49f3eb88` promoted that preview after backup `20260926T141745Z-93733a` captured 22 service datasets. The production Keycloak allocation passed its health check with zero restarts; the promotion completed its dependency allocation and all 19 HTTPS checks. From the Mac, discovery returned HTTP 200 with a trusted certificate and issuer `https://keycloak.studio.test/realms/master`. Live job inspection confirmed `start`, the full production hostname, and Forgejo's two-minute check-restart grace. |
| 16 | |
| 17 | Shale and Jellyfin currently use writable SQLite datasets, so their single-allocation `simple` rollouts cannot overlap old and new writers. The owner accepted a brief Shale restart after staged validation. A controlled restart of the Shale preview on the x86 VM returned 46 failed HTTPS probes from this Mac between 4.87 and 23.14 seconds into a 0.25-second sampling run. Nomad reported the task restart complete after two seconds, before the Caddy route served requests again. The preview subsequently returned HTTP 200 with trusted TLS and one running allocation. This measures a restart under x86 emulation, not a production release rollout or physical-host downtime. |
| 18 | |
| 19 | Open Speed Test has no writable mounts or fixed ports, so its job now uses Nomad's overlapping canary. The first VM promotion automatically promoted a healthy canary, but one of 200 Mac HTTPS probes received HTTP 502 when Caddy dialed the retiring allocation's closed port before a route reload. The router now gives generated reverse proxies a five-second retry window and remembers failed upstream connections for 30 seconds, using [Caddy's documented load balancing options](https://caddyserver.com/docs/caddyfile/directives/reverse_proxy). The next staged release validated the generated Caddyfile and automatically promoted another healthy canary; 510 consecutive Mac HTTPS probes from 03:19:41 to 03:22:11 UTC returned 200, spanning its 03:21:05–03:21:20 rollout. This is a sampled VM result, not a guarantee that every request in future rollouts will succeed. |
| 20 | |
| 21 | For a new clone stage with a PostgreSQL database requirement, Snow Globe now snapshots the service dataset and PostgreSQL dataset in one ZFS operation, which [OpenZFS creates atomically](https://openzfs.github.io/openzfs-docs/man/v2.1/8/zfs-snapshot.8.html). It starts an isolated PostgreSQL container from the latter snapshot and copies the database from that container into a stage database on the live PostgreSQL server. The 2026-09-27 HedgeDoc preview `evil-hedgedoc-preview-31768845` completed its Nomad deployment and returned HTTP 200 from the Mac. Its service snapshot remains for the preview; the temporary PostgreSQL snapshot, clone, and container were removed. This aligns the file and database fork point for these services, while applications remain responsible for their own crash recovery from the snapshots. |
| 22 | |
| 23 | Writable external ZFS mounts now join that same snapshot operation. The recreated YouTube Feed preview `yt-feed-preview-63d74b72` mounted cloned service data and a cloned writable Clover configuration folder, then completed its Nomad deployment. Both source snapshots report `createtxg=35540`, and the preview's VM health check passed. The Mac route redirected to authentication as configured. |
| 24 | |
| 25 | The preview's web container wrote a disposable file through `/yt-config` as its assigned `3118:3000` identity. It appeared in the cloned Clover configuration folder with that owner and was absent from `/srv/clover/Documents/Config/Youtube Downloader`; the file was removed afterward. This verifies writable external mount isolation and group access through the container, beyond snapshot metadata alone. |
| 26 | |
| 27 | An encrypted ZFS rehearsal on the VM created separate source and staging encryption roots, then cloned a source snapshot beneath the staging root. The clone mounted, retained the source encryption root and origin, initially used `0B`, and accepted a write without changing the source file. The temporary datasets and keys were destroyed afterward. Production staging therefore needs the source dataset key loaded while its clone exists; placing the clone below a separately encrypted staging parent does not rekey it. |
| 28 | |
| 29 | Promotion of `c1f546fe904f4c33` stopped during Keycloak configuration. The old task digest included every file in each service folder, so a change to PostgreSQL's `provide.py` restarted PostgreSQL despite no database job configuration change. Keycloak's watched database endpoint then restarted its existing task; its group-level health check restarted that task again before x86 startup completed. A fresh Keycloak allocation completed startup, and all 19 production routes passed. Snow Globe now hashes source files into a task environment variable only for services with a `prepare` script, since those scripts change startup files under managed volumes. Editing PostgreSQL's provider script left its rendered job identical in a temporary release copy; editing qBittorrent's prepare script changed its rendered job. Single-container Nomad services now associate their checks with the task, and setup waits for its route before calling the app API. All 26 generated jobs validated. The new Keycloak preview reached HTTPS health and completed setup with zero restarts; after a manual in-place restart, it had no health-triggered restart through its nine-minute startup. |
| 30 | |
| 31 | Release `c19bee75fad72499` promoted successfully after backup `20260927T042521Z-9fcd75`; all 19 routes passed. Keycloak's production allocation reached health after its slow x86 build and startup with zero restarts. The Mac reached Keycloak and Shale with HTTP 200 and Jellyfin with its expected login redirect. During pgAdmin startup, QEMU paused with a host disk I/O error. Removing the two detached NixOS installer ISOs and moving the unused arm64 demo disk to `/Volumes/Project/Studio VM Archives/clover-studio-vm-arm64.qcow2` freed host space; QEMU resumed and the same promotion process finished. The resulting release history and `/opt/studio/current` both point at `c19bee75fad72499`. |
| 32 | |
| 33 | Read-only `nomad job plan` checks against the active release returned no allocation creates or destroys for PostgreSQL, Keycloak, or Shale. Their rendered HCL files were byte-identical to `nomad job inspect -hcl`. Nomad still described each plan as one in-place update without a field-level diff, so that phrase alone is not evidence of a changed job or a task restart. [The Nomad CLI defines exit code 0 as no allocation creation or destruction](https://developer.hashicorp.com/nomad/commands/job/plan). Resubmitting the identical PostgreSQL job retained its allocation `ff3fc333` and version 14; Keycloak retained allocation `d1c5e27a`. |
| 34 | |
| 35 | Release `9ee7ba520b01b1e5` staged YouTube Triage with its writable external clone, passed the VM NixOS dry build, and promoted after backup `20260927T053411Z-183aca`. The NixOS switch raised the demo pool size to 16 GB; all 19 routes passed. PostgreSQL, Keycloak, Jellyfin, Shale, and Dawarich kept allocation IDs `ff3fc333`, `d1c5e27a`, `c2b6c232`, `25050918`, and `abad9408` through this promotion. The Mac received HTTP 302 at the protected preview URL and HTTP 200 from Shale. |
| 36 | |
| 37 | The manual CLI rolled back from `9ee7ba520b01b1e5` to `c19bee75fad72499` after backup `20260927T054434Z-67d6c0`, switched the NixOS host configuration, and passed all 19 routes. Promoting the same preview again backed up the rollback state as `20260927T055034Z-27d895`, restored NixOS release `9ee7ba520b01b1e5`, and passed the same route sweep. Those five sampled allocations kept their IDs through both switches; the Mac still received Shale HTTP 200 and the protected preview's HTTP 302. Data restore was not invoked by code rollback. |