| 1 | # Production data rollback |
| 2 | |
| 3 | Every promotion and code rollback backs up the active release before switching it. One ZFS snapshot operation captures the mounted service datasets and PostgreSQL data at the same instant. A temporary clone of that PostgreSQL snapshot runs without a network and produces the service-owned database dumps, including services with multiple database inputs. The clone is destroyed before the backup is published. `python3 tools/deploy.py backups` lists the backup IDs and services. A failed promotion can leave a completed backup even when release history has no new entry. |
| 4 | |
| 5 | The backup now also saves a private `nomad.snap` and its SHA-256 in the manifest. A disposable backup run against the VM's live Nomad server verified the file and digest, then removed it. This snapshot includes jobs, ACLs, and secret variables; keep the backup directory private. Service data restore does not apply the whole Nomad snapshot, which is reserved for an explicit cluster recovery. The backup directory is still on the OS disk until Snow Globe control state has durable storage. |
| 6 | |
| 7 | At inspection, the VM retained 25 release backups: their PostgreSQL dumps and manifests used `275 MB`, and their service snapshots numbered 550. No automatic retention policy deletes these backups yet; choose one before long-running production use. Earlier backup manifests do not contain `nomad.snap`. |
| 8 | |
| 9 | To restore one service, roll its code back first, then run: |
| 10 | |
| 11 | ```sh |
| 12 | python3 tools/deploy.py backups |
| 13 | python3 tools/deploy.py rollback <old-release-id> |
| 14 | python3 tools/deploy.py data-restore <service-id> --backup <backup-id> --discard-writes |
| 15 | ``` |
| 16 | |
| 17 | The backup's `fromRelease` must match the active release. Data restore stops that service, saves its current dataset and database as a safety copy, restores the chosen ZFS snapshot and PostgreSQL dump, then deploys and checks the service. **Changes written after the chosen backup are discarded from the live service.** The safety snapshot and dump remain on the host. Postgres itself is shared; restore its owner services individually. Shared Clover and Media mounts are outside this service restore; restoring an entire shared dataset would discard unrelated users' writes. |
| 18 | |
| 19 | The CLI checks that the backup snapshot and database dumps exist before stopping the service. A disposable missing-snapshot manifest on the VM was rejected while Redis Insight stayed running; the fixture was removed. |
| 20 | |
| 21 | On the VM, backup `20260927T030206Z-67f38b` captured 22 service snapshots and four database dumps from a cloned PostgreSQL snapshot. All dump hashes and snapshots verified; the temporary container and dataset were removed while the production PostgreSQL allocation stayed running. An injected dump failure after clone startup left the ZFS dataset/snapshot list and backup directory list unchanged, with no probe container. Earlier rollback and restore tests are recorded in the release history. |
| 22 | |
| 23 | # VM restore rehearsal |
| 24 | |
| 25 | On 2026-09-27, backup `20260927T033539Z-123cab` captured the running x86 VM's service datasets and PostgreSQL databases. A throwaway file and table were then added to `evil-hedgedoc`. Running `data.py restore 20260927T033539Z-123cab evil-hedgedoc --discard-writes` stopped the job, saved the pre-restore safety snapshot, restored the dataset and database, and redeployed it. Both probes disappeared, `https://md.evil.studio.test/_health` returned 200 from the Mac, and all 19 VM routes passed their health checks. This proves the explicit data restore path on the VM; the production NAS cutover remains separate. |
| 26 | |
| 27 | The `20260927T055034Z-27d895` backup retained all 22 service snapshots and four matching PostgreSQL dumps. Its Dawarich dump restored into a new database owned by `svc_dawarich` in an isolated PostgreSQL ZFS clone, with `pgcrypto` and `postgis` created first and `pg_restore --role=svc_dawarich`. The restored `points` table contained 22,666 rows. The probe container and clone were removed, the live PostgreSQL allocation stayed running, and Dawarich still returned HTTP 200. |