forked from EduCraft/curriculum-project-hub
docs: resolve backup recovery audit
This commit is contained in:
+325
@@ -0,0 +1,325 @@
|
|||||||
|
# Backup, migration, and disaster-recovery audit
|
||||||
|
|
||||||
|
## Verdict
|
||||||
|
|
||||||
|
The current Hub has no production backup or restore control plane. PostgreSQL
|
||||||
|
can be dumped and restored manually, but the service's authoritative state is
|
||||||
|
split between PostgreSQL and mutable Project workspaces, with no shared
|
||||||
|
consistency point, backup catalog, off-host retention, restore workflow,
|
||||||
|
recovery objectives, or recovery drill.
|
||||||
|
|
||||||
|
The current deployment path can also delete customer workspaces when the
|
||||||
|
documented relative workspace root is used: `deploy_platform.sh` applies
|
||||||
|
`rsync --delete` to the live Hub directory while `.env.example` places
|
||||||
|
workspaces below `./data/project-workspaces`. Project creation has the same
|
||||||
|
cross-store weakness at a smaller scale: it creates the workspace before the
|
||||||
|
database transaction and leaves an orphan directory when that transaction
|
||||||
|
fails.
|
||||||
|
|
||||||
|
Migration safety is incomplete even though the current PostgreSQL migration
|
||||||
|
engine rolled back a deliberately failing migration's DDL in the executed
|
||||||
|
probe. The failed migration record blocks later migrations, migrations execute
|
||||||
|
inside every service start, no compatibility declaration or production-data
|
||||||
|
rehearsal exists, and current history contains destructive changes that make an
|
||||||
|
old application query fail after schema upgrade. Application rollback is
|
||||||
|
therefore not database rollback.
|
||||||
|
|
||||||
|
## Contract and persistent-state baseline
|
||||||
|
|
||||||
|
The audit preserves the accepted contracts rather than inventing missing
|
||||||
|
product policy:
|
||||||
|
|
||||||
|
- `Organization` is the tenant root, but `User` is global and can belong to
|
||||||
|
multiple Organizations. A selective Organization restore is not equivalent
|
||||||
|
to restoring a self-contained database schema.
|
||||||
|
- Project workspaces are authoritative curriculum-engineering files bounded by
|
||||||
|
ADR-0018. PostgreSQL stores their absolute `Project.workspaceDir` locations;
|
||||||
|
it does not store their file content.
|
||||||
|
- Run, provider cursor, ingress, delivery, audit, and workspace commit/recovery
|
||||||
|
semantics remain split across the unresolved lifecycle tickets. A backup
|
||||||
|
design cannot silently choose their crash behavior.
|
||||||
|
- ADR-0022 requires storage ceilings and emergency `DRAIN`/`STOP_NOW`, but does
|
||||||
|
not decide backup frequency, retention, recovery objectives, or tenant-level
|
||||||
|
restore scope.
|
||||||
|
- Audit retention is still `OPEN`; regulatory backup retention and deletion
|
||||||
|
obligations require product/legal input already recorded in the map's Fog.
|
||||||
|
|
||||||
|
The recovery set has at least four classes:
|
||||||
|
|
||||||
|
| State | Current location | Recovery significance |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| Tenant and workflow state | PostgreSQL | Organizations, identities, membership, permissions, Feishu bindings, sessions, runs, locks, receipts, usage facts, and audit records |
|
||||||
|
| Customer artifacts | `HUB_PROJECT_WORKSPACE_ROOT` | Project source files, generated artifacts, inbound attachments, and any workspace-local session/history files |
|
||||||
|
| Credentials and configuration | `/srv/.../.secrets/platform.env` today; future org-scoped encrypted connections plus external key material | Database/Feishu/provider access and ability to decrypt restored Organization credentials |
|
||||||
|
| Release identity | application release, Prisma migrations, `cph`, runtime/toolchain manifest | Required to start a schema-compatible application and reproduce outputs after restore |
|
||||||
|
|
||||||
|
Caches, build output, and reinstallable dependencies should be classified as
|
||||||
|
regenerable rather than silently mixed into the protected recovery set.
|
||||||
|
|
||||||
|
## Executed evidence
|
||||||
|
|
||||||
|
All probes used isolated temporary paths or databases. No production or shared
|
||||||
|
test data was modified.
|
||||||
|
|
||||||
|
### 1. The deployment rsync pattern deletes an in-tree workspace
|
||||||
|
|
||||||
|
A temporary live Hub directory was given the documented workspace shape and a
|
||||||
|
sentinel customer file, then the exact deployment rsync include/exclude pattern
|
||||||
|
was applied:
|
||||||
|
|
||||||
|
```text
|
||||||
|
rsync -a --delete \
|
||||||
|
--exclude node_modules --exclude dist --exclude .env \
|
||||||
|
hub/ <temporary-live-hub>/
|
||||||
|
|
||||||
|
workspace_survived=false
|
||||||
|
reproduced: deploy rsync pattern deletes an in-tree workspace root
|
||||||
|
```
|
||||||
|
|
||||||
|
This is deterministic: `data/` is absent from the release source and is not an
|
||||||
|
exclude, so `--delete` removes it from the destination. The root cause is that
|
||||||
|
the documented persistent path is nested inside the mutable release tree and
|
||||||
|
the deploy contract does not distinguish release artifacts from state.
|
||||||
|
|
||||||
|
### 2. A rejected Project transaction leaves an orphan workspace
|
||||||
|
|
||||||
|
A minimal Node harness called the real built `createProjectFromOrgAdmin` with a
|
||||||
|
valid Organization admin and a missing Folder. The fake Prisma boundary rejected
|
||||||
|
inside the same transaction seam used by production. The harness asserted that
|
||||||
|
the Organization workspace directory must remain empty and was run twice:
|
||||||
|
|
||||||
|
```text
|
||||||
|
node .scratch/saas-production-readiness/probes/orphan-workspace.mjs
|
||||||
|
|
||||||
|
orphan_workspace_entries=project_4aa0b6792c4547059d2d92b505b9299e
|
||||||
|
exit_code=1
|
||||||
|
|
||||||
|
orphan_workspace_entries=project_190e21afb8854244872b1241a600aa58
|
||||||
|
exit_code=1
|
||||||
|
```
|
||||||
|
|
||||||
|
`createManagedProject` calls `mkdir` before `prisma.$transaction`; the failure
|
||||||
|
path has no compensating removal, staged rename, or orphan reconciler. The
|
||||||
|
throwaway harness was removed after reproduction.
|
||||||
|
|
||||||
|
### 3. PostgreSQL rolled back failed migration DDL but retained a blocking failure
|
||||||
|
|
||||||
|
An isolated database and throwaway Prisma project applied one migration that
|
||||||
|
created a table, divided by zero, then would have created another table. The
|
||||||
|
real repository Prisma CLI returned `P3018`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
migration_exit=1
|
||||||
|
before_failure=missing
|
||||||
|
after_failure=missing
|
||||||
|
20260710000000_partial_failure|finished=false|rolled_back=false|steps=0|logs=true
|
||||||
|
```
|
||||||
|
|
||||||
|
This specific PostgreSQL/Prisma 6.19.3 path did **not** leave partial DDL, so the
|
||||||
|
audit does not claim that it did. It did leave an unfinished
|
||||||
|
`_prisma_migrations` row, and Prisma's error explicitly refused new migrations
|
||||||
|
until operator recovery. Because systemd runs `migrate deploy` as
|
||||||
|
`ExecStartPre`, the application cannot start past that state. The throwaway
|
||||||
|
database and migration were removed. This one multi-statement result is not a
|
||||||
|
general atomicity guarantee: Prisma's first-party guidance says PostgreSQL
|
||||||
|
migrations are not transaction-wrapped by default, so each migration must make
|
||||||
|
its transaction and partial-failure behavior explicit.
|
||||||
|
|
||||||
|
### 4. The current schema is not backward-compatible with an older application
|
||||||
|
|
||||||
|
The application immediately before the typed-principal migration queried
|
||||||
|
`PermissionGrant.principal`. That migration drops the column in favor of
|
||||||
|
`principalType` plus `principalId`. Running the old read against the current
|
||||||
|
test database produced:
|
||||||
|
|
||||||
|
```text
|
||||||
|
ERROR: column "principal" does not exist
|
||||||
|
HINT: Perhaps you meant to reference the column "PermissionGrant.principalId".
|
||||||
|
old_query_exit=1
|
||||||
|
```
|
||||||
|
|
||||||
|
This is direct evidence that switching the application symlink back after that
|
||||||
|
migration cannot restore service. The migration history also deletes duplicate
|
||||||
|
grant rows and later drops `ExternalPrincipalMembership.source`; there are no
|
||||||
|
downgrade migrations.
|
||||||
|
|
||||||
|
### 5. A database-only custom dump restores cleanly
|
||||||
|
|
||||||
|
As a positive baseline, `pg_dump --format=custom` captured the current
|
||||||
|
`cph_hub_test` database and `pg_restore --exit-on-error --no-owner` restored it
|
||||||
|
into a new isolated database:
|
||||||
|
|
||||||
|
```text
|
||||||
|
restored_tables=24
|
||||||
|
finished_migrations=9
|
||||||
|
failed_migrations=0
|
||||||
|
dump_bytes=76113
|
||||||
|
```
|
||||||
|
|
||||||
|
All nine finished migration checksums in the source database also matched the
|
||||||
|
checked-in migration files. This proves a manual logical database round trip on
|
||||||
|
the current PostgreSQL 14.20 workstation. It does not prove workspace,
|
||||||
|
credentials, global database roles, production volume, cross-version restore,
|
||||||
|
application readiness, or an operational RPO/RTO.
|
||||||
|
|
||||||
|
## Critical findings
|
||||||
|
|
||||||
|
### 1. No backup, retention, or restore mechanism exists
|
||||||
|
|
||||||
|
The repository has no backup/restore command, systemd timer, off-host target,
|
||||||
|
encryption/immutability policy, retention policy, backup manifest/checksum,
|
||||||
|
success/failure alert, restore command, or scheduled recovery drill. The
|
||||||
|
installer assumes PostgreSQL and its role/database already exist. A database
|
||||||
|
dump alone does not reproduce those global prerequisites or the filesystem and
|
||||||
|
secret state above.
|
||||||
|
|
||||||
|
There is also no decided RPO, RTO, backup frequency, retention duration,
|
||||||
|
geographic copy rule, deletion-in-backups rule, or ownership/escalation path.
|
||||||
|
Those values are product and operations decisions, not safe constants to infer
|
||||||
|
from the single-host topology.
|
||||||
|
|
||||||
|
### 2. Persistent workspaces are inside the documented release tree
|
||||||
|
|
||||||
|
`.env.example` recommends `./data/project-workspaces`; systemd's working
|
||||||
|
directory is the live Hub directory; deployment mutates that same directory
|
||||||
|
with `rsync --delete`. No preflight forbids an in-release workspace root and no
|
||||||
|
exclude preserves it. This can turn a routine deploy into irreversible customer
|
||||||
|
data loss before build, migration, restart, or readiness runs.
|
||||||
|
|
||||||
|
The persistent workspace root must live outside every release directory, have
|
||||||
|
explicit ownership/mode and capacity, and be impossible for release retention
|
||||||
|
or `rsync --delete` to traverse. Immutable release work already exists as a
|
||||||
|
separate ticket, but its acceptance test must include this sentinel case.
|
||||||
|
|
||||||
|
### 3. PostgreSQL and workspaces have no consistency boundary
|
||||||
|
|
||||||
|
PostgreSQL can provide a transactionally consistent database snapshot, but the
|
||||||
|
Hub mutates workspace files independently and has no generation, journal,
|
||||||
|
commit marker, snapshot ID, or shared checkpoint that binds filesystem content
|
||||||
|
to the corresponding database state. A database dump and filesystem copy taken
|
||||||
|
at different instants can disagree about Project existence, active runs,
|
||||||
|
provider cursors, delivered messages, and file effects.
|
||||||
|
|
||||||
|
Project creation already demonstrates the same missing compensation boundary.
|
||||||
|
A replacement host has an additional problem: `Project.workspaceDir` is an
|
||||||
|
absolute path and runtime paths from the database are used directly; restore
|
||||||
|
has no same-mount validation, safe rebasing operation, or missing/orphan scan.
|
||||||
|
|
||||||
|
For the initial single-host topology, a quiesced backup checkpoint is the
|
||||||
|
smallest credible route: stop new admission, drain/cancel according to the
|
||||||
|
accepted run contract, reach durable zero in-flight work, capture database and
|
||||||
|
workspace artifacts under one manifest, then resume. An online scheme would
|
||||||
|
need a newly designed journal/generation protocol; it must not be implied by
|
||||||
|
two independent copy commands.
|
||||||
|
|
||||||
|
### 4. Schema mutation is coupled to service start
|
||||||
|
|
||||||
|
Both `npm start` and the systemd unit execute `prisma migrate deploy`; every
|
||||||
|
ordinary restart can therefore become a schema rollout. The release path has no
|
||||||
|
separate migration lock/gate, `migrate status` check, production-data rehearsal,
|
||||||
|
duration/lock budget, pre-migration recovery checkpoint, or compatibility
|
||||||
|
manifest. CI only applies the full migration history to a fresh empty
|
||||||
|
PostgreSQL database, so it does not exercise real upgrade volume, legacy values,
|
||||||
|
old/new application overlap, failure recovery, or rollback compatibility.
|
||||||
|
|
||||||
|
The current failed-migration behavior is fail-closed, which is better than
|
||||||
|
starting on an unknown schema, but there is no operator procedure for inspecting
|
||||||
|
the failure, deciding roll-forward versus mark-rolled-back, restoring the
|
||||||
|
checkpoint, or preventing a restart loop. Existing destructive migrations prove
|
||||||
|
that application rollback cannot be assumed safe.
|
||||||
|
|
||||||
|
### 5. Restore has no isolated verification or cutover gate
|
||||||
|
|
||||||
|
There is no procedure that restores into a non-production environment with
|
||||||
|
outbound Feishu/provider traffic disabled, verifies backup checksums and
|
||||||
|
migration state, validates workspace paths/content/permissions, checks tenant
|
||||||
|
relationships, decrypts a sample org connection, reconciles non-terminal work,
|
||||||
|
runs application readiness and representative journeys, and measures recovery
|
||||||
|
time/point before controlled cutover.
|
||||||
|
|
||||||
|
Starting the current server is not a neutral verifier: startup deletes all
|
||||||
|
Project locks and independently marks only `ACTIVE` runs failed. Restored data
|
||||||
|
must first pass the durable lifecycle reconciliation designed by the run-state
|
||||||
|
tickets; otherwise validation itself mutates the evidence and can leave
|
||||||
|
`WAITING_FOR_USER` inconsistent.
|
||||||
|
|
||||||
|
### 6. Full-service and Organization-selective recovery are different products
|
||||||
|
|
||||||
|
A full-service restore can recreate the shared database and all workspaces.
|
||||||
|
Restoring one Organization is much harder: users are global and may belong to
|
||||||
|
multiple Organizations, permission/audit/history rows span many related tables,
|
||||||
|
workspace paths are external, and Feishu/provider connections have external
|
||||||
|
state. The repository has no export/import identity mapping or conflict policy.
|
||||||
|
|
||||||
|
The initial production gate must at least prove full-service disaster recovery.
|
||||||
|
Whether Organization-selective restore/export is promised, and how retained
|
||||||
|
backups interact with customer deletion, remains an explicit product/legal
|
||||||
|
decision rather than an accidental property of `pg_restore` filters.
|
||||||
|
|
||||||
|
## Required recovery contract
|
||||||
|
|
||||||
|
Production readiness requires the following guarantees, with exact numerical
|
||||||
|
objectives supplied by the pending policy decision:
|
||||||
|
|
||||||
|
1. **Declared objectives and ownership.** RPO/RTO, backup frequency, retention,
|
||||||
|
copy locations, encryption/key custody, immutability, deletion handling,
|
||||||
|
full-service versus Organization restore scope, operator ownership, and
|
||||||
|
escalation are documented and testable.
|
||||||
|
2. **Complete backup set.** One cataloged manifest identifies release and
|
||||||
|
migration checksums, PostgreSQL artifact and global prerequisites, workspace
|
||||||
|
snapshot, configuration inventory, encrypted org-secret data, external key
|
||||||
|
recovery material, timestamps, sizes, and cryptographic checksums. Backup
|
||||||
|
data and decryption material are not kept only on the protected host or in
|
||||||
|
the same failure domain.
|
||||||
|
3. **One consistency point.** A backup is taken only after an observable
|
||||||
|
quiesce/checkpoint protocol, or by a proven online generation protocol that
|
||||||
|
binds database state and workspace effects. Accepted input, active runs,
|
||||||
|
queues, deliveries, locks, and provider cursors have explicit recovery
|
||||||
|
semantics.
|
||||||
|
4. **Safe migration rollout.** Migration is a release phase, not an incidental
|
||||||
|
service start. The gate checks status/checksums, rehearses on representative
|
||||||
|
production-shaped data, records old/new application compatibility, uses
|
||||||
|
expand-and-contract for destructive changes, requires a recent verified
|
||||||
|
recovery checkpoint, and provides a tested failed-migration procedure.
|
||||||
|
5. **Isolated restore verification.** Every backup class is periodically
|
||||||
|
restored to a clean isolated host/database with outbound integrations off.
|
||||||
|
Verification checks integrity, schema/migration state, global database
|
||||||
|
prerequisites, workspace completeness/path ownership, secret decryption,
|
||||||
|
tenant isolation, lifecycle reconciliation, application readiness, and
|
||||||
|
representative admin/Feishu/agent journeys.
|
||||||
|
6. **Observable operation.** Backup age, duration, artifact/checksum status,
|
||||||
|
off-host replication, retention deletion, WAL/archive continuity if used,
|
||||||
|
restore-drill result, measured RPO/RTO, and failures are logged/metricized
|
||||||
|
and alerted with bounded cardinality.
|
||||||
|
7. **Controlled cutover and fallback.** Disaster recovery has an explicit
|
||||||
|
authority, traffic/workload brake, DNS/proxy/Feishu reconnect sequence,
|
||||||
|
duplicate/gap checks, rollback/fallback point, verification checklist, and
|
||||||
|
incident evidence record.
|
||||||
|
|
||||||
|
## Required release evidence
|
||||||
|
|
||||||
|
- A deploy test proves an in-tree or release-overlapping workspace root is
|
||||||
|
rejected before any rsync/switch, while persistent sentinel files survive
|
||||||
|
build failure, successful release, rollback, and release pruning.
|
||||||
|
- Project creation, archive/deletion, failed transaction, crash, and restore
|
||||||
|
tests prove database/workspace allocation has no silent missing or orphan
|
||||||
|
state; the reconciler reports rather than hides inconsistencies.
|
||||||
|
- Every checked-in migration is rehearsed from the previous production-shaped
|
||||||
|
snapshot, with old/new compatibility and lock/runtime evidence. A deliberately
|
||||||
|
failed migration exercises the documented recovery route.
|
||||||
|
- The backup implementation proves completeness, checksum verification,
|
||||||
|
encryption, off-host copy, retention, monitoring, and fail-closed partial
|
||||||
|
artifact cleanup.
|
||||||
|
- A clean isolated host restores a representative database, workspaces,
|
||||||
|
secrets, and compatible release; lifecycle reconciliation and critical tenant
|
||||||
|
journeys pass with outbound integrations controlled.
|
||||||
|
- Scheduled drills publish achieved recovery point and recovery time against the
|
||||||
|
accepted objectives, and alert when no recent successful drill exists.
|
||||||
|
|
||||||
|
## Primary-source interpretation
|
||||||
|
|
||||||
|
The official PostgreSQL, Prisma, and systemd behavior supporting this audit is
|
||||||
|
collected with claim-level links in
|
||||||
|
[Backup and migration primary sources](./backup-migration-primary-sources.md).
|
||||||
|
That note distinguishes what a PostgreSQL-consistent dump proves from the
|
||||||
|
cross-store guarantees the Hub must build itself.
|
||||||
@@ -0,0 +1,574 @@
|
|||||||
|
# Backup, migration, and disaster-recovery primary-source research
|
||||||
|
|
||||||
|
This note is the primary-source evidence pack for the Wayfinder ticket
|
||||||
|
“Audit backup, migration, and disaster recovery safety.” External claims use
|
||||||
|
only PostgreSQL, Prisma, systemd, and rsync first-party documentation. Project
|
||||||
|
recommendations are explicitly labelled as derived implications rather than
|
||||||
|
upstream guarantees.
|
||||||
|
|
||||||
|
PostgreSQL links below use `docs/current`, which was PostgreSQL 18 when this
|
||||||
|
research was performed on 2026-07-10. A production implementation must pin the
|
||||||
|
deployed PostgreSQL major and verify the same pages for that major. The Hub
|
||||||
|
currently depends on Prisma `^6.19.3`, so Prisma 6-specific command examples
|
||||||
|
remain relevant even where the live documentation also describes Prisma 7.
|
||||||
|
|
||||||
|
## Executive conclusions
|
||||||
|
|
||||||
|
1. `pg_dump` can take a consistent **single-database** snapshot while ordinary
|
||||||
|
reads and writes continue, but it does not include the Hub's workspace tree,
|
||||||
|
cluster-global roles/tablespaces, or a WAL chain. PostgreSQL now explicitly
|
||||||
|
says it is generally not the right regular production-backup mechanism
|
||||||
|
outside simple cases. It is a useful logical/export layer, not proof of
|
||||||
|
whole-service disaster recovery.
|
||||||
|
([`pg_dump` description](https://www.postgresql.org/docs/current/app-pgdump.html#PG-DUMP-DESCRIPTION),
|
||||||
|
[continuous archiving overview](https://www.postgresql.org/docs/current/continuous-archiving.html#CONTINUOUS-ARCHIVING))
|
||||||
|
2. Database PITR requires a physical base backup plus a continuous, unbroken
|
||||||
|
WAL sequence from the backup start through the recovery target. A successful
|
||||||
|
base-backup command alone does not prove that chain exists.
|
||||||
|
([continuous archiving overview](https://www.postgresql.org/docs/current/continuous-archiving.html#CONTINUOUS-ARCHIVING),
|
||||||
|
[making a base backup](https://www.postgresql.org/docs/current/continuous-archiving.html#BACKUP-BASE-BACKUP))
|
||||||
|
3. `pg_verifybackup` is necessary but not sufficient: PostgreSQL explicitly
|
||||||
|
requires test restores followed by checks that the restored database works
|
||||||
|
and contains the expected data.
|
||||||
|
([`pg_verifybackup`](https://www.postgresql.org/docs/current/app-pgverifybackup.html#APP-PGVERIFYBACKUP))
|
||||||
|
4. Prisma `migrate deploy` applies pending migrations, but it does not detect
|
||||||
|
schema drift and, on PostgreSQL, Prisma does not wrap a migration in a
|
||||||
|
transaction by default. A failed migration can therefore leave partial
|
||||||
|
schema changes requiring explicit operator recovery.
|
||||||
|
([Prisma production workflow](https://www.prisma.io/docs/orm/prisma-migrate/workflows/development-and-production#production-and-testing-environments),
|
||||||
|
[Prisma transactional behaviour](https://www.prisma.io/blog/prisma-migrate-dx-primitives#what-if-schema-migrations-were-atomic),
|
||||||
|
[failed-migration recovery](https://www.prisma.io/docs/orm/prisma-migrate/workflows/patching-and-hotfixing#failed-migration))
|
||||||
|
5. systemd supplies process lifecycle orchestration, not an application-level
|
||||||
|
consistency barrier. It sends the configured termination signal, waits up
|
||||||
|
to `TimeoutStopSec=`, then can forcibly kill remaining processes; a
|
||||||
|
successful `systemctl stop` does not prove that an `ExecStop=` command or
|
||||||
|
application drain succeeded.
|
||||||
|
([`TimeoutStopSec=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#TimeoutStopSec=),
|
||||||
|
[`systemctl stop`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#stop%20PATTERN%E2%80%A6))
|
||||||
|
6. The Hub persists one logical service state across two stores: PostgreSQL
|
||||||
|
records `Project.workspaceDir`, runs, messages, locks, and file-change audit
|
||||||
|
rows, while workspaces contain project files. A workspace transcript helper
|
||||||
|
exists but is not wired into the production runner. PostgreSQL's snapshot/
|
||||||
|
PITR guarantees cover only PostgreSQL. A whole-service recovery point
|
||||||
|
therefore requires an application-defined database/workspace checkpoint or
|
||||||
|
a journaled/versioned reconciliation protocol; neither PostgreSQL nor
|
||||||
|
systemd creates one automatically.
|
||||||
|
([Hub schema](../../../hub/prisma/schema.prisma),
|
||||||
|
[workspace transcript helper](../../../hub/src/agent/transcript.ts),
|
||||||
|
[`pg_dump` scope](https://www.postgresql.org/docs/current/app-pgdump.html#PG-DUMP-DESCRIPTION))
|
||||||
|
|
||||||
|
## Current repository facts that make the boundaries material
|
||||||
|
|
||||||
|
- The systemd unit runs `prisma migrate deploy` as `ExecStartPre=` before every
|
||||||
|
service start and uses `Restart=on-failure`.
|
||||||
|
([`cph-hub.service`](../../../hub/deploy/cph-hub.service))
|
||||||
|
- The deploy script updates the live Hub directory using
|
||||||
|
`rsync -az --delete`, builds in that same directory, restarts the unit, and
|
||||||
|
accepts a constant liveness endpoint as deployment success.
|
||||||
|
([`deploy_platform.sh`](../../../hub/deploy/deploy_platform.sh))
|
||||||
|
- `Project.workspaceDir` is database state, but the referenced directory holds
|
||||||
|
independently mutable tenant content. A workspace JSONL transcript helper
|
||||||
|
exists but is not wired into the production runner; file contents still live
|
||||||
|
in the workspace while the database separately carries run/message/file-change
|
||||||
|
facts.
|
||||||
|
([Hub schema](../../../hub/prisma/schema.prisma),
|
||||||
|
[workspace transcript helper](../../../hub/src/agent/transcript.ts))
|
||||||
|
|
||||||
|
These are project facts, not upstream guarantees. The sections below establish
|
||||||
|
what the upstream tools do and, critically, what they do not do.
|
||||||
|
|
||||||
|
## PostgreSQL logical dump and restore
|
||||||
|
|
||||||
|
### Snapshot and concurrent-write semantics
|
||||||
|
|
||||||
|
- `pg_dump` makes a consistent export even while the database is being used;
|
||||||
|
it does not block ordinary readers or writers. It dumps only one database.
|
||||||
|
Cluster-global objects such as roles and tablespaces require `pg_dumpall` or
|
||||||
|
`pg_dumpall --globals-only`.
|
||||||
|
([`pg_dump` description](https://www.postgresql.org/docs/current/app-pgdump.html#PG-DUMP-DESCRIPTION),
|
||||||
|
[whole-cluster logical dumps](https://www.postgresql.org/docs/current/backup-dump.html#BACKUP-DUMP-ALL))
|
||||||
|
- `pg_dumpall` invokes `pg_dump` once per database. Each database is internally
|
||||||
|
consistent, but the snapshots are not synchronized across databases.
|
||||||
|
([whole-cluster logical dumps](https://www.postgresql.org/docs/current/backup-dump.html#BACKUP-DUMP-ALL))
|
||||||
|
- A parallel dump uses directory format, opens `njobs + 1` connections, and
|
||||||
|
uses synchronized snapshots so every worker sees the same data set. The
|
||||||
|
connection count and additional server load are part of its capacity cost.
|
||||||
|
([`pg_dump --jobs`](https://www.postgresql.org/docs/current/app-pgdump.html#PG-DUMP-OPTIONS))
|
||||||
|
- Concurrent DDL can still interfere. The parallel-dump leader holds
|
||||||
|
`ACCESS SHARE` locks; if a later exclusive-lock request queues ahead of a
|
||||||
|
worker, that worker uses `NOWAIT` and aborts the dump rather than deadlocking.
|
||||||
|
`--lock-wait-timeout` can bound the initial lock wait.
|
||||||
|
([`pg_dump --jobs` and lock behaviour](https://www.postgresql.org/docs/current/app-pgdump.html#PG-DUMP-OPTIONS))
|
||||||
|
- The ordinary snapshot is suitable for disaster-recovery dump semantics, but
|
||||||
|
it need not correspond to a serial execution of all transactions that later
|
||||||
|
commit. `--serializable-deferrable` waits for an anomaly-free snapshot;
|
||||||
|
PostgreSQL says that option is not beneficial for a dump intended only for
|
||||||
|
disaster recovery, but it can matter for a continuously updated reporting
|
||||||
|
copy.
|
||||||
|
([`--serializable-deferrable`](https://www.postgresql.org/docs/current/app-pgdump.html#PG-DUMP-OPTIONS))
|
||||||
|
- The default waits for dump files to be written safely to disk. Production
|
||||||
|
backups must not use `--no-sync`, because an immediate OS crash can otherwise
|
||||||
|
leave a corrupt dump.
|
||||||
|
([`pg_dump --no-sync`](https://www.postgresql.org/docs/current/app-pgdump.html#PG-DUMP-OPTIONS))
|
||||||
|
|
||||||
|
Derived Hub implications:
|
||||||
|
|
||||||
|
- An online database dump may be internally valid while referring to a
|
||||||
|
`workspaceDir` whose files are from another instant. “`pg_dump` succeeded” is
|
||||||
|
therefore not a whole-Hub consistency claim.
|
||||||
|
- Backup scheduling must either serialize against schema migration/DDL or
|
||||||
|
treat any lock wait, abort, non-zero exit, and stderr warning as a failed
|
||||||
|
backup. A partially written archive must never be published as restorable.
|
||||||
|
- If the deployed cluster contains more than the Hub database, a single Hub
|
||||||
|
dump is not a cluster backup. Even `pg_dumpall` does not create a single
|
||||||
|
cross-database snapshot.
|
||||||
|
|
||||||
|
### Restore error and atomicity semantics
|
||||||
|
|
||||||
|
- `pg_restore` defaults to continuing after SQL errors and reports an error
|
||||||
|
count at the end. `--exit-on-error` is required for fail-fast behaviour.
|
||||||
|
`--single-transaction` wraps all emitted commands in `BEGIN`/`COMMIT`, makes
|
||||||
|
the restore all-or-nothing, and implies `--exit-on-error`.
|
||||||
|
([`pg_restore` error and transaction options](https://www.postgresql.org/docs/current/app-pgrestore.html#APP-PGRESTORE-OPTIONS))
|
||||||
|
- Parallel restore uses separate connections for data loading, index creation,
|
||||||
|
and constraint creation. It supports only custom/directory archives and
|
||||||
|
cannot be combined with `--single-transaction`; it is not a whole-restore
|
||||||
|
atomic operation.
|
||||||
|
([`pg_restore --jobs`](https://www.postgresql.org/docs/current/app-pgrestore.html#APP-PGRESTORE-OPTIONS))
|
||||||
|
- A single giant restore transaction can be impractical because it locks every
|
||||||
|
restored object and may exhaust the lock table. PostgreSQL 18 offers
|
||||||
|
`--transaction-size=N` as a bounded intermediate option; it implies
|
||||||
|
`--exit-on-error` but is not all-or-nothing.
|
||||||
|
([`pg_restore --transaction-size`](https://www.postgresql.org/docs/current/app-pgrestore.html#APP-PGRESTORE-OPTIONS))
|
||||||
|
- Plain SQL restored by `psql` also continues after errors unless
|
||||||
|
`ON_ERROR_STOP` is enabled; a single transaction is separately required for
|
||||||
|
all-or-nothing behaviour.
|
||||||
|
([restoring an SQL dump](https://www.postgresql.org/docs/current/backup-dump.html#BACKUP-DUMP-RESTORE))
|
||||||
|
- PostgreSQL recommends restoring into a truly empty database based on
|
||||||
|
`template0`; otherwise local additions in `template1` can cause duplicate
|
||||||
|
definition errors. Required roles must exist before owner/ACL restoration.
|
||||||
|
([restoring an SQL dump](https://www.postgresql.org/docs/current/backup-dump.html#BACKUP-DUMP-RESTORE))
|
||||||
|
- A dump can execute arbitrary code selected by superusers of the source
|
||||||
|
database. Partial dump/restore does not remove that trust boundary.
|
||||||
|
([`pg_dump` security warning](https://www.postgresql.org/docs/current/app-pgdump.html#PG-DUMP-DESCRIPTION))
|
||||||
|
- A `pg_dump` binary can dump older servers but refuses a server newer than its
|
||||||
|
own major version. Output is intended to load into newer PostgreSQL versions,
|
||||||
|
but loading into an older major is not guaranteed.
|
||||||
|
([`pg_dump` cross-version notes](https://www.postgresql.org/docs/current/app-pgdump.html#PG-DUMP-NOTES))
|
||||||
|
|
||||||
|
Derived Hub implications:
|
||||||
|
|
||||||
|
- Restore should target a new isolated database/cluster. If a non-atomic or
|
||||||
|
parallel restore reports any error, discard that target and start again;
|
||||||
|
never promote a partially restored database.
|
||||||
|
- The runbook must record the restore atomicity strategy and exact PostgreSQL
|
||||||
|
client/server versions. A green process exit without explicit fail-fast
|
||||||
|
options is insufficient evidence.
|
||||||
|
- A restored database must remain disconnected from real Feishu/provider
|
||||||
|
side-effect endpoints until application-level validation completes.
|
||||||
|
|
||||||
|
## Physical backup, WAL, and PITR
|
||||||
|
|
||||||
|
### Base backups
|
||||||
|
|
||||||
|
- `pg_basebackup` backs up a running PostgreSQL cluster without blocking other
|
||||||
|
clients. A full backup is an exact copy of the entire cluster and can seed
|
||||||
|
PITR or a standby; it cannot select one database or object.
|
||||||
|
([`pg_basebackup` description](https://www.postgresql.org/docs/current/app-pgbasebackup.html#APP-PGBASEBACKUP))
|
||||||
|
- It uses the replication protocol. The account needs `REPLICATION` or
|
||||||
|
superuser privilege, `pg_hba.conf` must permit it, and `max_wal_senders`
|
||||||
|
needs capacity for the backup and, when streaming WAL, another connection.
|
||||||
|
([`pg_basebackup` requirements](https://www.postgresql.org/docs/current/app-pgbasebackup.html#APP-PGBASEBACKUP))
|
||||||
|
- Default `-X stream` obtains WAL over a second connection while copying. With
|
||||||
|
`-X fetch`, WAL is collected only at the end; if required WAL was recycled,
|
||||||
|
the backup fails and is unusable.
|
||||||
|
([`pg_basebackup --wal-method`](https://www.postgresql.org/docs/current/app-pgbasebackup.html#APP-PGBASEBACKUP))
|
||||||
|
- The default waits for backup files to be durable. `--no-sync` is explicitly
|
||||||
|
unsuitable for production because an OS crash can corrupt the backup.
|
||||||
|
([`pg_basebackup --no-sync`](https://www.postgresql.org/docs/current/app-pgbasebackup.html#APP-PGBASEBACKUP))
|
||||||
|
- A manifest is generated by default, with per-file checksums. A SHA manifest
|
||||||
|
can detect malicious modification only if the manifest itself is held in a
|
||||||
|
separate trusted location or otherwise authenticated.
|
||||||
|
([`pg_basebackup` manifest options](https://www.postgresql.org/docs/current/app-pgbasebackup.html#APP-PGBASEBACKUP))
|
||||||
|
|
||||||
|
Derived minimum backup record:
|
||||||
|
|
||||||
|
- source cluster/system identifier and PostgreSQL version;
|
||||||
|
- backup tool version, start/end time, start/end LSN and WAL range;
|
||||||
|
- manifest digest and its independently protected location;
|
||||||
|
- destination object/version, byte count, exit status, and verification state;
|
||||||
|
- an explicit `VALID`, `INVALID`, or `QUARANTINED` state so a failed leftover
|
||||||
|
directory cannot enter the retention set merely because files exist.
|
||||||
|
|
||||||
|
### WAL archival and recovery
|
||||||
|
|
||||||
|
- PITR combines a physical backup with WAL replay. It can stop at a chosen time
|
||||||
|
and reach a consistent database state. Logical dumps do not contain enough
|
||||||
|
information for WAL replay.
|
||||||
|
([continuous archiving overview](https://www.postgresql.org/docs/current/continuous-archiving.html#CONTINUOUS-ARCHIVING))
|
||||||
|
- Successful recovery requires a continuous sequence of archived WAL extending
|
||||||
|
back at least to the base-backup start. PostgreSQL tells operators to set up
|
||||||
|
and test WAL archiving **before** taking the first base backup.
|
||||||
|
([continuous archiving overview](https://www.postgresql.org/docs/current/continuous-archiving.html#CONTINUOUS-ARCHIVING))
|
||||||
|
- The archive command must return zero if and only if durable archival
|
||||||
|
succeeded. Non-zero causes retry. Archive implementations should refuse to
|
||||||
|
overwrite an existing file; a repeated name may report success only when the
|
||||||
|
existing durable content is identical, and differing content must fail.
|
||||||
|
([WAL archiving](https://www.postgresql.org/docs/current/continuous-archiving.html#BACKUP-ARCHIVING-WAL))
|
||||||
|
- Repeated archive failure accumulates WAL in `pg_wal`; filling that filesystem
|
||||||
|
causes PostgreSQL to panic and remain offline until space is freed. PostgreSQL
|
||||||
|
explicitly advises monitoring archival progress.
|
||||||
|
([WAL archiving failure behaviour](https://www.postgresql.org/docs/current/continuous-archiving.html#BACKUP-ARCHIVING-WAL))
|
||||||
|
- WAL is archived after a segment is completed, so low traffic can make the
|
||||||
|
last committed change wait a long time for archival. `archive_timeout` can
|
||||||
|
bound that age by forcing segment switches, at the cost of archive growth.
|
||||||
|
([WAL archiving latency](https://www.postgresql.org/docs/current/continuous-archiving.html#BACKUP-ARCHIVING-WAL))
|
||||||
|
- WAL archival does not restore manually edited PostgreSQL configuration files
|
||||||
|
(`postgresql.conf`, `pg_hba.conf`, `pg_ident.conf`); those need a separate
|
||||||
|
file backup.
|
||||||
|
([WAL archiving caveat](https://www.postgresql.org/docs/current/continuous-archiving.html#BACKUP-ARCHIVING-WAL))
|
||||||
|
- PostgreSQL's recovery procedure keeps clients out, restores the base backup,
|
||||||
|
configures `restore_command` and `recovery.signal`, replays WAL, and requires
|
||||||
|
inspection of database contents before restoring ordinary access.
|
||||||
|
([PITR recovery procedure](https://www.postgresql.org/docs/current/continuous-archiving.html#BACKUP-PITR-RECOVERY))
|
||||||
|
- Recovery can target a timestamp, LSN, transaction ID, named restore point, or
|
||||||
|
the earliest consistent point. Only one target may be set. Timeline history
|
||||||
|
is required to navigate recovery branches, and a configured target not
|
||||||
|
reached before WAL exhaustion is a fatal recovery failure.
|
||||||
|
([recovery target settings](https://www.postgresql.org/docs/current/runtime-config-wal.html#RUNTIME-CONFIG-WAL-RECOVERY-TARGET),
|
||||||
|
[timelines](https://www.postgresql.org/docs/current/continuous-archiving.html#BACKUP-TIMELINES))
|
||||||
|
|
||||||
|
Derived Hub implications:
|
||||||
|
|
||||||
|
- The database RPO is the age of the newest **durably archived, continuous**
|
||||||
|
WAL position, not the age of the newest base backup and not merely the name
|
||||||
|
in `last_archived_wal`.
|
||||||
|
- Archive keys need a collision-proof cluster/system-identifier namespace and
|
||||||
|
create-if-absent semantics; unrelated clusters must never overwrite a shared
|
||||||
|
WAL name.
|
||||||
|
- Monitoring must cover archive failures, time since durable success, archive
|
||||||
|
lag, WAL continuity, `pg_wal` growth, free space, and base-backup age.
|
||||||
|
- PITR restores PostgreSQL only. It does not rewind workspace files. A database
|
||||||
|
recovery target has no whole-Hub meaning until paired with a workspace
|
||||||
|
recovery point and reconciliation rule.
|
||||||
|
|
||||||
|
## Verification is a ladder, not a checksum
|
||||||
|
|
||||||
|
- `pg_verifybackup` validates the backup manifest and system identifier,
|
||||||
|
detects missing/extra/changed files, checks file checksums, and parses the
|
||||||
|
required WAL with `pg_waldump`. WAL verification is currently supported only
|
||||||
|
for plain-format base backups.
|
||||||
|
([`pg_verifybackup` checks](https://www.postgresql.org/docs/current/app-pgverifybackup.html#APP-PGVERIFYBACKUP))
|
||||||
|
- PostgreSQL explicitly says these checks cannot cover everything a running
|
||||||
|
server will encounter and still requires test restores plus checks that the
|
||||||
|
database works and contains correct data.
|
||||||
|
([`pg_verifybackup` limitation](https://www.postgresql.org/docs/current/app-pgverifybackup.html#APP-PGVERIFYBACKUP))
|
||||||
|
- `pg_amcheck` can check supported heap/index structures, but currently only
|
||||||
|
ordinary/TOAST tables, materialized views, sequences, and B-tree indexes are
|
||||||
|
supported; other relation types are silently skipped. Stronger B-tree checks
|
||||||
|
can block concurrent modifications.
|
||||||
|
([`pg_amcheck`](https://www.postgresql.org/docs/current/app-pgamcheck.html#APP-PGAMCHECK))
|
||||||
|
|
||||||
|
Derived recovery-evidence ladder:
|
||||||
|
|
||||||
|
1. **Artifact verification:** backup command succeeded; manifest and all
|
||||||
|
required WAL validate; any failure marks the artifact unusable.
|
||||||
|
2. **Isolated technical restore:** restore to a clean non-production target,
|
||||||
|
reach the intended LSN/time/timeline, and prove PostgreSQL starts without
|
||||||
|
exposing it to users.
|
||||||
|
3. **Database validation:** verify migration history/schema identity, tenant
|
||||||
|
counts and relationships, critical constraints/sequences, and appropriate
|
||||||
|
corruption checks.
|
||||||
|
4. **Workspace validation:** restore the paired snapshot/version, validate file
|
||||||
|
hashes/ownership/containment, and reconcile database file-change/run state
|
||||||
|
against actual files.
|
||||||
|
5. **Application validation:** run read-only tenant-isolation and core workflow
|
||||||
|
smoke tests with external side effects disabled.
|
||||||
|
6. **Recovery report:** persist the exact backup, WAL range, workspace point,
|
||||||
|
release/schema versions, checks, observed RPO/RTO, operator, and failures.
|
||||||
|
Only a fully passing report may authorize promotion.
|
||||||
|
|
||||||
|
## Prisma Migrate production semantics
|
||||||
|
|
||||||
|
### What `migrate deploy` proves and does not prove
|
||||||
|
|
||||||
|
- In production, `migrate deploy` compares applied migration checksums with the
|
||||||
|
migration directory, warns when an applied migration was modified, and
|
||||||
|
applies pending migrations.
|
||||||
|
([Prisma 6 production workflow](https://www.prisma.io/docs/orm/v6/prisma-migrate/workflows/development-and-production#production-and-testing-environments))
|
||||||
|
- It does **not** warn when an already-applied migration is missing, detect
|
||||||
|
production schema drift, reset the database, generate Prisma Client, or use a
|
||||||
|
shadow database.
|
||||||
|
([Prisma 6 production workflow](https://www.prisma.io/docs/orm/v6/prisma-migrate/workflows/development-and-production#production-and-testing-environments))
|
||||||
|
- The migration directory is the deployable history and must be committed in
|
||||||
|
full, including `migration_lock.toml`. Prisma warns against editing or
|
||||||
|
deleting migrations already applied in production; `migrate deploy` executes
|
||||||
|
migration SQL rather than deriving a new migration from the Prisma schema.
|
||||||
|
([Prisma migration histories](https://www.prisma.io/docs/orm/prisma-migrate/understanding-prisma-migrate/migration-histories))
|
||||||
|
- Prisma uses a database advisory lock for production migration commands so
|
||||||
|
multiple migration commands do not run simultaneously. The documented lock
|
||||||
|
timeout is ten seconds and not configurable; after timeout, the caller must
|
||||||
|
retry. The lock can be disabled by environment variable.
|
||||||
|
([Prisma 6 advisory locking](https://www.prisma.io/docs/orm/v6/prisma-migrate/workflows/development-and-production#advisory-locking))
|
||||||
|
|
||||||
|
Derived Hub implication: `migrate deploy` exit zero proves that this invocation
|
||||||
|
applied or found no pending migration. It does not prove the live schema equals
|
||||||
|
the expected migration end state, that the new and previous application builds
|
||||||
|
are both compatible with it, or that the migration is operationally safe under
|
||||||
|
production data volume and traffic.
|
||||||
|
|
||||||
|
The advisory lock is a migration-command concurrency guard, not a substitute
|
||||||
|
for the repository's deployment lock or application/schema compatibility. As a
|
||||||
|
derived safety rule, production preflight should reject
|
||||||
|
`PRISMA_SCHEMA_DISABLE_ADVISORY_LOCK` unless an explicit reviewed mechanism
|
||||||
|
provides the same serialization guarantee.
|
||||||
|
|
||||||
|
### Transactions and failed migrations
|
||||||
|
|
||||||
|
- Prisma's first-party migration guidance states that PostgreSQL migrations
|
||||||
|
are not wrapped in a transaction by default. Operators may opt in by adding
|
||||||
|
`BEGIN;`/`COMMIT;` to a migration, but large transactional migrations can hold
|
||||||
|
locks longer and consume more database resources.
|
||||||
|
([Prisma transactional behaviour](https://www.prisma.io/blog/prisma-migrate-dx-primitives#what-if-schema-migrations-were-atomic))
|
||||||
|
- A failed migration records its error in `_prisma_migrations.logs`; further
|
||||||
|
migrations cannot deploy until the failed state is explicitly recovered.
|
||||||
|
Recovery is either (a) revert/repair partial steps, mark the migration rolled
|
||||||
|
back, and redeploy, or (b) manually complete the exact steps and mark it
|
||||||
|
applied.
|
||||||
|
([Prisma failed-migration recovery](https://www.prisma.io/docs/orm/prisma-migrate/workflows/patching-and-hotfixing#failed-migration))
|
||||||
|
- Prisma error codes `P3009` and `P3018` explicitly report that failed
|
||||||
|
migrations block new migration application until recovery.
|
||||||
|
([Prisma error reference](https://www.prisma.io/docs/orm/reference/error-reference#p3009))
|
||||||
|
- `migrate resolve --rolled-back` changes migration-history state so Prisma may
|
||||||
|
apply the migration again; it does not itself undo SQL already executed.
|
||||||
|
`--applied` records success without executing the migration SQL.
|
||||||
|
([failed-migration recovery](https://www.prisma.io/docs/orm/prisma-migrate/workflows/patching-and-hotfixing#failed-migration),
|
||||||
|
[`migrate resolve`](https://www.prisma.io/docs/cli/migrate/resolve))
|
||||||
|
- Prisma can generate a down SQL script, but its documentation limits the
|
||||||
|
simple rollback workflow to a **failed** up migration. To reverse a successful
|
||||||
|
migration, the supported history is a new forward migration generated from a
|
||||||
|
reverted Prisma schema. A schema down script also cannot restore lost data,
|
||||||
|
application changes, or manually authored SQL automatically.
|
||||||
|
([generating down migrations](https://www.prisma.io/docs/orm/prisma-migrate/workflows/generating-down-migrations#considerations-when-generating-down-migrations))
|
||||||
|
|
||||||
|
Derived Hub implications:
|
||||||
|
|
||||||
|
- Every migration needs a declared atomicity choice. If it is intentionally
|
||||||
|
non-transactional, the migration must be resumable/idempotent or ship an
|
||||||
|
exact, tested partial-failure procedure. Blindly rerunning a partially applied
|
||||||
|
script is unsafe.
|
||||||
|
- A database backup checkpoint before a destructive or difficult-to-reverse
|
||||||
|
migration is necessary, but restoring that checkpoint is a whole-database
|
||||||
|
recovery event—not a substitute for a compatible rollout design.
|
||||||
|
- Service restart cannot be treated as application rollback after a successful
|
||||||
|
schema migration. The previous binary may no longer understand the new
|
||||||
|
schema, and a down migration cannot recreate deleted data.
|
||||||
|
|
||||||
|
### Compatibility and rollout ordering
|
||||||
|
|
||||||
|
- Prisma documents expand-and-contract as the way to avoid downtime for fields
|
||||||
|
in active use: add the new representation, deploy code that writes both while
|
||||||
|
reading the old, migrate and verify data, switch reads, stop old writes, then
|
||||||
|
remove the old representation.
|
||||||
|
([Prisma expand-and-contract workflow](https://www.prisma.io/docs/orm/prisma-migrate/workflows/customizing-migrations#example-use-the-expand-and-contract-pattern-to-evolve-the-schema-without-downtime))
|
||||||
|
- Prisma also warns that generated “renames” may be expressed as add/drop and
|
||||||
|
lose data unless the SQL is reviewed and customized before application.
|
||||||
|
([customizing migrations](https://www.prisma.io/docs/orm/prisma-migrate/workflows/customizing-migrations#example-rename-a-field))
|
||||||
|
|
||||||
|
Derived deployment contract:
|
||||||
|
|
||||||
|
1. Migration SQL is immutable after production application and reviewed for
|
||||||
|
destructive operations, lock duration, data volume, and transaction choice.
|
||||||
|
2. It is rehearsed against a recent, anonymized production-scale restore and
|
||||||
|
its runtime/locks/disk growth are recorded.
|
||||||
|
3. Each release declares a schema compatibility window. Expand releases remain
|
||||||
|
compatible with both old and new application versions; contract/destructive
|
||||||
|
releases occur only after old binaries can no longer run.
|
||||||
|
4. Migration status/drift is checked before changing the live release. A
|
||||||
|
failed migration blocks deployment and enters an operator runbook; it must
|
||||||
|
never be hidden by repeated restart attempts.
|
||||||
|
5. Rollback means switching to a binary proven compatible with the current
|
||||||
|
schema or deploying a new forward migration. Restoring a backup is the last
|
||||||
|
resort and must also restore/reconcile the workspace recovery point.
|
||||||
|
|
||||||
|
## systemd stop, restart, and dependency semantics
|
||||||
|
|
||||||
|
- When `ExecStartPre=`, `ExecStart=`, or `ExecStartPost=` fails or times out,
|
||||||
|
service startup does not continue to the normal running process. `ExecStop=`
|
||||||
|
is skipped because startup never succeeded; `ExecStopPost=` is the hook that
|
||||||
|
still runs after a failed startup.
|
||||||
|
([systemd start/stop command semantics](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#ExecStartPre=),
|
||||||
|
[`ExecStopPost=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#ExecStopPost=),
|
||||||
|
[permanent upstream man-page source](https://github.com/systemd/systemd/blob/924552cfbacd51be0229ac6ce95fe3b69c7b7923/man/systemd.service.xml#L525-L587))
|
||||||
|
- `Restart=on-failure` covers non-zero exits and failures/timeouts in service
|
||||||
|
control processes, including `ExecStartPre=`. An explicit systemd stop does
|
||||||
|
not trigger automatic restart. Restarts are subject to unit start-rate
|
||||||
|
limiting.
|
||||||
|
([`Restart=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#Restart=),
|
||||||
|
[permanent upstream source](https://github.com/systemd/systemd/blob/924552cfbacd51be0229ac6ce95fe3b69c7b7923/man/systemd.service.xml#L856-L995))
|
||||||
|
- With no `ExecStop=`, systemd sends `KillSignal=` (default `SIGTERM`).
|
||||||
|
`TimeoutStopSec=` bounds the wait; after it expires, remaining processes can
|
||||||
|
be forcibly killed with `SIGKILL`/`FinalKillSignal=`.
|
||||||
|
([`ExecStop=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#ExecStop=),
|
||||||
|
[`TimeoutStopSec=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#TimeoutStopSec=))
|
||||||
|
- Default `KillMode=control-group` applies termination to all remaining
|
||||||
|
processes in the unit cgroup. systemd explicitly discourages `process` or
|
||||||
|
`none` because children can escape the service lifecycle and remain running
|
||||||
|
after the service is considered stopped.
|
||||||
|
([`KillMode=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.kill.html#KillMode=))
|
||||||
|
- `systemctl stop` waits for the job unless `--no-block` is used, but the command
|
||||||
|
does not fail merely because an `ExecStop=` command failed: the manager still
|
||||||
|
forcibly terminates the unit.
|
||||||
|
([`systemctl stop`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#stop%20PATTERN%E2%80%A6),
|
||||||
|
[permanent upstream source](https://github.com/systemd/systemd/blob/924552cfbacd51be0229ac6ce95fe3b69c7b7923/man/systemctl.xml#L444-L459))
|
||||||
|
- `systemctl restart` is a stop followed by a start; `ExecStop=` and
|
||||||
|
`ExecStopPost=` participate. It does not necessarily flush every unit resource
|
||||||
|
in the same way as separate completed `stop` then `start` operations.
|
||||||
|
([restart semantics](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#ExecStop=),
|
||||||
|
[`systemctl restart`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#restart%20PATTERN%E2%80%A6))
|
||||||
|
- `After=` is only an ordering relationship; it does not pull in another unit.
|
||||||
|
Requirement and ordering dependencies are independent. During shutdown the
|
||||||
|
start order is reversed, so a Hub ordered `After=postgresql.service` stops
|
||||||
|
before PostgreSQL when both are in the shutdown transaction.
|
||||||
|
([`Requires=` and ordering](https://www.freedesktop.org/software/systemd/man/latest/systemd.unit.html#Requires=),
|
||||||
|
[`Before=`/`After=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.unit.html#Before=))
|
||||||
|
- `systemctl freeze` only uses the cgroup freezer to suspend process execution;
|
||||||
|
its documented contract contains no filesystem flush or stable-image
|
||||||
|
guarantee. It is not a substitute for an application drain or a filesystem
|
||||||
|
snapshot barrier.
|
||||||
|
([permanent `systemctl freeze` man-page source](https://github.com/systemd/systemd/blob/924552cfbacd51be0229ac6ce95fe3b69c7b7923/man/systemctl.xml#L653-L675))
|
||||||
|
|
||||||
|
Derived assessment of the current unit:
|
||||||
|
|
||||||
|
- A Prisma migration failure in the current `ExecStartPre=` prevents the Hub
|
||||||
|
from starting, then `Restart=on-failure` can retry startup until rate-limited.
|
||||||
|
That is fail-closed, but it is not controlled migration recovery and can
|
||||||
|
repeatedly hit a database already left in Prisma's failed state.
|
||||||
|
- `KillSignal=SIGTERM` plus `TimeoutStopSec=30` is only a deadline. Until the Hub
|
||||||
|
installs a real termination handler that makes readiness false, stops ingress,
|
||||||
|
drains/checkpoints active work, closes Feishu and database clients, and reports
|
||||||
|
success, the unit cannot claim a graceful consistency barrier.
|
||||||
|
- A backup runbook must issue a controlled drain, verify durable work reached a
|
||||||
|
reconciled terminal/checkpoint state, stop the Hub, verify the unit and its
|
||||||
|
cgroup are inactive, and only then snapshot mutable workspace state. A zero
|
||||||
|
`systemctl stop` exit by itself is insufficient evidence.
|
||||||
|
- Recovery should not start the normal unit until the restored database and
|
||||||
|
workspace have been validated. The current unit would run pending migrations
|
||||||
|
immediately on first start, mutating the restored database before validation
|
||||||
|
and possibly destroying the ability to inspect the original recovered state.
|
||||||
|
|
||||||
|
## Filesystem and rsync boundaries
|
||||||
|
|
||||||
|
- PostgreSQL says a plain filesystem copy of a live data directory is unusable:
|
||||||
|
the database must be shut down, because tools such as `tar` do not create an
|
||||||
|
atomic snapshot and server buffering remains. A correctly implemented frozen
|
||||||
|
snapshot can work while PostgreSQL runs, but all volumes containing data,
|
||||||
|
WAL, and tablespaces must be captured simultaneously; otherwise shut down or
|
||||||
|
use the continuous-archive base-backup protocol.
|
||||||
|
([PostgreSQL filesystem backups](https://www.postgresql.org/docs/current/backup-file.html))
|
||||||
|
- For a PostgreSQL rsync backup, the documented safe pattern is a first pass
|
||||||
|
while running, followed by shutdown and a second `rsync --checksum`; the
|
||||||
|
second pass is what makes the result consistent.
|
||||||
|
([PostgreSQL filesystem backups](https://www.postgresql.org/docs/current/backup-file.html))
|
||||||
|
- At the filesystem layer, `fsfreeze` is documented to let in-progress
|
||||||
|
filesystem operations complete, halt new modifications, and flush dirty data,
|
||||||
|
metadata, and logs before returning. That is still only a filesystem-volume
|
||||||
|
guarantee; it does not replace PostgreSQL's backup protocol and does not make
|
||||||
|
separately frozen volumes simultaneous.
|
||||||
|
([util-linux `fsfreeze` manual source](https://github.com/util-linux/util-linux/blob/master/sys-utils/fsfreeze.8.adoc))
|
||||||
|
- rsync's `--delete` defaults to deletion during transfer on modern peers, so
|
||||||
|
the receiver tree changes incrementally. `--delay-updates` merely holds each
|
||||||
|
updated file and renames files into place rapidly at the end; its own manual
|
||||||
|
says this only attempts to make updates “a little more atomic” and points to
|
||||||
|
a separate parallel-tree/switch approach for stronger atomicity.
|
||||||
|
([rsync `--delete`](https://rsync.samba.org/ftp/rsync/rsync.1.html#opt--delete),
|
||||||
|
[rsync `--delay-updates`](https://rsync.samba.org/ftp/rsync/rsync.1.html#opt--delay-updates))
|
||||||
|
|
||||||
|
Derived Hub implications:
|
||||||
|
|
||||||
|
- The current in-place `rsync --delete` deployment is neither an immutable
|
||||||
|
release switch nor a backup snapshot. A failed/interrupted transfer can leave
|
||||||
|
a mixed live tree; service stop/restart cannot reconstruct the deleted prior
|
||||||
|
artifact.
|
||||||
|
- The workspace snapshot mechanism remains an explicit design choice. Whichever
|
||||||
|
storage technology is selected must prove atomicity over all workspace
|
||||||
|
volumes or use a quiesced/no-writer interval, then publish an immutable
|
||||||
|
version/checksum manifest. Merely copying a changing directory is not enough.
|
||||||
|
|
||||||
|
## Derived minimum whole-Hub recovery contract
|
||||||
|
|
||||||
|
The primary sources support the following minimum contract; exact RPO, RTO,
|
||||||
|
retention, storage vendor, and numeric schedules remain product/operations
|
||||||
|
decisions and must not be invented by implementation.
|
||||||
|
|
||||||
|
### Backup path
|
||||||
|
|
||||||
|
1. Maintain verified PostgreSQL base backups and a monitored continuous WAL
|
||||||
|
archive, or obtain and test an explicitly equivalent managed-PostgreSQL
|
||||||
|
guarantee. Keep logical dumps as a separate portability/selective-restore
|
||||||
|
layer, not the sole PITR mechanism.
|
||||||
|
2. Keep PostgreSQL configuration, deployment/release manifests, migration
|
||||||
|
history, and workspace snapshots in the backup inventory; WAL alone does not
|
||||||
|
cover them.
|
||||||
|
3. Create a whole-Hub checkpoint only after ingress is drained and active
|
||||||
|
workspace writers are stopped/reconciled, unless the eventual storage layer
|
||||||
|
supplies a proven atomic snapshot/journal spanning the necessary stores.
|
||||||
|
4. Persist one signed/immutable checkpoint manifest joining PostgreSQL
|
||||||
|
backup/WAL position, workspace snapshot/version, application revision,
|
||||||
|
migration digest, tool versions, checksums, timestamps, and validation state.
|
||||||
|
5. Treat database/WAL/workspace backups as cross-tenant sensitive data. Apply
|
||||||
|
least-privilege access, encryption, audit, deletion/retention, and an
|
||||||
|
independently protected integrity manifest; WAL contains effectively all
|
||||||
|
database changes and PostgreSQL advises protecting it from unauthorized
|
||||||
|
readers.
|
||||||
|
([WAL archive confidentiality](https://www.postgresql.org/docs/current/continuous-archiving.html#BACKUP-ARCHIVING-WAL))
|
||||||
|
|
||||||
|
### Migration path
|
||||||
|
|
||||||
|
1. Serialize deploys and backup checkpoints; run a drift/status gate before
|
||||||
|
migration, not just `migrate deploy`.
|
||||||
|
2. Rehearse each migration on a recent production-scale restore and record
|
||||||
|
runtime, locks, disk/WAL growth, transaction choice, and failure recovery.
|
||||||
|
3. Require expand/contract compatibility for online changes. A release must
|
||||||
|
declare which old/new binaries can run against each schema phase.
|
||||||
|
4. Separate migration orchestration from ordinary service restart so a failed
|
||||||
|
migration produces one durable incident state and an operator decision—not
|
||||||
|
an opaque restart loop.
|
||||||
|
5. Verify a restorable checkpoint before destructive change. Do not describe
|
||||||
|
backup restore as a normal “rollback”; it is whole-service disaster recovery.
|
||||||
|
|
||||||
|
### Restore and promotion path
|
||||||
|
|
||||||
|
1. Restore into an isolated environment with real external side effects
|
||||||
|
disabled.
|
||||||
|
2. Restore the selected PostgreSQL base backup/WAL target and the checkpoint's
|
||||||
|
exact workspace version; do not run pending Prisma migrations implicitly.
|
||||||
|
3. Prove artifact integrity, database startup, recovery target/timeline,
|
||||||
|
migration/schema identity, tenant invariants, workspace hashes/containment,
|
||||||
|
and read-only application flows.
|
||||||
|
4. Record observed recovery point and recovery time. A restore that starts but
|
||||||
|
misses its target, loses a tenant workspace, has partial migration state, or
|
||||||
|
fails any invariant is not promotable.
|
||||||
|
5. Promote explicitly, then enable ingress and external integrations. Preserve
|
||||||
|
the pre-recovery cluster/WAL and recovery report until the incident is closed.
|
||||||
|
|
||||||
|
## Decisions and evidence still required
|
||||||
|
|
||||||
|
These values are not fixed by the upstream tools and remain `OPEN` for the SaaS
|
||||||
|
production contract:
|
||||||
|
|
||||||
|
- whole-Hub RPO and RTO, separately for ordinary deletion, database corruption,
|
||||||
|
host loss, region/provider loss, and migration failure;
|
||||||
|
- base-backup, logical-dump, workspace-snapshot, WAL, and restore-drill cadence;
|
||||||
|
- immutable/off-host storage, encryption/key ownership, retention and verified
|
||||||
|
deletion policy, including org-deletion and legal-retention interactions;
|
||||||
|
- exact workspace checkpoint/journal technology and database/workspace
|
||||||
|
reconciliation semantics for in-flight or partially written runs;
|
||||||
|
- whether restore promotion is whole-platform only or supports a safe
|
||||||
|
organization/project export-import path (PostgreSQL physical PITR is
|
||||||
|
cluster-wide);
|
||||||
|
- migration transaction policy, maximum lock/downtime budget, compatibility
|
||||||
|
window, and who can authorize contract/destructive phases;
|
||||||
|
- evidence from at least one production-like full restore and one PITR drill,
|
||||||
|
including deliberate missing-WAL, corrupted-backup, failed-migration,
|
||||||
|
interrupted-workspace-write, and restart-during-recovery probes.
|
||||||
@@ -1,7 +1,7 @@
|
|||||||
# Audit backup, migration, and disaster recovery safety
|
# Audit backup, migration, and disaster recovery safety
|
||||||
|
|
||||||
Type: research
|
Type: research
|
||||||
Status: open
|
Status: resolved
|
||||||
|
|
||||||
## Question
|
## Question
|
||||||
|
|
||||||
@@ -9,3 +9,40 @@ What backup, restore, workspace/database consistency, migration compatibility,
|
|||||||
and disaster-recovery guarantees are needed to deploy current persistent data
|
and disaster-recovery guarantees are needed to deploy current persistent data
|
||||||
safely, and what evidence or tooling is missing to prove them?
|
safely, and what evidence or tooling is missing to prove them?
|
||||||
|
|
||||||
|
## Answer
|
||||||
|
|
||||||
|
The Hub has no production backup/restore control plane. PostgreSQL can be
|
||||||
|
logically dumped and restored, but authoritative state is split across the
|
||||||
|
database, mutable Project workspaces, credentials/key material, and a
|
||||||
|
schema-compatible release with no shared checkpoint or recovery manifest.
|
||||||
|
RPO/RTO, retention, off-host/immutable storage, encryption/key custody,
|
||||||
|
full-service versus Organization-selective restore, monitoring, and recovery
|
||||||
|
drills are all undefined.
|
||||||
|
|
||||||
|
Executed probes proved that the current deploy `rsync --delete` removes the
|
||||||
|
documented in-tree workspace root, a failed Project transaction leaves an
|
||||||
|
orphan workspace, and an older application query fails after a destructive
|
||||||
|
migration. A deliberately failed Prisma migration rolled back that probe's DDL
|
||||||
|
but left a blocking failed migration record; the official contract does not
|
||||||
|
make PostgreSQL migrations transactional by default. A positive custom-format
|
||||||
|
round trip restored 24 tables and all nine current migration records, proving
|
||||||
|
only database-level portability on the local PostgreSQL 14.20 setup.
|
||||||
|
|
||||||
|
The complete state map, probe commands/results, root causes, required recovery
|
||||||
|
contract, and release evidence are in the
|
||||||
|
[backup, migration, and disaster-recovery audit](../assets/backup-migration-disaster-recovery-audit.md).
|
||||||
|
Official PostgreSQL, Prisma, systemd, and rsync semantics with claim-level links
|
||||||
|
are in [backup and migration primary sources](../assets/backup-migration-primary-sources.md).
|
||||||
|
|
||||||
|
The newly-clear decision and implementation frontier is:
|
||||||
|
|
||||||
|
- [Decide recovery objectives, backup retention, and restore scope](39-decide-recovery-objectives.md)
|
||||||
|
- [Implement coherent backup sets and verified restore](40-verified-backup-restore.md)
|
||||||
|
- [Make the Project workspace lifecycle recoverable](41-recoverable-project-workspace-lifecycle.md)
|
||||||
|
- [Establish migration compatibility and failed-migration recovery](14-migration-compatibility-and-recovery.md)
|
||||||
|
|
||||||
|
Persistent-path provisioning and release-tree isolation remain owned by
|
||||||
|
[Provision a runnable non-root Hub service](10-provision-nonroot-service.md)
|
||||||
|
and [Make application releases atomic and rollbackable](12-atomic-releases-and-rollback.md).
|
||||||
|
Full recovery/cutover procedures and drill evidence remain owned by
|
||||||
|
[Write and drill production incident recovery runbooks](31-write-and-drill-incident-runbooks.md).
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
Type: task
|
Type: task
|
||||||
Status: open
|
Status: open
|
||||||
Blocked by: 01, 02, 03, 04, 05, 06, 07, 09, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38
|
Blocked by: 01, 02, 03, 04, 05, 06, 07, 09, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41
|
||||||
|
|
||||||
## Question
|
## Question
|
||||||
|
|
||||||
|
|||||||
@@ -8,4 +8,6 @@ Status: open
|
|||||||
Make clean-host installation create or validate the dedicated service identity,
|
Make clean-host installation create or validate the dedicated service identity,
|
||||||
home/cache/state/workspace paths and ownership, generate a complete production
|
home/cache/state/workspace paths and ownership, generate a complete production
|
||||||
configuration, honor the configured bind host, and fail preflight before unit
|
configuration, honor the configured bind host, and fail preflight before unit
|
||||||
installation when any required value or prerequisite is invalid.
|
installation when any required value or prerequisite is invalid. The persistent
|
||||||
|
workspace root must be outside all release trees, and preflight must reject any
|
||||||
|
path that overlaps rsync, release switching, rollback, or release pruning.
|
||||||
|
|||||||
@@ -9,4 +9,5 @@ Blocked by: 10, 11
|
|||||||
Replace in-place rsync/build with deployment locking, immutable staged release
|
Replace in-place rsync/build with deployment locking, immutable staged release
|
||||||
directories, a revision manifest, atomic current/previous switching, bounded
|
directories, a revision manifest, atomic current/previous switching, bounded
|
||||||
retention, and an explicit application rollback that preserves the live release
|
retention, and an explicit application rollback that preserves the live release
|
||||||
on build, preflight, switch, or readiness failure.
|
on build, preflight, switch, or readiness failure. Prove deploy, rollback, and
|
||||||
|
release pruning cannot traverse or delete the persistent workspace root.
|
||||||
|
|||||||
+5
-3
@@ -8,6 +8,8 @@ Blocked by: 06
|
|||||||
|
|
||||||
Define and enforce the application/schema compatibility rule for each release,
|
Define and enforce the application/schema compatibility rule for each release,
|
||||||
require expand-and-contract sequencing for destructive changes, add migration
|
require expand-and-contract sequencing for destructive changes, add migration
|
||||||
status and production-like rehearsal gates, and provide an operator procedure
|
status/checksum and production-like rehearsal gates, and move schema mutation
|
||||||
for recovering a partially failed Prisma migration without pretending that an
|
out of ordinary service start into an explicit release phase. Require a recent
|
||||||
application rollback reverses the database.
|
verified recovery checkpoint, bounded lock/runtime evidence, and an operator
|
||||||
|
procedure for resolving or restoring a failed Prisma migration without
|
||||||
|
pretending that an application rollback reverses the database.
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
Type: task
|
Type: task
|
||||||
Status: open
|
Status: open
|
||||||
Blocked by: 06, 12, 13, 14, 23, 24, 25, 26, 27, 28, 29, 30
|
Blocked by: 06, 12, 13, 14, 23, 24, 25, 26, 27, 28, 29, 30, 39, 40, 41
|
||||||
|
|
||||||
## Question
|
## Question
|
||||||
|
|
||||||
@@ -13,4 +13,6 @@ deploy/rollback, host restart during work, and suspected tenant exposure. Give
|
|||||||
each alert an owner and safe read-first commands, require confirmed/audited
|
each alert an owner and safe read-first commands, require confirmed/audited
|
||||||
scope for destructive or replay actions, execute every runbook against
|
scope for destructive or replay actions, execute every runbook against
|
||||||
production-shaped staging, and archive evidence that detection, mitigation,
|
production-shaped staging, and archive evidence that detection, mitigation,
|
||||||
reconciliation, verification, escalation, and alert clearing all work.
|
reconciliation, verification, escalation, and alert clearing all work. Include
|
||||||
|
database/workspace backup failure, failed migration, full-host restore and
|
||||||
|
cutover, duplicate/gap review, fallback, and achieved RPO/RTO evidence.
|
||||||
|
|||||||
@@ -0,0 +1,14 @@
|
|||||||
|
# Decide recovery objectives, backup retention, and restore scope
|
||||||
|
|
||||||
|
Type: grilling
|
||||||
|
Status: open
|
||||||
|
|
||||||
|
## Question
|
||||||
|
|
||||||
|
For the initial SaaS service, decide measurable RPO/RTO targets, backup and
|
||||||
|
restore-drill frequency, retention and deletion-in-backups policy, off-host
|
||||||
|
copy/failure-domain requirements, logical export versus physical base-backup
|
||||||
|
and continuous-WAL/PITR coverage, encryption and key custody, operator ownership,
|
||||||
|
and whether the supported promise is full-service disaster recovery only or
|
||||||
|
also Organization-selective export/restore. Record compliance-dependent items
|
||||||
|
as explicit OPEN inputs rather than inventing defaults.
|
||||||
@@ -0,0 +1,16 @@
|
|||||||
|
# Implement coherent backup sets and verified restore
|
||||||
|
|
||||||
|
Type: task
|
||||||
|
Status: open
|
||||||
|
Blocked by: 10, 12, 14, 19, 23, 24, 25, 26, 29, 39, 41
|
||||||
|
|
||||||
|
## Question
|
||||||
|
|
||||||
|
Implement the decided recovery policy as cataloged, checksummed, encrypted,
|
||||||
|
off-host backup sets covering PostgreSQL logical exports plus physical base
|
||||||
|
backup/continuous WAL or a proven managed equivalent, Project workspaces,
|
||||||
|
compatible release/migration identity, configuration inventory, org-secret
|
||||||
|
ciphertext, and separately recoverable key material. Establish a proven
|
||||||
|
cross-store checkpoint, retention and alerting, partial-artifact cleanup,
|
||||||
|
isolated restore with outbound integrations disabled, lifecycle reconciliation,
|
||||||
|
critical SaaS journey verification, and measured RPO/RTO drills.
|
||||||
+14
@@ -0,0 +1,14 @@
|
|||||||
|
# Make the Project workspace lifecycle recoverable
|
||||||
|
|
||||||
|
Type: task
|
||||||
|
Status: open
|
||||||
|
Blocked by: 10, 26, 34
|
||||||
|
|
||||||
|
## Question
|
||||||
|
|
||||||
|
Move persistent workspaces outside every mutable release tree and make Project
|
||||||
|
workspace allocation, database commit, archive/deletion, relocation, restore,
|
||||||
|
and orphan/missing-path reconciliation explicit and crash-safe. Prove failed
|
||||||
|
Project creation leaves no silent orphan, deployment/rollback/pruning cannot
|
||||||
|
delete customer files, restored absolute paths are validated or safely rebased,
|
||||||
|
and every inconsistency is observable rather than auto-hidden.
|
||||||
@@ -40,6 +40,7 @@ rollback and recovery, and no known critical security or data-integrity gaps.
|
|||||||
- [Audit run lifecycle and data integrity under failure](issues/03-audit-run-lifecycle-integrity.md) — retain the database-enforced project-lock uniqueness and normal happy path, but replace receipt-before-success, process-local queues, independent run/session/lock effects, silent delivery/audit loss, and partial startup reset with durable, recoverable state machines; explicitly decide failed-run workspace semantics.
|
- [Audit run lifecycle and data integrity under failure](issues/03-audit-run-lifecycle-integrity.md) — retain the database-enforced project-lock uniqueness and normal happy path, but replace receipt-before-success, process-local queues, independent run/session/lock effects, silent delivery/audit loss, and partial startup reset with durable, recoverable state machines; explicitly decide failed-run workspace semantics.
|
||||||
- [Define production observability and incident recovery](issues/04-define-observability-recovery.md) — keep the existing Fastify/Prisma/systemd signal fragments, but production requires dependency-aware readiness, typed and correlated errors, durable-work telemetry, an explicit SLO/retention/alerting control plane, and recovery runbooks proven by fault drills.
|
- [Define production observability and incident recovery](issues/04-define-observability-recovery.md) — keep the existing Fastify/Prisma/systemd signal fragments, but production requires dependency-aware readiness, typed and correlated errors, durable-work telemetry, an explicit SLO/retention/alerting control plane, and recovery runbooks proven by fault drills.
|
||||||
- [Define initial abuse and capacity controls](issues/05-define-abuse-capacity-controls.md) — require layered fail-closed limits, durable Organization-fair admission, explicit backpressure, bounded files/storage/runs, soft usage alerts, and audited emergency workload brakes; calibrate numerical ceilings under production-like load.
|
- [Define initial abuse and capacity controls](issues/05-define-abuse-capacity-controls.md) — require layered fail-closed limits, durable Organization-fair admission, explicit backpressure, bounded files/storage/runs, soft usage alerts, and audited emergency workload brakes; calibrate numerical ceilings under production-like load.
|
||||||
|
- [Audit backup, migration, and disaster recovery safety](issues/06-audit-backup-migration-recovery.md) — separate persistent state from releases, define a database/workspace checkpoint and recovery objectives, make migrations compatibility-gated, and prove complete off-host backup sets through isolated restore and cutover drills.
|
||||||
|
|
||||||
## Fog
|
## Fog
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user