Files
curriculum-project-hub/.scratch/saas-production-readiness/assets/clean-host-deployment-audit.md
T

9.7 KiB

Clean-host deployment and rollback audit

Verdict

The current hub/deploy path is not a production deployment mechanism. It can update a manually-prepared host in the happy path, but a clean supported host cannot become a runnable Hub from the documented inputs, and an upgrade failure can leave the live application tree partially replaced with no application rollback path.

The accepted initial topology remains viable: one Linux/systemd host, PostgreSQL, Feishu WebSocket ingress, and a non-root Hub service. The blocker is the release contract and its verification, not a need to replace that topology.

Executed evidence

Static gates

  • bash -n hub/deploy/install_service.sh and bash -n hub/deploy/deploy_platform.sh pass.
  • ShellCheck reports only informational SC2029 findings for intentional client-side expansion of remote paths. It does not find a shell syntax defect.
  • A source-to-installer comparison finds six requireEnv settings in hub/src/server.ts, but the generated production environment omits HUB_PROJECT_WORKSPACE_ROOT and HUB_SESSION_SECRET.

Seeded configuration cannot boot the built server

The installer was run against a temporary BASE, with the real built Hub as HUB_DIR. It exited successfully after creating platform.env. The generated file was then used as the production configuration while supplying the model and Feishu values that its message asks the operator to fill. The built server exited with status 1:

[hub] missing required env: HUB_PROJECT_WORKSPACE_ROOT

After that is supplied, HUB_SESSION_SECRET is the next required omission. HUB_PUBLIC_BASE_URL is also absent; its code fallback is localhost, which is not a valid public OAuth callback for a production deployment.

The service identity is referenced but not provisioned

The installer was executed with SERVICE_USER=definitely-missing-cph-user and mocked install/systemctl boundaries. It returned success and rendered:

User=definitely-missing-cph-user
WorkingDirectory=.../hub
EnvironmentFile=.../.secrets/platform.env

There is no getent, id, useradd, or equivalent validation/provisioning path. It also creates no service home, cache, state, or project-workspace directory. This contradicts ADR-0018's implemented deployment invariant that the Hub runs as a dedicated unprivileged service user. A genuinely clean host will fail at systemd user resolution; a manually-created user can still fail later when the agent SDK, cph cache, or project workspace needs a writable path.

A failed build mutates the live tree before failure

deploy_platform.sh was run with mocked SSH and rsync boundaries. The first remote build was forced to exit 42. The observed order was:

ssh:mkdir -p '/srv/curriculum-project-hub/hub'
rsync:live-tree-mutated
ssh:cd '/srv/curriculum-project-hub/hub' && npm ci && npm run build

The deploy command correctly propagated status 42, but it had already applied rsync --delete to the one live directory. There is no staged release, immutable release identity, current/previous pointer, deployment lock, or rollback command. npm ci also replaces that live directory's dependency tree before the new build has passed.

Blocking gaps

1. Incomplete clean-host bootstrap

install_service.sh writes a unit containing User=cph-hub without creating or validating that account. It does not create owned state/cache/workspace paths, and its generated environment omits required settings. On its first invocation it deliberately exits after seeding the environment, after which deploy_platform.sh proceeds to restart a unit that has not yet been installed. The documented workflow is therefore neither one-shot nor complete.

2. The runtime product is not fully delivered

The deploy uploads only hub/. The target prerequisite list does not include a preinstalled cph, and no release step builds or installs crates/cph-cli. The Hub starts without checking it: the first agent cph check or cph build returns exit 127 while /api/healthz remains green. Bubblewrap is asserted by the unit, but Node/npm, curl, PostgreSQL connectivity, service-user namespace support, writable storage, and cph --version are not checked as one preflight.

3. Releases are in-place and have no application rollback

Source, dependencies, build output, and the running service all share one directory. A build failure, interrupted rsync, or unsuccessful restart can leave a mixed tree. No release manifest records the git revision, Node version, dependency lock digest, migration set, or cph version that was deployed. Concurrent deploys are not serialized.

4. The deployment health check is only liveness

/api/healthz returns a constant success response after the HTTP server starts. It does not verify database access, migration state, workspace writability, cph, the agent sandbox, or Feishu listener startup. The listener's start promise is not awaited or reflected in readiness. A release can therefore be declared successful while its defining workflows are unavailable.

The service unit also claims graceful shutdown, but server.ts installs no SIGTERM/SIGINT handler and never calls app.close() or closes the Feishu listener. Node's default SIGTERM handling terminates the process; Fastify's documented graceful path begins only when fastify.close() is invoked. The full data-integrity consequence belongs to the run-lifecycle audit, but the current deployment contract must not claim draining behavior it does not have.

References:

5. Application rollback and schema migration are not coordinated

The unit runs prisma migrate deploy immediately before every start. Current migration history includes column drops, index replacement, and data deduplication; there is no declaration that the previous application remains compatible after each migration. Switching application code back after a successful migration is therefore not generally safe.

Prisma documents that migrate deploy applies pending migrations but does not detect schema drift, and a partially failed migration requires explicit operator recovery. Its recommended zero-downtime direction is an expand-and-contract sequence. The repository has no pre-deploy migration status gate, production-data rehearsal, backup checkpoint, expand/contract rule, or failed-migration runbook.

References:

6. Network binding configuration is misleading

The installer writes HOST=127.0.0.1, but server.ts ignores it and binds 0.0.0.0. An operator following the generated configuration can unintentionally expose the Hub directly instead of binding it behind the intended TLS reverse proxy. Public URL validation and proxy topology are not part of deployment preflight.

Required release contract

A production-capable implementation on the accepted topology should establish these invariants:

  1. Provisioned identity and storage. Installation creates or validates a non-login cph-hub user and explicit persistent state, cache, home, and workspace paths with tested ownership.
  2. Complete validated configuration. A preflight validates every required secret/value, the public HTTPS URL, bind address, PostgreSQL connectivity, writable storage, bwrap, and the exact compatible cph binary before the live service is touched.
  3. Immutable staged releases. Each revision is built in a new release directory, identified by revision and manifest. The live service points to a current symlink only after build and preflight pass. Deploys are locked.
  4. Complete product artifact. The release carries both the Hub and the matching cph executable; readiness checks the binary/version rather than deferring failure to an agent run.
  5. Real readiness. Liveness remains cheap, while readiness proves the dependencies needed to accept work. Deployment success waits for readiness, not merely an open HTTP port.
  6. Application rollback. A failed switch returns current to the previous release and verifies readiness. Database changes are never silently reversed.
  7. Migration compatibility. Every release states whether its migration set is backward-compatible with the previous application. Future destructive changes use expand/contract releases; failed migrations follow an explicit recovery runbook tied to backup evidence.
  8. Observable lifecycle. Release revision, migration status, readiness failures, restart/drain outcomes, and rollback results are present in command output and service logs.

Verification route

The implementation is not complete until an automated Linux test drives these scenarios with the real scripts and a production-like PostgreSQL instance:

  • first install on a host without the service user or directories;
  • incomplete configuration fails before unit installation/restart;
  • staged build failure leaves the current release byte-for-byte unchanged;
  • successful deploy reports revision and passes readiness;
  • failed new release switches back to the previous application;
  • missing/wrong cph, unwritable workspace, unavailable database, and failed migration each produce a distinct non-zero deployment failure;
  • repeated installation/deployment is idempotent;
  • concurrent deployments serialize or one fails clearly.