Files
curriculum-project-hub/hub/deploy/README.md
T

5.5 KiB

Alpha Silo service installation

The supervised alpha runs one Organization per named Silo. The supported host has systemd, PostgreSQL, Node.js 24+, pg_isready, runuser, setpriv, bubblewrap, socat, pg_dump, tar, sha256sum, and a compatible cph.

Each Silo needs its own logical database/role, instance id, service user, environment, keyring, workspace, port, domain, and host filesystem quota. Application code lives in immutable versioned directories under /srv/curriculum-project-hub/releases/; releases contain no customer state.

Install an instance

sudo BASE=/srv/curriculum-project-hub \
  HUB_DIR=/srv/curriculum-project-hub/releases/<release-id>/hub \
  INSTANCE_ID=org-a \
  PORT=8788 \
  MEMORY_MAX=16G CPU_QUOTA=400% TASKS_MAX=512 \
  bash /srv/curriculum-project-hub/releases/<release-id>/hub/deploy/install_service.sh

The first invocation creates root-owned 0600 files below /srv/curriculum-project-hub/.secrets/org-a/ and exits 78. Copy the keyring to a separately protected recovery destination and fill every blank in platform.env. Before rerunning the installer, migrate the empty Silo database:

sudo bash -c '
  set -euo pipefail
  set -a; . /srv/curriculum-project-hub/.secrets/org-a/platform.env; set +a
  node /srv/curriculum-project-hub/releases/<release-id>/hub/node_modules/prisma/build/index.js \
    migrate deploy \
    --schema /srv/curriculum-project-hub/releases/<release-id>/hub/prisma/schema.prisma
'

Then rerun the installer so it can complete preflight and install the stopped unit. Resource numbers above are examples, not defaults; the operator must set measured ceilings explicitly.

The installer refuses overlapping release/persistent paths, root service users, missing Node/cph/PostgreSQL/bubblewrap prerequisites, invalid credentials, and unreachable database roles. The installed unit is cph-hub-org-a.service and uses systemd LoadCredential; the service user cannot read the source keyring.

Default state paths are:

/var/lib/cph-hub/org-a/home
/var/lib/cph-hub/org-a/state
/var/lib/cph-hub/org-a/workspaces
/var/cache/cph-hub/org-a

Bootstrap the only Organization

Prepare a root-owned 0600 JSON file containing organization, owner, feishu, provider, and optionally teams. Secrets never appear in command-line arguments.

{
  "organization": { "id": "org_a", "slug": "org-a", "name": "Org A" },
  "owner": { "openId": "ou_owner", "displayName": "Owner" },
  "feishu": {
    "appId": "cli_xxx",
    "appSecret": "...",
    "botOpenId": "ou_bot"
  },
  "provider": {
    "providerId": "openrouter",
    "baseUrl": "https://openrouter.ai/api",
    "authToken": "..."
  },
  "teams": [{ "slug": "teachers", "name": "Teachers" }]
}
sudo bash -c '
  set -euo pipefail
  set -a; . /srv/curriculum-project-hub/.secrets/org-a/platform.env; set +a
  node /srv/curriculum-project-hub/releases/<release-id>/hub/dist/deployment/bootstrap-silo-cli.js \
    --config-file /root/org-a-bootstrap.json \
    --keyring-file /srv/curriculum-project-hub/.secrets/org-a/secret-keyring.json
'

Bootstrap encrypts and activates both connection records and is rerunnable when provider configuration fails after Organization creation. Hub startup refuses to become healthy unless the sole Organization, active provider, active Feishu application, and Feishu listener are ready. Delete the bootstrap input after the recovery copy is current. Production has no process-global Feishu/provider credential fallback.

Start and rotate

sudo systemctl start cph-hub-org-a.service
sudo systemctl status cph-hub-org-a.service

KEK rotation requires the named unit to be stopped. Source its instance env so HUB_SYSTEMD_UNIT and DATABASE_URL identify the same Silo:

sudo bash -c '
  set -euo pipefail
  systemctl stop cph-hub-org-a.service
  set -a; . /srv/curriculum-project-hub/.secrets/org-a/platform.env; set +a
  node /srv/curriculum-project-hub/hub/dist/deployment/rotate-secret-kek.js \
    --keyring-file /srv/curriculum-project-hub/.secrets/org-a/secret-keyring.json
  systemctl start cph-hub-org-a.service
'

Stop the unit and run backup_silo.sh with distinct root-owned destinations:

sudo INSTANCE_ID=org-a \
  ENV_FILE=/srv/curriculum-project-hub/.secrets/org-a/platform.env \
  KEYRING_FILE=/srv/curriculum-project-hub/.secrets/org-a/secret-keyring.json \
  BUSINESS_BACKUP_DIR=/backup/business \
  RECOVERY_BACKUP_DIR=/separate-recovery \
  bash hub/deploy/backup_silo.sh

The business set contains the PostgreSQL custom dump and workspace archive. The separate recovery set contains the keyring and environment. Both include checksums; neither destination may be the live host's only disk.

Restore into a separate drill database/workspace, verify checksums, then run:

set -a; . /path/to/restored/platform.env; set +a
node hub/dist/deployment/restore-preflight.js \
  --keyring-file /path/to/restored/secret-keyring.json

Traffic stays disabled until the sole Organization and every Feishu/provider envelope authenticate, the workspace exists, and an end-to-end test run passes.

For the first supervised demo, install → bootstrap → start → health check is the deployment gate. A completed off-host restore drill is an Alpha hardening item, not a prerequisite for bringing up that first controlled Silo.

The default bind remains loopback; expose it only through a TLS reverse proxy. The platform admin surface is not part of the alpha and must not be exposed.