Files
curriculum-project-hub/hub/deploy/README.md
T

152 lines
5.5 KiB
Markdown

# Alpha Silo service installation
The supervised alpha runs one Organization per named Silo. The supported host
has systemd, PostgreSQL, Node.js 24+, `pg_isready`, `runuser`, `setpriv`,
bubblewrap, `socat`, `pg_dump`, `tar`, `sha256sum`, and a compatible `cph`.
Each Silo needs its own logical database/role, instance id, service user,
environment, keyring, workspace, port, domain, and host filesystem quota.
Application code lives in immutable versioned directories under
`/srv/curriculum-project-hub/releases/`; releases contain no customer state.
## Install an instance
```sh
sudo BASE=/srv/curriculum-project-hub \
HUB_DIR=/srv/curriculum-project-hub/releases/<release-id>/hub \
INSTANCE_ID=org-a \
PORT=8788 \
MEMORY_MAX=16G CPU_QUOTA=400% TASKS_MAX=512 \
bash /srv/curriculum-project-hub/releases/<release-id>/hub/deploy/install_service.sh
```
The first invocation creates root-owned `0600` files below
`/srv/curriculum-project-hub/.secrets/org-a/` and exits 78. Copy the keyring to
a separately protected recovery destination and fill every blank in
`platform.env`. Before rerunning the installer, migrate the empty Silo database:
```sh
sudo bash -c '
set -euo pipefail
set -a; . /srv/curriculum-project-hub/.secrets/org-a/platform.env; set +a
node /srv/curriculum-project-hub/releases/<release-id>/hub/node_modules/prisma/build/index.js \
migrate deploy \
--schema /srv/curriculum-project-hub/releases/<release-id>/hub/prisma/schema.prisma
'
```
Then rerun the installer so it can complete preflight and install the stopped
unit. Resource numbers above are examples, not defaults; the operator must set
measured ceilings explicitly.
The installer refuses overlapping release/persistent paths, root service users,
missing Node/cph/PostgreSQL/bubblewrap prerequisites, invalid credentials, and
unreachable database roles. The installed unit is `cph-hub-org-a.service` and
uses systemd `LoadCredential`; the service user cannot read the source keyring.
Default state paths are:
```text
/var/lib/cph-hub/org-a/home
/var/lib/cph-hub/org-a/state
/var/lib/cph-hub/org-a/workspaces
/var/cache/cph-hub/org-a
```
## Bootstrap the only Organization
Prepare a root-owned `0600` JSON file containing
`organization`, `owner`, `feishu`, `provider`, and optionally `teams`. Secrets
never appear in command-line arguments.
```json
{
"organization": { "id": "org_a", "slug": "org-a", "name": "Org A" },
"owner": { "openId": "ou_owner", "displayName": "Owner" },
"feishu": {
"appId": "cli_xxx",
"appSecret": "...",
"botOpenId": "ou_bot"
},
"provider": {
"providerId": "openrouter",
"baseUrl": "https://openrouter.ai/api",
"authToken": "..."
},
"teams": [{ "slug": "teachers", "name": "Teachers" }]
}
```
```sh
sudo bash -c '
set -euo pipefail
set -a; . /srv/curriculum-project-hub/.secrets/org-a/platform.env; set +a
node /srv/curriculum-project-hub/releases/<release-id>/hub/dist/deployment/bootstrap-silo-cli.js \
--config-file /root/org-a-bootstrap.json \
--keyring-file /srv/curriculum-project-hub/.secrets/org-a/secret-keyring.json
'
```
Bootstrap encrypts and activates both connection records and is rerunnable when
provider configuration fails after Organization creation. Hub startup refuses
to become healthy unless the sole Organization, active provider, active Feishu
application, and Feishu listener are ready. Delete the bootstrap input after
the recovery copy is current. Production has no process-global Feishu/provider
credential fallback.
## Start and rotate
```sh
sudo systemctl start cph-hub-org-a.service
sudo systemctl status cph-hub-org-a.service
```
KEK rotation requires the named unit to be stopped. Source its instance env so
`HUB_SYSTEMD_UNIT` and `DATABASE_URL` identify the same Silo:
```sh
sudo bash -c '
set -euo pipefail
systemctl stop cph-hub-org-a.service
set -a; . /srv/curriculum-project-hub/.secrets/org-a/platform.env; set +a
node /srv/curriculum-project-hub/hub/dist/deployment/rotate-secret-kek.js \
--keyring-file /srv/curriculum-project-hub/.secrets/org-a/secret-keyring.json
systemctl start cph-hub-org-a.service
'
```
## Backup and restore drill (recommended after the demo is running)
Stop the unit and run `backup_silo.sh` with distinct root-owned destinations:
```sh
sudo INSTANCE_ID=org-a \
ENV_FILE=/srv/curriculum-project-hub/.secrets/org-a/platform.env \
KEYRING_FILE=/srv/curriculum-project-hub/.secrets/org-a/secret-keyring.json \
BUSINESS_BACKUP_DIR=/backup/business \
RECOVERY_BACKUP_DIR=/separate-recovery \
bash hub/deploy/backup_silo.sh
```
The business set contains the PostgreSQL custom dump and workspace archive. The
separate recovery set contains the keyring and environment. Both include
checksums; neither destination may be the live host's only disk.
Restore into a separate drill database/workspace, verify checksums, then run:
```sh
set -a; . /path/to/restored/platform.env; set +a
node hub/dist/deployment/restore-preflight.js \
--keyring-file /path/to/restored/secret-keyring.json
```
Traffic stays disabled until the sole Organization and every Feishu/provider
envelope authenticate, the workspace exists, and an end-to-end test run passes.
For the first supervised demo, install → bootstrap → start → health check is the
deployment gate. A completed off-host restore drill is an Alpha hardening item,
not a prerequisite for bringing up that first controlled Silo.
The default bind remains loopback; expose it only through a TLS reverse proxy.
The platform admin surface is not part of the alpha and must not be exposed.