Files
curriculum-project-hub/hub/deploy/NEW_SILO_RUNBOOK.md
T

5.3 KiB

New Alpha Silo runbook

Use this runbook for a new Organization. A Silo is not merely another row in the existing database: it has an independent PostgreSQL role/database, Linux service identity, systemd unit, secret directory/keyring, workspace, skill store, loopback port, domain, Feishu app and provider credential.

The repeatable entry point is:

bash hub/deploy/new_silo.sh

It gathers values and writes a private deployment bundle below ~/.cph-silo-plans/<instance-id>/. The directory and all generated files are mode 0700/0600. Never commit, paste into chat, or copy that directory into an immutable release. Run the wizard once per Organization; do not edit a copy from another Organization.

Inputs to collect

The platform operator chooses the host, release, unique instance id, short workspace path, unique loopback port, database name/role, resource ceilings and public domain. The Organization administrator supplies:

  • Organization display name and slug;
  • Feishu App ID, App Secret and bot Open ID;
  • the first OWNER's Open ID and display name;
  • an Organization-exclusive provider token and provider base URL.

The Feishu app is scoped to this Silo. OAuth users authenticated by that app are automatically admitted to this Organization; OWNER remains the initial privileged membership used for controlled administration and bootstrap. An empty initial team list does not block the Alpha.

Obtain a person's Open ID from the Feishu user-get documentation page by clicking the user_id value picker and selecting the person. Configure the redirect URL shown by the generated OPERATE.md; it is required for first-time OAuth admission.

Host prerequisites

Before the first Silo on a host, install Node.js 24+, npm, rsync, PostgreSQL server/client, pg_isready, systemd, bubblewrap, socat, runuser, setpriv, pg_dump, tar, sha256sum, Nginx, Certbot, Typst and a compatible cph binary. Configure outbound proxying independently at host/service level and verify both GitHub and the selected model provider through it. The wizard does not install or select proxy nodes.

Use a deployment account with only the required passwordless sudo operations. The application itself always runs as the installer-created non-root cph-<instance-id> user. PostgreSQL must have a separate login role and logical database per Silo even when all Silo databases share one PostgreSQL server.

Execute a generated bundle

Open the bundle's OPERATE.md and perform its numbered gates in order:

  1. DNS, Feishu redirect URL, permissions and event subscription.
  2. Dedicated PostgreSQL role and database.
  3. Immutable release publication. Exit 78 is expected only when the first installer call seeds this instance's keyring and environment template.
  4. Root-owned secret installation and off-host keyring recovery copy.
  5. Prisma migration, stopped service installation and idempotent bootstrap.
  6. Explicit runtime role/skill installation.
  7. Nginx/TLS, service start, internal/external health and Feishu acceptance.
  8. First off-host backup.

Every command must fail fast. Do not add || true around install, migration, bootstrap, Nginx validation, health, or backup checks. If a check fails, retain the unit logs and the exact failed stage before changing configuration.

Default runtime role and skills

Roles and skills are dynamic Silo state, not release contents. The release contains only the management CLI. Stage approved skill directories on the host and install them with agent_config.sh; PostgreSQL records role definitions and skill selections while the versioned skill content lives in the Silo state directory and is included in backups.

For the current Alpha, upsert the default role with the agreed education assistant system prompt, model selection, tools policy and approved skill list. Keep the prompt in a root-controlled staging file, pass it via --system-prompt-file, then verify with agent_config.sh list. Do not sync an operator's entire personal skill directory: each enabled skill must be reviewed and named explicitly. Typst being installed on the host and the Typst skill being enabled are separate gates.

Acceptance gate

A Silo is ready only when all of the following pass:

  • its systemd service is active and both loopback and TLS health endpoints pass;
  • startup preflight sees exactly the configured Organization plus active Feishu and provider connections;
  • OWNER completes OAuth and can interact with the bot;
  • a non-OWNER completes OAuth and can interact with the same app/Organization;
  • two Feishu groups bind distinct projects/sessions;
  • a restart preserves persisted session cursor behavior;
  • the configured provider/model succeeds through the host proxy;
  • every enabled skill is listed, and a Typst task succeeds if Typst is enabled;
  • the first business backup and separate recovery backup are stored off-host.

Rollback boundary

For a failed code release, point only this instance back to the previous immutable release and rerun its installer with the same instance parameters. Do not roll back a database after migrations unless that release's documented database compatibility permits it. Preserve the environment, keyring, database, workspace and skill store. For destructive recovery, stop traffic and use the separate restore procedure; never substitute another Silo's state.