32 KiB
Production observability and incident-recovery audit
Date: 2026-07-10
Scope: the hub HTTP service, Feishu WebSocket ingress and outbound calls, PostgreSQL/Prisma, agent execution, the cph subprocess, project workspaces, and the packaged deployment path. This is a release-gate audit, not a proposal to change product semantics. Where the Lean contract is open, this report asks for an operational decision rather than inventing one.
Verdict
No-go for unattended SaaS production deployment. The service exposes a liveness-shaped endpoint, emits some structured component logs, and persists useful run/audit fields, but it cannot currently answer the minimum production questions:
- Can this instance accept a real request now (database reachable, Feishu listener authenticated, required runtime executable and workspace usable)?
- Which tenant, request, Feishu event, run, subprocess, and outbound message participated in one failed user action?
- How many requests/runs/messages are failing, stuck, delayed, retried, or duplicated?
- Which condition should page an operator, and which documented recovery action is safe?
- After a restart, deploy, rollback, or dependency outage, what evidence proves that work was neither lost nor silently acknowledged?
Three fault injections demonstrate that the current /api/healthz and logging surface can be green while the service is already unable to serve its core paths:
- After startup, cutting PostgreSQL sockets left two consecutive
/api/healthzrequests at HTTP 200. - An invalid Feishu app ID was rejected by the installed Feishu SDK while
/api/healthzremained HTTP 200. - With PostgreSQL cut,
/api/mereturned HTTP 400bad_request, exposed a raw Prisma invocation in the response, produced no request-scoped application error log, and only emitted two unscoped Prisma error lines.
Together with the already-confirmed lifecycle and deployment failures, this means green process state, green health checks, and even a visible user error do not prove that the service or its durable workflows are healthy.
Evidence and positive signals already present
The starting point is not empty. These are useful foundations to retain:
- Fastify's built-in HTTP logging records request/response facts with an internal request ID through the server logger (
hub/src/server.ts). - The same Fastify logger is threaded into the Feishu trigger/runtime, while Prisma emits its configured error/warn events (and optional query events) (
hub/src/server.ts,hub/src/db.ts). AgentRunpersists status, timestamps, error/result fields, token/cost fields, and links to project/session;ProjectAgentLockpersists a heartbeat;AuditEntryandAgentMessageprovide potential durable trace anchors (hub/prisma/schema.prisma).- Agent execution emits intermediate activity to the streaming Feishu card, and lock heartbeat activity is attempted during execution (
hub/src/agent/runner.ts,hub/src/feishu/trigger.ts). - Deployment is managed by systemd, so restart policy, exit status, and journal collection have an existing operational host surface (
hub/deploy/cph-hub.service). - Database migrations are explicit Prisma artifacts rather than ad-hoc startup DDL (
hub/prisma/migrations). - Existing tests cover many individual authorization, trigger, lock, queue, runner, and API behaviors. They are valuable regression evidence, although they are not yet a deploy/recovery acceptance suite.
These signals are fragmented rather than joined into an operational contract. Several writes are best-effort or swallowed, identifiers are not propagated end-to-end, there is no metrics/alert surface, and the health endpoint does not test dependencies.
Failure-to-signal matrix
| Surface | Failure or degraded condition | Current externally visible signal | What is missing / why the signal is unsafe | Minimum production signal and recovery evidence |
|---|---|---|---|---|
| HTTP | PostgreSQL becomes unreachable after startup | Affected request fails; /api/healthz stays 200 |
Health is process liveness only. /api/me misclassifies the dependency failure as client 400 bad_request, leaks raw Prisma call text, and creates no request-scoped error record. An LB/operator cannot distinguish bad input from platform outage. |
Separate liveness and readiness. Readiness performs a bounded DB probe and returns non-2xx on failure. Unexpected exceptions map to sanitized 5xx with a stable error code and request_id; structured log contains the same ID and normalized Prisma category. Alert on readiness failure and 5xx/error-code rate. |
| HTTP | Handler latency, event-loop stall, exhausted connections, overload | Access log only after a response is completed | No in-flight count, latency histogram, timeout counter, saturation signal, or request correlation. A hanging request may emit no timely evidence. | Request count, duration, in-flight requests, timeout/cancellation count, response class, and event-loop/runtime/process saturation. Bound request duration and log start/end/error with one request_id. |
| Feishu ingress | Invalid/revoked app credentials; WebSocket connect/auth failure | SDK rejects/logs asynchronously; process and /api/healthz remain green |
Startup success is reported before listener readiness. No listener state, last-connected/last-event timestamp, reconnect counter, or readiness dependency. The installed SDK already exposes onReady, onError, onReconnecting, onReconnected, getConnectionStatus, and close, but the Hub does not retain or wire the WSClient. systemd can perpetually supervise a process that cannot receive work. |
Listener state machine (starting/connected/degraded/stopped), last successful connection/event, reconnect/auth-failure counters, bounded startup readiness gate, and alert on auth failure or prolonged disconnect. Credential contents must never be logged. |
| Feishu ingress | Duplicate delivery, handler failure after receipt, process restart with queued/batched work | Event receipt/dedup row and selected warning logs | The run-lifecycle audit proved receipt-before-success and in-memory batching/queue loss. There is no durable processing state/lease, retry age, dead-letter view, or alert for accepted-but-unstarted work. | Durable ingress states and timestamps, attempts, last normalized error, event/chat/project correlation, queue depth/oldest age, dead-letter count. Recovery drill proves restart resumes or deliberately redrives exactly the intended work. See run lifecycle audit. |
| Feishu outbound | API timeout/ambiguous success, rate limit, invalid credentials, permanent card patch/send failure | Some helpers log a warning; several message helpers catch failure and return null |
No idempotent delivery record, delivery state, attempt count, provider request/message ID, retry delay, terminal failure metric, or user-visible reconciliation path. Fault injection proved a run can be marked completed although nothing was delivered. | Durable outbox/delivery state; idempotency key; provider message/request ID; attempts, latency and normalized provider error; queue depth/oldest age/dead-letter signals; recovery command to retry or reconcile safely. |
| PostgreSQL | Startup unavailable | Process startup rejects | This is a useful fail-fast signal, but it is not joined to a deploy gate with bounded logs and rollback criteria. | Installer/supervisor waits for explicit readiness, captures failure reason, and aborts/rolls back without replacing last-known-good service. |
| PostgreSQL | Runtime outage, pool saturation, slow queries, migration incompatibility | Raw Prisma errors; request symptoms | No readiness failure, pool/saturation metric, query latency/timeout signal, migration compatibility assertion, or tenant/request/run correlation. Raw client internals can cross the HTTP boundary. | DB probe latency/result, Prisma error code/category counters, bounded query timing for critical operations, connection/pool saturation if available, schema migration identifier, sanitized API mapping, and DB outage/runbook drill. Prisma documents structured client error types/codes and configurable event-based logging; use those instead of returning raw exception strings (Prisma error reference, Prisma logging). |
| Agent run | SDK/model error, credentials invalid, provider throttling, malformed stream | Run may become FAILED; the card may show an error and trigger finalization may warn |
The installed SDK emits API retry, rate-limit, tool-progress, runtime-init/MCP-status, provider request ID, duration, and error-category messages, but runner.ts handles only stream/assistant/user/result messages and drops the other lifecycle signals. Provider/model/attempt/latency/normalized error are not consistently available as metrics or correlated logs. Some persistence and audit operations are best-effort. |
Run start/finish transition event with run_id, session_id, tenant/project, provider/model, status, duration, turns, tokens/cost, normalized error and retryability; capture the SDK's bounded retry/rate-limit/tool/runtime signals; aggregate run outcome/latency/stuck gauges without high-cardinality IDs as metric labels. |
| Agent run | No result/hang, missed heartbeats, process death, terminal DB write failure | Lock heartbeat may stop; restart sweep mutates some state | No wall-clock deadline/watchdog, stale-lock/stale-active-run metric, readiness impact, page, or documented reconciliation. Lifecycle fault injections proved active run without lock, terminal run with lock, and timeout misclassification are possible. | Deadline and durable lease semantics, gauges for active/waiting/stale runs and locks, oldest age, transition-failure counter, reconciliation command and drill. Startup recovery must be transactional and evidence-producing. See run lifecycle audit. |
cph subprocess |
Missing binary, permission failure, timeout, non-zero exit, malformed output | In production, cph_check/cph_build grant the Claude SDK's Bash tool; tool results are forwarded only to the in-memory streaming card. The direct agent/cph.ts wrapper captures exit/stdout/stderr but is used only by its integration tests, not by the production runner. |
There is no production-owned cph invocation boundary, startup/readiness validation of the exact deployed artifact, version/build identity, persisted operation result, duration metric, bounded/redacted stderr contract, or reliable distinction between spawn failure and domain diagnostics. |
Readiness checks executable path/version and required filesystem permissions without mutating domain state. Introduce an observable production invocation boundary or capture the SDK Bash/tool lifecycle with operation, run/project correlation, duration, exit category and bounded redacted diagnostic; metrics count timeout/spawn/non-zero/protocol errors. Deployment evidence pins and smoke-tests the artifact. |
| Workspace | Missing/unwritable root, wrong ownership, disk/inode exhaustion, symlink/root escape, partial agent writes | Error appears only when a command/file operation happens | No readiness check, free-space/inode monitoring, workspace mutation journal/commit marker, or recovery procedure. Security audit already proved containment failures; lifecycle audit proved a run can fail after workspace side effects. | Bounded non-destructive readiness check for configured root, owner/mode, read/write/rename/fsync as required, free space and inode signals, per-project mutation correlation, containment validation, backup/restore and partial-write reconciliation drill. See tenant/auth security audit and run lifecycle audit. |
| Deploy | Bad env/service user/artifact, failed migration/build, unhealthy new process | Shell/systemd exit in some cases; process may restart | Clean-host audit found missing required env/service identity/artifact and in-place, non-atomic release replacement. Health is liveness-only and there is no last-known-good rollback unit. | Immutable versioned release, preflight, migration compatibility gate, explicit readiness deadline, release/build/schema identity endpoint/log, automated rollback criteria, last-known-good artifact and tested rollback. See clean-host deployment audit. |
| Deploy | SIGTERM/restart during active work | systemd stops/restarts the process | No graceful drain/stop readiness, bounded shutdown, durable queue resume, or post-restart reconciliation proof. In-memory queues are lost. | On termination: readiness false, stop ingress, drain or checkpoint durable work, enforce timeout, emit shutdown summary. Recovery drill sends work across restart and verifies a single durable outcome and delivery. systemd supports explicit notification/readiness/watchdog protocols, but the current unit does not implement one (systemd service types, sd_notify). |
Logs and end-to-end traceability
Current gaps
- Fastify generates an internal request ID for its access logs, but it is not returned to the client as a support handle or propagated into application errors, Prisma diagnostics, Feishu work, runs, or subprocess calls.
- Feishu
event_id, provider request/message IDs,chat_id,project_id,organization_id,run_id,session_id, and subprocess invocation are not propagated as one causally connected context. - The DB-cut experiment produced raw Prisma diagnostics without an application error record that identified route, request, response class, or tenant context.
- Several critical failure paths catch and ignore persistence errors or only warn, including audit/message persistence and some card/lock finalization paths. This creates absence of evidence precisely when durable state is least trustworthy.
- The production runner discards SDK
api_retry,rate_limit_event,tool_progress, runtime initialization/MCP status, and related system messages even though the installed SDK exposes the fields needed to diagnose provider and tool degradation. - Exception strings can cross the API boundary. Besides leaking implementation details, this makes the client-visible error vocabulary unstable and impossible to alert on reliably.
- Logs have no documented retention, rotation, ingestion, access-control, redaction, or tenant-support lookup contract. Journald availability alone is not a log strategy.
Minimum structured event envelope
Every boundary and durable state transition should emit one structured event with a stable event_name, timestamp, severity, service, environment, release, and the applicable correlation fields:
request_id, feishu_event_id, organization_id, project_id, chat_id, actor_user_id, run_id, session_id, delivery_id, provider_request_id, and cph_invocation_id.
Raw prompt/message/file contents, credentials, cookies, authorization headers, provider tokens, full Prisma invocation text, and unrestricted stdout/stderr must not be logged. Tenant/project/user identifiers are operationally necessary but access-controlled personal/customer metadata; metric labels must not use high-cardinality IDs. OpenTelemetry's HTTP semantic conventions specify the stable http.server.request.duration instrument, an opt-in http.server.active_requests instrument, and low-cardinality route/error guidance (HTTP metrics semantic conventions).
At minimum, events are required for:
- request accepted/completed/failed/canceled;
- Feishu listener connected/disconnected/auth-failed/reconnecting;
- ingress received/deduplicated/queued/claimed/started/retried/dead-lettered;
- run and lock state transitions, transition-write failures, heartbeat stale detection, and recovery decisions;
- provider/agent start/result/error with normalized category;
cphspawn/result/timeout/protocol failure;- workspace preflight and mutation/commit/reconciliation decisions;
- outbound delivery queued/attempted/confirmed/retried/dead-lettered;
- deploy, schema migration, readiness transition, drain, shutdown, startup recovery, and rollback.
Errors must be allowed to fail the affected operation rather than disappear. If a non-critical telemetry sink is unavailable, that telemetry failure itself needs a bounded local signal/counter; durable business audit records are not optional telemetry and must follow their explicit transactional contract.
Health, readiness, and startup contract
The service needs three distinct concepts:
- Liveness: the process/event loop is responsive. This endpoint must be cheap and must not depend on remote services. The present
/api/healthzcan serve this purpose if it is named/documented as liveness only. - Readiness: the instance can accept the traffic assigned to it. A bounded probe must include at least PostgreSQL connectivity/schema compatibility, Feishu listener authenticated/connected state for ingress-serving instances, required runtime configuration, exact
cphartifact availability/version, and workspace root access/capacity. A failed dependency makes readiness non-2xx without killing the process and losing diagnostics. - Startup/deploy acceptance: migrations/preflight complete, listener reaches ready state within a deadline, one authenticated HTTP/database smoke path and one controlled runtime/workspace smoke path pass, and the deployed release/schema identity matches the intended version. Only then is traffic enabled and the previous release eligible for cleanup.
Readiness responses should expose only component names, stable state/error codes, and timestamps—not secrets or raw exceptions. Full diagnostic context belongs in correlated restricted logs. The deployment audit must decide whether Feishu and HTTP run in one instance role; if roles are split later, readiness dependencies become role-specific rather than globally optional.
Minimum metrics and alerts
Exact SLO values are product/operations decisions and are currently unspecified. The release nevertheless needs the following bounded-cardinality instruments and a reviewed initial threshold for each page/ticket alert.
Traffic and failures
- HTTP request count, duration, active requests, status/error-code class, timeouts/cancellations.
- Feishu listener connected state, reconnect/auth failures, time since last successful connection/event.
- Ingress queue/receipt counts by durable state, processing attempts, oldest pending age, dead-letter count.
- Outbound delivery attempts/results/latency, retry/dead-letter count, oldest pending age, provider rate-limit responses.
- Prisma operation failures by normalized code/category and bounded critical-query duration; DB readiness result/latency.
Workflow integrity and capacity
- Runs started and terminal outcomes, run duration, active/waiting/stale count and oldest age.
- Locks held/stale/orphaned and transition/reconciliation failures.
- Agent/provider failures by provider/model/error category; turns/tokens/cost as operational/business counters with tenant detail kept out of metric labels.
cphinvocation count/duration/result category/timeouts.- Workspace free bytes/inodes, preflight failures, write/commit/reconciliation failures.
- Process uptime/restarts, event-loop delay, memory, CPU, file descriptors, unhandled rejection/uncaught exception, graceful shutdown outcome.
- Release/build/schema identity as an info metric or startup event, never as an unbounded label set.
Required release-time alerts
Before production traffic, alerts and routing must exist for:
- instance readiness false beyond the rollout/startup grace period;
- Feishu auth failure or prolonged listener disconnect;
- sustained server-side HTTP errors or request timeouts;
- database readiness failure / critical Prisma failure spike;
- oldest ingress/outbound item, active run, or lock exceeds its decided deadline;
- dead-letter count non-zero or durable state-transition/reconciliation failure;
- agent/provider or
cphfailure/timeout spike; - workspace disk/inodes below headroom threshold or workspace preflight fails;
- restart loop, failed deployment, failed migration, or rollback;
- telemetry pipeline blind (expected heartbeat/scrape/log ingestion absent).
Each alert needs an owner, severity, user impact statement, dashboard/log query, first safe diagnostic command, mitigation, escalation criterion, and recovery verification. A page without a runbook is not a production control.
Incident recovery runbooks required before release
At minimum, checked-in runbooks must cover these scenarios. Commands must default to read-only inspection; destructive/redrive/force-release actions require explicit confirmation, identify tenant/project/run scope, emit an audit record, and be idempotent or document why they are not.
- PostgreSQL unavailable/degraded — identify connection vs credential vs migration vs capacity failure; stop accepting work through readiness; restore service; reconcile active runs, locks, ingress receipts, and outbox; verify no raw errors escape HTTP.
- Feishu listener disconnected/credentials revoked — distinguish auth from transient network/reconnect; restore credentials safely; verify last event and controlled inbound round trip; inspect ingress gap and redrive rules.
- Ingress backlog/dead letter — list oldest/attempts/error by tenant-safe scope; decide retry vs reject; prove dedup/idempotency; verify one resulting run.
- Stuck/orphaned run or lock — inspect run/session/lock/provider/workspace/audit/outbox together; use lease/reconciliation operation rather than direct row deletion; prove terminal state and next queued work behavior.
- Outbound delivery failure — reconcile provider message ID and durable outbox; retry with idempotency key; confirm user-visible delivery without duplicating a card/message.
- Agent/provider or
cphoutage — identify provider credential/rate-limit/runtime/protocol failure; halt or queue new runs; test a canary; resume and reconcile affected work. - Workspace capacity/corruption/containment event — quarantine affected project, snapshot evidence, restore/rollback documented unit, verify ownership/containment and repository state before resuming.
- Failed migration/deployment — preserve old release, collect preflight/migration/readiness diagnostics, follow compatibility-aware application/database rollback, and verify release/schema identity plus canaries.
- Process/host restart during work — drain when planned; after unplanned restart run durable reconciliation; verify receipts, runs, locks, workspace effects and deliveries converge exactly once at the business level.
- Suspected tenant data exposure — disable affected path/tenant access without deleting evidence, preserve correlated audit/log records, rotate exposed credentials, determine affected org/project/users, and follow an explicitly owned incident-notification decision process.
Minimum release evidence
Production readiness is not proved by a green unit suite alone. A release candidate needs an archived, attributable evidence bundle containing all of the following:
Build and supply identity
- immutable application and
cphartifact digests, source commit, build timestamp/toolchain, dependency lockfiles and vulnerability policy result; - deployed configuration manifest listing required keys and instance role, with secret values redacted;
- migration set/schema identity and a compatibility statement for both forward deploy and rollback;
- systemd unit and installer/preflight version used on a clean host.
Automated gates
- typecheck, complete unit/integration suite, migration-from-supported-version tests, and Lean
lake buildwhenspec/changes; - clean-host install test with the real packaged artifacts and declared service user/directories;
- negative startup tests for missing/invalid env, DB unavailable, invalid Feishu credentials, missing/wrong
cph, and unwritable/full workspace; - readiness tests proving each dependency failure independently makes readiness fail while liveness remains diagnostically available;
- API error-mapping tests proving unexpected dependency exceptions return sanitized 5xx/stable codes and create a correlated error log;
- lifecycle fault tests from the run-lifecycle audit, including process kill/restart, DB transition failures, duplicate ingress, ambiguous outbound completion, timeout, and concurrent trigger/session creation;
- tenant isolation/security regression suite, symlink/root containment tests, and dependency audit policy gate.
Staging drills against production-shaped dependencies
- authenticated Feishu inbound → durable receipt → one run → workspace effect → durable terminal state → one outbound confirmation, searchable end-to-end by a correlation ID;
- PostgreSQL socket cut/restoration while serving: readiness changes, alert fires, requests return sanitized server errors, state reconciles, alert clears;
- Feishu credential rejection/disconnect/reconnect: readiness/alert and controlled recovery are visible;
- SIGTERM and forced process kill during queued, active, and outbound-pending work: restart reconciliation proves no silent loss and no unintended duplicate user effect;
cphtimeout/non-zero/malformed response, provider throttle/error, workspace disk/permission failure, and outbound rate-limit/permanent failure drills;- failed deploy/readiness and rollback drill on a clean host, followed by migration/data compatibility and canary verification;
- backup restore of representative database and workspace data, with measured recovery point/time and tenant isolation checks.
Operations acceptance
- dashboards expose all release metrics above and show release/schema identity;
- every required alert is injected once in staging, reaches the named owner, links to its checked-in runbook, and clears only after verified recovery;
- log search can pivot from request/event to org/project/run/session/delivery without message content or secrets; retention and access controls are documented and tested;
- on-call ownership, escalation, incident severity, maintenance/rollback authority, backup ownership, and recovery objectives are explicitly decided;
- a signed release checklist records who reviewed each piece of evidence and where immutable test/drill outputs are stored.
Prioritized remediation / release gates
P0 — before any external production traffic
- Implement distinct bounded liveness, readiness, and startup/deploy acceptance contracts; make deployment traffic activation depend on readiness.
- Install centralized unexpected-error handling that returns sanitized stable 5xx responses and emits request-correlated structured errors. Never return raw Prisma/SDK/process exception text.
- Make Feishu connection/auth state observable and readiness-relevant; make listener startup rejection fail the startup gate.
- Implement the durable ingress, transactional run/lock/session, durable outbox, workspace commit/reconciliation, and durable audit/history work already opened by the lifecycle audit. Observability cannot compensate for acknowledged work that has no recoverable state.
- Add end-to-end correlation and stable structured transition events; remove silent catches from critical paths.
- Add the minimum metrics, alerts, dashboards, and checked-in recovery runbooks, then execute the fault/restart/rollback drills.
- Replace in-place deployment with the clean-host audit's immutable release/readiness/rollback contract and prove it on a clean host.
- Close the critical tenant/agent/workspace/Feishu authorization and dependency findings from the security audit before declaring readiness.
P1 — required for routine operations immediately after the P0 architecture is in place
- Define initial SLOs/error budgets and tune alert thresholds using staging/load evidence.
- Add capacity/load/soak tests for DB connections, event-loop/process resources, ingress/outbound backlog, agent concurrency, provider limits, workspace disk/inodes, and log/metric pipelines.
- Exercise backup/restore, credential rotation, tenant-support trace lookup, and incident evidence preservation.
- Define data retention/deletion/access-control policy for logs, metrics, traces, audits, messages, receipts, deliveries, runs, and workspaces.
Decision points that remain OPEN
The code/spec does not currently authorize choosing these product/operations semantics implicitly:
- readiness dependency roles (combined HTTP+Feishu process versus independently deployable roles);
- maximum acceptable ingress, run, lock, and outbound ages; retry budgets; terminal/dead-letter semantics;
- workspace mutation commit/rollback unit and what a user should observe after partial agent work;
- audit durability/availability trade-off and exact event taxonomy;
- SLOs, recovery point objective, recovery time objective, retention, regional/data-residency constraints, and incident-notification ownership;
- whether deploy migration compatibility must support application rollback across one or more releases.
These decisions should be recorded in the upstream contract/ADR layer before implementation when they affect domain-visible behavior. Operational mechanisms that only expose already-defined behavior can be implemented and tested without inventing domain semantics.
Source notes
Primary local evidence:
hub/src/server.ts,hub/src/admin/errors.ts— process startup, health/listener/restart behavior, and HTTP exception-to-response mapping.hub/src/db.ts— current Prisma event/log configuration.hub/src/feishu/client.ts,hub/src/feishu/trigger.ts,hub/src/feishu/messageBatcher.ts,hub/src/feishu/triggerQueue.ts— ingress/listener, batching, dedup, queue, run and outbound behavior.hub/src/agent/runner.ts,hub/src/agent/roleTools.ts,hub/src/agent/cph.ts,hub/src/lock.ts,hub/src/audit.ts— production SDK/Bash tool routing, the legacy direct cph wrapper, persistence, heartbeat/lock, and audit behavior.hub/prisma/schema.prisma— durable state and available correlation fields.hub/deploy— installer/systemd/deploy behavior.- Clean-host deployment audit, tenant/auth security audit, and run-lifecycle integrity audit — prior deterministic audits and their recorded commands/results.
Official references used only for interface semantics and implementation guidance:
- Prisma Client error reference and Prisma logging.
- OpenTelemetry HTTP metric semantic conventions.
- systemd service types/readiness and
sd_notify.
No secondary production-readiness checklist was treated as authority. Recommendations are derived from the local failure evidence and the primary interfaces the service actually uses.