# Run lifecycle and data-integrity audit ## Verdict The current Hub does **not** have a crash-safe or database-failure-safe run lifecycle. The normal happy path is covered by tests and the database correctly enforces one lock per project, but the surrounding workflow is a chain of independent effects: inbound receipt, in-memory batching/queueing, session/run creation, lock acquisition, workspace mutation, Feishu delivery, session cursor update, run terminal update, audit write, and lock release. There is no durable coordinator or reconciliation loop across those effects. Deterministic fault injection confirmed that the current code can permanently lose accepted work, leave a terminal run holding a lock, leave an `ACTIVE` run without a lock, split one logical session into duplicate active rows, declare a run complete without delivering its result, and create an audit entry referring to no run. A hung provider call also has no wall-clock deadline or usable lease, so one run can block a project until an operator restarts the service or edits the database. These are production blockers for the accepted single-host topology; they do not require HA to reproduce. ## Contract baseline The audit used the following upstream decisions and intentionally did not invent semantics where the contract is open: - `Spec.System.Run`: `active` and `waitingForUser` are non-terminal; `completed`, `failed`, `timedOut`, and `canceled` are terminal. - `Spec.System.Lock` and ADR-0002: each project has at most one lock, its owner is an `AgentRun`, every lock holder is non-terminal, and the lock is released when the run completes, fails, times out, or is canceled. - ADR-0002: many runs reuse the same long-lived `AgentSession`. - ADR-0017: the active session identity is project + provider + role + model, and its provider cursor is persisted for the next run. - `Spec.System.Audit`: every `AuditEntry` is associated with a run. Audit event content, retention, and query dimensions remain `OPEN`. - The contract does **not** decide whether filesystem changes made by a failed, canceled, timed-out, or process-killed run should commit, remain as partial work, or roll back. That choice is surfaced below rather than assumed. ## Positive evidence The core happy path is real: - PostgreSQL uniqueness enforces at most one `ProjectAgentLock` per project and at most one project lock per run. - `currentLockRunId` rejects a terminal run that still holds a lock instead of silently treating it as healthy. - Normal trigger completion updates the provider cursor, writes a terminal run state, and releases the lock. Interrupt uses an `AbortController` and maps to `CANCELED` on the normal path. - A sequential Feishu redelivery with the same event ID is skipped, so the ordinary duplicate case does not start a second run. - The targeted lifecycle baseline passed before fault injection: ```text npm run check 6 test files passed 71 tests passed ``` Those protections are necessary, but they only cover operations in which every dependency succeeds in the expected order. ## Executed fault-injection evidence A temporary Vitest harness exercised the real trigger, Prisma test database, message batcher, and runner seams. It asserted the production invariant rather than the current behavior. The harness was run repeatedly and produced the same result each time; the final run was: ```text npm test -- --run test/run-lifecycle-audit.repro.test.ts Test Files 1 failed (1) Tests 8 failed (8) Duration 918ms ``` The throwaway harness was removed after diagnosis. Its eight minimal probes and observed values were: | Injected condition | Required invariant | Observed current result | | --- | --- | --- | | Authorization database error after Feishu receipt creation, followed by the same event redelivery | The failed event is retried | Authorizer calls remained `1`; redelivery was skipped | | Batched-run setup callback throws, then `flushAll` retries | The pending batch remains recoverable | Callback attempts remained `1`; batch had already been deleted | | Two active sessions with the same project/provider/role/model tuple | Database rejects the duplicate | Second insert succeeded | | Audit insert uses a nonexistent run ID | Database enforces `AuditEntry → AgentRun` | Insert succeeded with arbitrary `runId` | | Every Feishu message create/reply fails | Completion is not reported as delivered | Run became `COMPLETED`; delivered cards were `0` | | First lock release delete fails | Terminal run is eventually unlocked | Lock count remained `1` after run became `COMPLETED` | | Every terminal `AgentRun.update` fails | A non-terminal run retains/reacquires its lock | Run remained `ACTIVE`; lock count became `0` | | Claude SDK returns `error_max_turns` | Runner returns `length`, trigger stores `TIMED_OUT` | Runner returned `failed` | This is a tight reproduction at the actual call seams; none of the failures depends on timing or external network access. ## Critical findings ### 1. Accepted Feishu work can be permanently lost `trigger.ts` commits `FeishuEventReceipt` before message validation, binding lookup, authorization, file download, command execution, batching, or run creation. The receipt has no processing state, lease, attempt count, error, or completion marker. Any later exception leaves a permanent row, so Feishu's at-least-once redelivery is treated as already completed. The two in-process buffers compound that behavior: - `MessageBatcher.flushPendingBatch` deletes the batch before invoking its callback. Timer-triggered callback failures are swallowed with `.catch(() => undefined)`. - `TriggerQueue` and `batchContexts` are process-local maps. A restart loses all queued prompts and pending debounce batches. Queue draining removes an item before run setup is durable, so a setup error also loses it. There is no shutdown hook that flushes or rejects these buffers visibly. The service can therefore acknowledge and forget user work even when PostgreSQL is healthy again milliseconds later. ### 2. Run state and lock ownership can contradict each other Run creation, lock acquisition, terminal update, and lock release are separate Prisma calls. Their failure handling is asymmetric: - a lock-acquisition race creates a session and an `AgentRun` first, then marks the losing run failed; - terminal update failure is caught, a second best-effort failure update is attempted, and `finally` releases the lock regardless; - lock release failure is only logged once, with no retry or reconciler. Fault injection reached both forbidden outcomes: `COMPLETED + lock` and `ACTIVE + no lock`. The first makes `currentLockRunId` throw the explicit ADR-0002 invariant error on all later triggers; the second permits another run to start while the database still claims the earlier run is live. In the injected terminal-write case the earlier execution had already returned, but the same unlocked/non-terminal shape created by an overlapping startup can coexist with a still-running process. Startup does not repair these states safely. `server.ts` first deletes **all** locks and then, in a separate statement, marks only `ACTIVE` runs failed. `WAITING_FOR_USER` is left non-terminal and unlocked, contrary to the contract. The two statements are not transactional, and running a second Hub process would destructively clear the live first process's locks and fail its runs. ### 3. A provider hang can block a project indefinitely `maxTurns` is passed to the SDK, but there is no wall-clock run deadline. The SDK max-turn result is currently mapped to runner status `failed`, making the trigger's `length → TIMED_OUT` branch unreachable for the actual max-turn case. `ProjectAgentLock.expiresAt` is never populated or consumed. `heartbeatAt` is a best-effort update only after an assistant message, not a periodic heartbeat, and there is no stale-lock/run sweeper. A query that stalls before its first assistant message remains `ACTIVE` while holding the project lock forever. `WAITING_FOR_USER` is declared but never written; approval waits remain `ACTIVE` in an in-memory `ApprovalManager`. Systemd sends `SIGTERM`, but the application has no signal handler, listener close, in-flight drain, run cancellation, Prisma disconnect, or final recovery transaction. A subsequent successful restart happens to fail `ACTIVE` runs, but that is not a graceful lifecycle and does not cover all non-terminal states. ### 4. Workspace effects have no defined commit or recovery boundary The Claude SDK's built-in Write/Bash tools mutate the project workspace while the run is still `ACTIVE`. A process kill, timeout, cancellation, database failure, or later delivery failure can therefore leave real filesystem changes behind while the database says the run failed or never completed. No runtime code writes `AgentFileChange`: the only capture callback is in the legacy AI-SDK `writeFileTool`, while the production runner uses Claude SDK built-ins and never injects that collector. The schema comment claiming these changes are captured for review is not true. The transcript comments are also stale: the production runner neither loads nor appends the JSONL transcript; continuity comes from the provider cursor. `AgentMessage` writes are best-effort and silently swallowed. The session cursor, message projection, workspace state, run terminal state, and audit history can consequently describe different versions of the same execution. Before implementing a mechanism, the product/contract must decide whether partial workspace changes are committed, quarantined, or rolled back. ## High findings ### 5. Session identity and cursor continuity are not concurrency-safe `startAgentRun` performs `findFirst` then `create` for the active project/provider/role/model session before acquiring the project lock. The schema has only a non-unique index for that tuple. The database probe inserted two active rows with the same identity. Concurrent triggers can therefore fork a logical conversation before one loses the project-lock race. Future `findFirst` selection is unspecified and may resume the wrong provider cursor. On the completion path, the session cursor is updated separately and before the run terminal update, so a partial database failure can advance the conversation while leaving the originating run active or failed. ### 6. Outbound delivery is silent and not recoverable The Feishu send helpers catch message errors and return `null`; `sendText` returns `void`. `StreamingAgentCard.finish` treats a failed create/patch as a boolean and normal run completion proceeds. The fault probe recorded a `COMPLETED` run with zero delivered cards. There is no durable outbox, provider message ID record, delivery state, idempotency key, retry schedule, or reconciliation job. An operator cannot tell whether the user saw the result, and retrying a larger workflow later would have no durable key with which to prevent duplicate messages. ### 7. Audit rows violate the pinned relation and audit failures are invisible `AuditEntry.runId` is nullable, has no foreign key, and accepts arbitrary strings. This directly diverges from the pinned `Spec.System.Audit` relation that every audit entry belongs to a run. Pre-run permission/security events are valid product needs, but they require an explicit event model rather than weakening the one relation the contract fixes. `writeAudit` catches every Prisma error without logging, metric, retry, or caller signal, and existing unit tests encode that swallowing as desired. Lifecycle history can be absent exactly when PostgreSQL is unhealthy, with no evidence that evidence was lost. The content and failure policy are `OPEN` and need a decision, but silent unobservable loss is incompatible with operating and debugging the service. ## Failure-mode summary - **Process termination/restart:** loses batches, queued prompts, approvals, active controller handles, and unrecorded message/file history; may leave partial workspace changes; recovery ignores `WAITING_FOR_USER`. - **Timeout/hang:** no wall deadline, periodic heartbeat, lease, or sweeper; max-turn termination is mislabeled `FAILED` rather than `TIMED_OUT`. - **Duplicate Feishu delivery:** sequential success is deduplicated, but any failure after receipt becomes permanent loss; receipt check/create is also not one transactional claim operation. - **Database failure:** can produce either terminal-with-lock or active-without-lock, lose session/message/audit facts, or advance only the provider cursor. - **Concurrent requests/processes:** the project-lock unique key prevents two normal lock rows, but active sessions can fork; a second process's startup reset can invalidate the first process's live lifecycle. - **Outbound failure:** run completion and user-visible delivery are unrelated, and no durable recovery/idempotency record exists. ## Required release evidence Production readiness requires automated proof of all of the following: 1. Each accepted inbound event has a durable state machine and can be retried after process/DB failure without duplicate run creation. 2. Queueing and batching either survive restart or have an explicit, observable rejection/acknowledgement contract that cannot silently lose work. 3. Every database fault point preserves `lock exists ⇒ holder non-terminal` and prevents a live non-terminal run from continuing without its lock. 4. A wall deadline, lease/heartbeat, watchdog, and startup reconciler converge every non-terminal state, including human approval waits. 5. One active session identity is enforced in PostgreSQL and cursor advancement is serialized with the run transition that produced it. 6. Feishu delivery uses a durable, idempotent outbox/reconciliation path and exposes delivered, pending, and failed delivery states separately from model completion. 7. The decided failed-run workspace policy is implemented and tested across kill, cancel, timeout, and database failure; file/message history reflects the real effects or explicitly reports missing evidence. 8. The audit/event model conforms to the pinned run relation, records pre-run security events in an explicitly modeled place, and makes write loss observable and recoverable according to the accepted policy. 9. Graceful shutdown and two-process/startup-race tests prove that one process cannot clear another live process's locks or accept new work while draining.