# Write and drill production incident recovery runbooks Type: task Status: open Blocked by: 06, 12, 13, 14, 23, 24, 25, 26, 27, 28, 29, 30, 39, 40, 41 ## Question Check in operator runbooks for PostgreSQL outage/migration failure, Feishu disconnect and credential rejection, ingress/outbound backlog or dead letter, stuck run/lock, provider or cph failure, workspace capacity/corruption, deploy/rollback, host restart during work, and suspected tenant exposure. Give each alert an owner and safe read-first commands, require confirmed/audited scope for destructive or replay actions, execute every runbook against production-shaped staging, and archive evidence that detection, mitigation, reconciliation, verification, escalation, and alert clearing all work. Include database/workspace backup failure, failed migration, full-host restore and cutover, duplicate/gap review, fallback, and achieved RPO/RTO evidence. Also drill ADR-0023 platform-administrator lockout and Platform-owned Feishu Application credential loss through the dual-factor offline recovery command, expiring Emergency Platform Grant, session revocation, immutable Platform Audit, and explicit recovery closeout; never introduce a standing break-glass account.