Files
curriculum-project-hub/.scratch/saas-production-readiness/issues/31-write-and-drill-incident-runbooks.md
T

23 lines
1.2 KiB
Markdown

# Write and drill production incident recovery runbooks
Type: task
Status: open
Blocked by: 06, 12, 13, 14, 23, 24, 25, 26, 27, 28, 29, 30, 39, 40, 41
## Question
Check in operator runbooks for PostgreSQL outage/migration failure, Feishu
disconnect and credential rejection, ingress/outbound backlog or dead letter,
stuck run/lock, provider or cph failure, workspace capacity/corruption,
deploy/rollback, host restart during work, and suspected tenant exposure. Give
each alert an owner and safe read-first commands, require confirmed/audited
scope for destructive or replay actions, execute every runbook against
production-shaped staging, and archive evidence that detection, mitigation,
reconciliation, verification, escalation, and alert clearing all work. Include
database/workspace backup failure, failed migration, full-host restore and
cutover, duplicate/gap review, fallback, and achieved RPO/RTO evidence.
Also drill ADR-0023 platform-administrator lockout and Platform-owned Feishu
Application credential loss through the dual-factor offline recovery command,
expiring Emergency Platform Grant, session revocation, immutable Platform Audit,
and explicit recovery closeout; never introduce a standing break-glass account.