Files
curriculum-project-hub/.scratch/saas-production-readiness/issues/31-write-and-drill-incident-runbooks.md
T

19 lines
919 B
Markdown

# Write and drill production incident recovery runbooks
Type: task
Status: open
Blocked by: 06, 12, 13, 14, 23, 24, 25, 26, 27, 28, 29, 30, 39, 40, 41
## Question
Check in operator runbooks for PostgreSQL outage/migration failure, Feishu
disconnect and credential rejection, ingress/outbound backlog or dead letter,
stuck run/lock, provider or cph failure, workspace capacity/corruption,
deploy/rollback, host restart during work, and suspected tenant exposure. Give
each alert an owner and safe read-first commands, require confirmed/audited
scope for destructive or replay actions, execute every runbook against
production-shaped staging, and archive evidence that detection, mitigation,
reconciliation, verification, escalation, and alert clearing all work. Include
database/workspace backup failure, failed migration, full-host restore and
cutover, duplicate/gap review, fallback, and achieved RPO/RTO evidence.