forked from EduCraft/curriculum-project-hub
23 lines
1.2 KiB
Markdown
23 lines
1.2 KiB
Markdown
# Write and drill production incident recovery runbooks
|
|
|
|
Type: task
|
|
Status: open
|
|
Blocked by: 06, 12, 13, 14, 23, 24, 25, 26, 27, 28, 29, 30, 39, 40, 41
|
|
|
|
## Question
|
|
|
|
Check in operator runbooks for PostgreSQL outage/migration failure, Feishu
|
|
disconnect and credential rejection, ingress/outbound backlog or dead letter,
|
|
stuck run/lock, provider or cph failure, workspace capacity/corruption,
|
|
deploy/rollback, host restart during work, and suspected tenant exposure. Give
|
|
each alert an owner and safe read-first commands, require confirmed/audited
|
|
scope for destructive or replay actions, execute every runbook against
|
|
production-shaped staging, and archive evidence that detection, mitigation,
|
|
reconciliation, verification, escalation, and alert clearing all work. Include
|
|
database/workspace backup failure, failed migration, full-host restore and
|
|
cutover, duplicate/gap review, fallback, and achieved RPO/RTO evidence.
|
|
Also drill ADR-0023 platform-administrator lockout and Platform-owned Feishu
|
|
Application credential loss through the dual-factor offline recovery command,
|
|
expiring Emergency Platform Grant, session revocation, immutable Platform Audit,
|
|
and explicit recovery closeout; never introduce a standing break-glass account.
|