# Define production observability and incident recovery Type: research Status: open Blocked by: 01, 02, 03 ## Question Which logs, metrics, readiness signals, audit correlations, alerts, and operator recovery actions are missing for a production operator to detect and diagnose failures across HTTP, Feishu, PostgreSQL, agent execution, cph builds, and workspace storage?