Incident Response & Disaster Recovery
How incidents are detected and handled, who escalates, and how the service recovers.
- Severity — five standard operational levels.
- Detection — a 5-minute polling cycle (monitoring, heartbeat, ErrorLogs).
- SLA — availability 99.95%–99.9999% by agreement; target resolution within 4 hours.
- Roles — AI automates detection/triage/runbook execution; a pre-approval gate covers high-risk actions; the human FDE Squad handles new or high-risk cases.
- Escalation — a Slack channel with a per-customer named contact, 24/7 coverage, escalating operator → manager → end user.
- Post-incident — follows the customer’s standard incident report; if none exists, an md report is provided.
- Disaster recovery — RTO 4h; RPO by agreement; DB and source backed up daily (cloud storage or customer location); an FDE Box failure is fixed within 30 min or the box is swapped; the GIIP Engine runs Azure region-redundant; DR is tested every 6 months.