Disaster Recovery Playbook
Техническое руководство из исходного проекта Notty. Примеры, параметры и эксплуатационные ограничения.
This playbook documents the current recovery flow for the operational features added in Epic 9.
It is intentionally aligned with the existing backend capabilities in packages/server/src/lib/backup.ts,
packages/server/src/lib/operational-alerts.ts, packages/server/src/lib/jobs/, and packages/server/src/lib/webhooks.ts.
Scope
This playbook covers:
- backup creation and retention;
- restore from a completed backup snapshot;
- handling failed scheduled jobs;
- handling failed webhook deliveries and dead-letter backlog;
- handling security and operational alerts;
- post-recovery validation.
Recovery Objectives
- RPO: depends on the active backup schedule configured in
/api/admin/backups/schedule. - RTO: depends on database size, media volume, and whether restore is run as dry-run or full restore.
Operators should record the actual recovery window used during an incident review.
Source of Truth
- Backup catalog:
GET /api/admin/backups - Backup schedule:
GET /api/admin/backups/schedule - Restore endpoint:
POST /api/admin/backups/:id/restore - Jobs overview:
GET /api/admin/jobs,GET /api/admin/jobs/stats - Scheduled jobs:
GET /api/admin/scheduled-jobs - Operational alerts:
GET /api/admin/operational-alerts,GET /api/admin/operational-alerts/stats - Webhook health and logs:
GET /api/webhooksGET /api/webhooks/:id/logsGET /api/webhooks/dead-letterPOST /api/webhooks/:id/replay/:logIdPUT /api/webhooks/:id/health
Backup Strategy
Manual backup
Create an immediate snapshot before risky maintenance or schema/data migration:
curl -X POST "$BASE_URL/api/admin/backups" \
-H "Authorization: Bearer $ADMIN_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "pre-maintenance-backup",
"retentionDays": 30,
"async": false
}'
Scheduled backup
Use the schedule config endpoint to keep automated snapshots enabled:
curl -X PUT "$BASE_URL/api/admin/backups/schedule" \
-H "Authorization: Bearer $ADMIN_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"enabled": true,
"cronExpression": "0 2 * * *",
"retentionDays": 30,
"sections": {
"schemas": true,
"components": true,
"config": true,
"content": true,
"media": true
}
}'
Retention
- Default retention is 30 days.
- Expired backups are cleaned up by the built-in
backup.cleanupschedule. - Backup schedule config is stored internally and must not be treated as a real backup snapshot.
Restore Procedure
1. Stabilize the system
Before restore:
- stop write traffic if possible;
- pause or disable external automations that would enqueue more work;
- note active alerts, failed jobs, and dead-letter webhook entries;
- choose the latest known-good completed backup from
GET /api/admin/backups.
2. Run dry-run if only partial verification is needed
Use dryRun when validating the selected backup and restore sections:
curl -X POST "$BASE_URL/api/admin/backups/42/restore" \
-H "Authorization: Bearer $ADMIN_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"dryRun": true,
"sections": {
"schemas": true,
"components": true,
"config": true,
"content": true,
"media": true
}
}'
3. Execute full restore
When the selected snapshot is confirmed:
curl -X POST "$BASE_URL/api/admin/backups/42/restore" \
-H "Authorization: Bearer $ADMIN_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"dryRun": false
}'
Expected result:
- restore returns
success: true; - audit log contains
backup.restore; - operational alerts can be acknowledged/resolved after validation;
- the backup remains in the catalog for later forensics.
Incident Playbooks
Failed scheduled jobs
- Inspect current state in
GET /api/admin/jobsandGET /api/admin/jobs/stats. - Check whether the failure is transient, configuration-driven, or data-driven.
- Retry safe jobs with
POST /api/admin/jobs/:id/retry. - Cancel poisoned or obsolete jobs with
POST /api/admin/jobs/:id/cancel. - If a scheduler definition is wrong, fix it via
/api/admin/scheduled-jobs/:id. - Acknowledge or resolve linked operational alerts after recovery.
Failed webhook deliveries
- Inspect webhook health and logs in
GET /api/webhooks/:id/logs. - Review the dead-letter queue in
GET /api/webhooks/dead-letter. - Fix destination availability, signature mismatch, or payload contract issues.
- Replay recoverable deliveries with
POST /api/webhooks/:id/replay/:logId. - Reset webhook health counters with
PUT /api/webhooks/:id/healthwhen the destination is healthy again. - Resolve the related operational alert.
Backup failures
- Check
GET /api/admin/operational-alertsforbackup_failure. - Verify backup schedule config and database/storage availability.
- Trigger a manual backup to confirm recovery.
- Keep the alert open until a completed snapshot is created successfully.
Security events
- Review security-related alerts and recent audit log entries.
- Contain the source: rotate tokens, revoke compromised credentials, or block offending IPs upstream.
- Confirm the platform is stable before replaying jobs or webhooks.
- Record incident details in the postmortem.
Post-Restore Validation
After restore or incident recovery, verify:
GET /api/admin/backupsreturns the expected snapshot metadata;GET /api/admin/jobs/statsshows no growing failed backlog;GET /api/webhooks/dead-letteris empty or understood;GET /api/admin/operational-alerts/statsshows alerts moving back to acknowledged/resolved;- critical content types can be read through the content API;
- scheduled publishing and backup schedules are still present;
- preview invalidation and webhook delivery resume normally.
Rollback Guidance
If a restore introduces a worse state:
- stop writes again;
- identify the previous known-good backup;
- repeat the restore flow with that backup;
- compare audit log, alert state, and job/webhook backlog before reopening traffic.
Postmortem Checklist
- What failed?
- Which backup was used?
- What was the real RPO/RTO?
- Were jobs or webhooks replayed?
- Were any alerts missing or too noisy?
- Does the schedule or retention policy need adjustment?
Источник: docs/disaster-recovery-playbook.md. Снимок документации исходного проекта. Технический справочник сохраняет язык оригинала.