Разделы документации
ОбзорБыстрый стартРедакции и возможностиМодели и поляРедактор контентаМедиатекаЛокализацияПубликация и работа командыAPI, SDK и генерация типовРасширения и инструментыРабочие проектыЗадания, вебхуки и наблюдаемостьАудит и управление даннымиСоветники и доверие к плагинамКорпоративный входПространства, квоты и масштабированиеCommerce и PortalПрава и безопасностьРазвёртывание и обновленияЛицензии и установка пакетовТекущие ограниченияПомощь и диагностикаДанные в кабинетеCore CMSDeveloper PlatformProduction UseWorkflowOperationsComplianceAI AssistantsPlugin TrustEnterprise IdentityEnterprise ScaleEnterprise DeploymentCommerce BundlePortal BundleNotty CMS DocumentationAuth & SecurityContent ModelingDeploymentEcosystem & Packaging ConventionsEditions and First-party ModulesExtensibilityGetting StartedMedia ManagementModule Extraction PathDraft & PublishUpgrade GuideWebhooks & IntegrationsCookbook: Blog with Next.jsCookbook: Custom PluginCookbook: Multilingual SiteOperations DocsBackup AutomationDeployment BlueprintsRunbook — Восстановление БД из бэкапаRunbook — Плановый деплойRunbook — Реакция на инцидентRunbook — Откат релизаRunbook — Горизонтальное масштабированиеRunbook — Ротация секретовRunbook — Major upgradeSecrets ManagementNotty CMS — Capability MapNotty Configuration ModelGenerated App ContractComponents and Dynamic ZonesMiddleware SystemPerformance & Scaling ToolkitDisaster Recovery PlaybookDistribution Model
Документация / Технический справочник

Disaster Recovery Playbook

Техническое руководство из исходного проекта Notty. Примеры, параметры и эксплуатационные ограничения.

По функциям модуляОбновлено 2026-09-30

This playbook documents the current recovery flow for the operational features added in Epic 9. It is intentionally aligned with the existing backend capabilities in packages/server/src/lib/backup.ts, packages/server/src/lib/operational-alerts.ts, packages/server/src/lib/jobs/, and packages/server/src/lib/webhooks.ts.

Scope

This playbook covers:

  • backup creation and retention;
  • restore from a completed backup snapshot;
  • handling failed scheduled jobs;
  • handling failed webhook deliveries and dead-letter backlog;
  • handling security and operational alerts;
  • post-recovery validation.

Recovery Objectives

  • RPO: depends on the active backup schedule configured in /api/admin/backups/schedule.
  • RTO: depends on database size, media volume, and whether restore is run as dry-run or full restore.

Operators should record the actual recovery window used during an incident review.

Source of Truth

  • Backup catalog: GET /api/admin/backups
  • Backup schedule: GET /api/admin/backups/schedule
  • Restore endpoint: POST /api/admin/backups/:id/restore
  • Jobs overview: GET /api/admin/jobs, GET /api/admin/jobs/stats
  • Scheduled jobs: GET /api/admin/scheduled-jobs
  • Operational alerts: GET /api/admin/operational-alerts, GET /api/admin/operational-alerts/stats
  • Webhook health and logs:
    • GET /api/webhooks
    • GET /api/webhooks/:id/logs
    • GET /api/webhooks/dead-letter
    • POST /api/webhooks/:id/replay/:logId
    • PUT /api/webhooks/:id/health

Backup Strategy

Manual backup

Create an immediate snapshot before risky maintenance or schema/data migration:

curl -X POST "$BASE_URL/api/admin/backups" \
  -H "Authorization: Bearer $ADMIN_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "pre-maintenance-backup",
    "retentionDays": 30,
    "async": false
  }'

Scheduled backup

Use the schedule config endpoint to keep automated snapshots enabled:

curl -X PUT "$BASE_URL/api/admin/backups/schedule" \
  -H "Authorization: Bearer $ADMIN_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "enabled": true,
    "cronExpression": "0 2 * * *",
    "retentionDays": 30,
    "sections": {
      "schemas": true,
      "components": true,
      "config": true,
      "content": true,
      "media": true
    }
  }'

Retention

  • Default retention is 30 days.
  • Expired backups are cleaned up by the built-in backup.cleanup schedule.
  • Backup schedule config is stored internally and must not be treated as a real backup snapshot.

Restore Procedure

1. Stabilize the system

Before restore:

  • stop write traffic if possible;
  • pause or disable external automations that would enqueue more work;
  • note active alerts, failed jobs, and dead-letter webhook entries;
  • choose the latest known-good completed backup from GET /api/admin/backups.

2. Run dry-run if only partial verification is needed

Use dryRun when validating the selected backup and restore sections:

curl -X POST "$BASE_URL/api/admin/backups/42/restore" \
  -H "Authorization: Bearer $ADMIN_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "dryRun": true,
    "sections": {
      "schemas": true,
      "components": true,
      "config": true,
      "content": true,
      "media": true
    }
  }'

3. Execute full restore

When the selected snapshot is confirmed:

curl -X POST "$BASE_URL/api/admin/backups/42/restore" \
  -H "Authorization: Bearer $ADMIN_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "dryRun": false
  }'

Expected result:

  • restore returns success: true;
  • audit log contains backup.restore;
  • operational alerts can be acknowledged/resolved after validation;
  • the backup remains in the catalog for later forensics.

Incident Playbooks

Failed scheduled jobs

  1. Inspect current state in GET /api/admin/jobs and GET /api/admin/jobs/stats.
  2. Check whether the failure is transient, configuration-driven, or data-driven.
  3. Retry safe jobs with POST /api/admin/jobs/:id/retry.
  4. Cancel poisoned or obsolete jobs with POST /api/admin/jobs/:id/cancel.
  5. If a scheduler definition is wrong, fix it via /api/admin/scheduled-jobs/:id.
  6. Acknowledge or resolve linked operational alerts after recovery.

Failed webhook deliveries

  1. Inspect webhook health and logs in GET /api/webhooks/:id/logs.
  2. Review the dead-letter queue in GET /api/webhooks/dead-letter.
  3. Fix destination availability, signature mismatch, or payload contract issues.
  4. Replay recoverable deliveries with POST /api/webhooks/:id/replay/:logId.
  5. Reset webhook health counters with PUT /api/webhooks/:id/health when the destination is healthy again.
  6. Resolve the related operational alert.

Backup failures

  1. Check GET /api/admin/operational-alerts for backup_failure.
  2. Verify backup schedule config and database/storage availability.
  3. Trigger a manual backup to confirm recovery.
  4. Keep the alert open until a completed snapshot is created successfully.

Security events

  1. Review security-related alerts and recent audit log entries.
  2. Contain the source: rotate tokens, revoke compromised credentials, or block offending IPs upstream.
  3. Confirm the platform is stable before replaying jobs or webhooks.
  4. Record incident details in the postmortem.

Post-Restore Validation

After restore or incident recovery, verify:

  • GET /api/admin/backups returns the expected snapshot metadata;
  • GET /api/admin/jobs/stats shows no growing failed backlog;
  • GET /api/webhooks/dead-letter is empty or understood;
  • GET /api/admin/operational-alerts/stats shows alerts moving back to acknowledged/resolved;
  • critical content types can be read through the content API;
  • scheduled publishing and backup schedules are still present;
  • preview invalidation and webhook delivery resume normally.

Rollback Guidance

If a restore introduces a worse state:

  1. stop writes again;
  2. identify the previous known-good backup;
  3. repeat the restore flow with that backup;
  4. compare audit log, alert state, and job/webhook backlog before reopening traffic.

Postmortem Checklist

  • What failed?
  • Which backup was used?
  • What was the real RPO/RTO?
  • Were jobs or webhooks replayed?
  • Were any alerts missing or too noisy?
  • Does the schedule or retention policy need adjustment?

Источник: docs/disaster-recovery-playbook.md. Снимок документации исходного проекта. Технический справочник сохраняет язык оригинала.