Business Continuity and Disaster Recovery Policy template (SOC 2)

Sets recovery objectives by system tier, defines how a disaster is declared and recovered from, and requires the plan to be tested and maintained.

Policy P14 of 22 · Owner: Engineering Lead · 5 criteria in the crosswalk · Apache-2.0

Free companion template: Business continuity plan template (Word), the plan this policy calls for: RTO and RPO per service, scenario playbooks and a restore test log.

What this policy is for

The Business Continuity and Disaster Recovery Policy is one of the 22 governance policies in the Policyseed set. It is owned by the Engineering Lead, approved by the executive you name in the intake, and reviewed on the cadence you choose. Like every policy in the set it has nine numbered sections: purpose, scope, roles, policy statements, procedures, exceptions, enforcement, review cadence and a revision history table. Sections 4 and 5 carry the substance; those are the sections the Audit Kit rewrites for your named tools.

SOC 2 criteria this policy addresses

The Policyseed crosswalk maps 5 criteria to this policy. Each row names the section an auditor would read for that criterion; the criterion pages explain what it asks in plain words and list evidence examples.

CriterionWhat it coversWhere in this policy
CC7.5Recovering from incidentsSection 5 (Procedures)
CC9.1Mitigating business disruption riskSection 4 (Policy Statements)
A1.1Capacity managementSection 4 (Policy Statements)
A1.2Environmental protections, backups and recovery infrastructureSection 4 (Policy Statements)
A1.3Recovery testingSection 5 (Procedures)

Evidence auditors typically ask for

A policy is tested against artifacts. These are examples from the crosswalk for the criteria above, written for a small SaaS company; the Audit Kit ships the full list as an evidence checklist with an owner per row.

  • Post-mortem document for each significant incident with root cause and follow-up actions
  • Remediation tracker entries created from post-mortem action items and their closure dates
  • Backup restore test record used as part of incident recovery or recovery testing
  • Business continuity and disaster recovery plan with recovery time and recovery point objectives per system
  • Risk register entries for availability and disruption risks with treatments
  • Cyber insurance policy summary or documented decision not to carry it
  • Capacity or utilization dashboard screenshot from the monitoring tool (CPU, memory, storage, database connections)
  • Alert rules for capacity thresholds with the person or channel they notify

How to put the Business Continuity and Disaster Recovery Policy in place

Adopting the document is the easy part; the auditor tests whether it runs. These steps take it from a signed policy to a control with evidence, for a cloud company of 5 to 200 people.

  1. Build the system inventory and run the business impact analysis: for each system, record its dependencies on your cloud provider and vendors, estimate the impact of an outage at 1, 4 and 24 hours and 5 days, and assign Tier 1, 2 or 3.
  2. Have Executive Management approve the tier list together with an RTO and RPO for each tier, and save the approval (a signed document or a dated ticket) as evidence.
  3. Write a recovery procedure for every Tier 1 and Tier 2 system that an engineer who did not build the system could follow. Store it in your source control platform beside the infrastructure code, and keep an offline copy (for example an encrypted export in a separate cloud account or a printed binder) for when source control is down.
  4. Make sure Tier 1 systems can survive the loss of one availability zone by using multi-zone databases and load balancing, and confirm your infrastructure code can rebuild the environment in another region or account.
  5. Set up backups with copies in a separate account or region, meeting each tier's RPO under the Backup and Recovery Policy, and cross-reference the two policies in your test records.
  6. Set up an alternative communication channel that does not depend on your primary chat or email (for example a separate chat workspace or a phone tree), build the contact roster covering personnel, Executive Management, cloud provider support, key vendors, counsel and insurers, and refresh it every quarter.
  7. Write down the declaration criteria and who may declare, and keep a declaration template with fields for time, reason and recovery lead. Prepare customer notice wording and a status page procedure so the first notice goes out within the time the policy sets.
  8. Once a year, run the technical test: restore a Tier 1 environment from infrastructure code and backups into an isolated account or region (for example, rebuild the stack with Terraform and restore last night's RDS snapshot into a scratch account), timestamp each step, and record whether the RTO and RPO were met.
  9. Once a year, run a tabletop with Executive Management, Engineering and People Operations using a scenario such as loss of a cloud region, ransomware on the primary database or loss of a critical vendor. Record attendees, decisions, gaps and actions.
  10. After every test or declared disaster, complete the review within ten business days with a timeline, achieved RTO and RPO, deviations and corrective actions with owners and due dates. Add unresolved Tier 1 gaps to the risk register, and name a trained alternate for every role in a recovery procedure.

Common audit findings and how to avoid them

What the auditor findsHow to avoid it
The plan exists, but there is no evidence it was tested during the audit period, or only a tabletop was held with no technical restore.Run and document both the annual technical restore and the annual tabletop. Keep the runbook used, timestamps, restore verification output, attendee list and sign-off.
The test did not compare achieved recovery time and data age against the approved RTO and RPO, or the objectives were never formally approved.Keep the approved tier list with RTO and RPO, and record achieved times and restored-data age in every test record, along with any variance and its explanation.
The business impact analysis is out of date, does not cover a recently added customer-facing system, or has no approval.Re-run the analysis at least annually and whenever a new customer-facing system is added, and keep the dated inventory, tier list and Executive Management approval.
Corrective actions from the last test or incident have no owner, no due date or are still open with no risk register entry.Track each action in the ticketing system with an owner and due date, keep the post-test review that lists them, and show the risk register entry for any open Tier 1 gap.
The contact roster or recovery procedures are stale: they name people who have left, or reference systems that no longer exist.Keep the quarterly roster refresh record and the annual procedure review record, and reassign recovery roles within the period the policy sets when someone departs.

Full text: the Business Continuity and Disaster Recovery Policy rendered for Northwind Cloud Inc

Below is the complete template as the generator renders it for Northwind Cloud Inc, a fictional 11-50-person remote company running Northwind on AWS, Vercel, GitHub and Okta. Every name, tool and date comes from that sample intake; your answers replace them. Section headings carry anchors so the criterion pages can link straight to the section they cite.

Sample document

Northwind Cloud Inc Business Continuity and Disaster Recovery Policy

1. Purpose

This policy ensures that Northwind Cloud Inc can keep delivering Northwind, or restore it within defined targets, when a disruption occurs. It sets recovery objectives by system tier, assigns authority to declare a disaster, defines how recovery is executed and communicated, and requires the plan to be exercised so that targets are evidenced rather than assumed. Because the Availability criteria are within the scope of Northwind Cloud Inc's SOC 2 examination, the objectives in this policy are commitments that must be supported by test results and monitoring evidence.

2. Scope

This policy applies to all systems, data, personnel, facilities and vendors that Northwind Cloud Inc relies on to deliver Northwind and operate the business, including production infrastructure in AWS and Vercel, code and pipelines in GitHub and GitHub Actions, identity services in Okta, corporate SaaS and the people who operate them. It covers disruptions of any cause: infrastructure outages, security incidents, data corruption, loss of a key vendor, and loss of a workplace.

Systems are classified into three tiers. The tier sets the recovery time objective (RTO, the maximum time to restore service) and the recovery point objective (RPO, the maximum data loss measured in time). Engineering maintains the authoritative tier list in the system inventory.

TierDescriptionTypical systemsRTORPO
Tier 1 - CriticalLoss stops customers using Northwind or exposes customer dataProduction application and API, primary database, authentication and sessions, DNS and edge, secrets management, core accounts in AWS and Vercel4 hours1 hour
Tier 2 - ImportantNeeded to operate, support and change Northwind within one business dayRepositories in GitHub, GitHub Actions pipelines, Datadog monitoring, alerting and on-call tooling, customer support tooling, Okta administration, billing24 hours24 hours
Tier 3 - DeferrableCan be unavailable for several days without customer impactInternal wikis, analytics and reporting, development and staging environments5 business days7 days

3. Roles and Responsibilities

  • Executive Management approves this policy and the recovery objectives, may declare a disaster and approve emergency spend, alternative sites or vendors, decides on customer and public communication, and reviews the results of every test.
  • Security Owner (Dana Whitfield, CTO) owns the business impact analysis, ensures security controls remain in force during recovery, coordinates with the Incident Response Policy when the disruption is security-related, and keeps the plan, contact roster and vendor dependencies current.
  • Engineering owns the recovery procedures for every Tier 1 and Tier 2 system, maintains the infrastructure code and backups that make recovery possible, staffs the recovery team, executes and documents tests, and reports achieved recovery times against the objectives.
  • People Operations maintains the personnel contact roster, accounts for the safety and availability of personnel during a disruption, arranges cover for unavailable staff, and communicates working arrangements.
  • All Personnel know how to reach the alternative communication channel, keep their contact details current, and follow instructions from the recovery lead during a disruption.

4. Policy Statements

  • 4.1 Northwind Cloud Inc performs a business impact analysis at least annually, identifying each system that supports Northwind and the business, the impact of its loss over time, its dependencies and its tier. Tier assignments and RTO and RPO targets are approved by Executive Management.
  • 4.2 Every Tier 1 and Tier 2 system has a written recovery procedure that a competent engineer who did not build the system can follow, stored in GitHub alongside the infrastructure code and in an offline copy for use when GitHub is unavailable.
  • 4.3 Tier 1 systems are designed to survive the loss of a single availability zone or data centre without exceeding the RTO, using the redundancy features of AWS and Vercel, and infrastructure is defined as code so that environments can be rebuilt in an alternative region or account.
  • 4.4 Data required to meet each tier's RPO is backed up under the Backup and Recovery Policy, with copies in a separate account or region so that one administrative error or compromise cannot destroy both primary and backup.
  • 4.5 A disaster may be declared by Executive Management, the Security Owner or the Engineering on-call lead when a Tier 1 system has been unavailable or degraded for one hour, the RTO is at risk, or the primary environment cannot be trusted. The declaration records the time, reason and recovery lead.
  • 4.6 During a declared disaster the recovery lead may provision infrastructure, engage vendor support, restore from backups and incur emergency spend within limits set by Executive Management without standard change approval; every action is logged and reviewed afterwards.
  • 4.7 Security controls, including access control, encryption and logging, remain in force during recovery. Emergency access granted to recover a system is time-limited, logged and revoked when recovery is complete.
  • 4.8 Northwind Cloud Inc maintains an alternative communication channel that does not depend on its primary systems, and a contact roster covering personnel, Executive Management, support at AWS and Vercel, Supabase, Stripe, Datadog and Slack, outside counsel and insurers; personnel learn how to reach the channel during onboarding.
  • 4.9 Because Northwind Cloud Inc operates a Remote work model, loss of any single workplace does not interrupt operations; personnel work from an alternative location with their managed device. No production capability may depend on physical access to an office.
  • 4.10 Tier 1 recovery is tested at least annually by restoring a production-equivalent environment from infrastructure code and backups, and the whole plan is exercised at least annually through a tabletop with Executive Management; achieved recovery times are recorded against the objectives.
  • 4.11 Critical vendors, including Supabase, Stripe, Datadog and Slack where they support Tier 1 or Tier 2 systems, are assessed for continuity risk under the Vendor and Third-Party Risk Management Policy, and every Tier 1 dependency has a documented exit or fallback plan.
  • 4.12 Customers are kept informed during any disruption affecting Northwind through the status page or the channel committed in their contract, with updates at least hourly during a Tier 1 disruption.
  • 4.13 After every declared disaster and every test, Engineering completes a review within ten business days recording the timeline, achieved RTO and RPO, deviations from the plan and corrective actions with owners and due dates; unresolved gaps affecting Tier 1 objectives are entered in the risk register.
  • 4.14 Personnel with a role in recovery are trained on their responsibilities at least annually, and no Tier 1 recovery procedure depends on a single named individual.

5. Procedures

  • 5.1 Business impact analysis. Each year, and whenever a new customer-facing system is added, the Security Owner and Engineering review the system inventory, confirm each system's tier, document dependencies on the services of AWS and Vercel and on Supabase, Stripe, Datadog and Slack, and estimate the impact of an outage at 1, 4 and 24 hours and 5 days, and Executive Management approves the resulting tier list.
  • 5.2 Plan maintenance. Engineering reviews each Tier 1 and Tier 2 recovery procedure at least annually and after any material architecture change, confirming that the infrastructure code in GitHub still builds the environment and that restore steps match the backup tooling, currently AWS Backup. The Security Owner refreshes the contact roster quarterly.
  • 5.3 Declaration and mobilisation. When the criteria in 4.5 are met, the declaring person records the declaration, names the recovery lead, opens the alternative channel and notifies Executive Management. The recovery lead assembles the team, assigns a scribe and confirms the affected systems, tiers and start time; security-related disruptions run on one timeline shared with the Incident Commander.
  • 5.4 Recovery execution. The recovery team follows the written procedure for each affected system in tier order: identity and secrets first, then data from the most recent backup meeting the RPO, then application infrastructure from GitHub through GitHub Actions (or a manual pipeline if GitHub Actions is unavailable), then edge and DNS. The scribe records each step, its timing and any deviation.
  • 5.5 Verification. Before a recovered system is returned to customers, the team verifies data integrity against expected record counts or checksums, runs the automated and smoke tests, confirms that logging to Datadog and alerting work, and confirms with the Security Owner that access controls and encryption are in force.
  • 5.6 Communication. The recovery lead posts an initial customer notice within 30 minutes of declaration and updates at least hourly for Tier 1 disruptions, using wording approved by Executive Management. People Operations confirms personnel safety and communicates working arrangements. Customers with contractual notification terms are notified within those terms.
  • 5.7 Return to normal. When objectives are met and verification is complete, the recovery lead declares the end of the disruption, confirms that temporary infrastructure, emergency credentials and firewall exceptions are removed, ensures backups are running against the recovered environment, and schedules the review required by 4.13.
  • 5.8 Annual technical test. Engineering restores a Tier 1 environment from infrastructure code and backups into an isolated account or region, measures elapsed time and the age of restored data, and records whether RTO and RPO were achieved, with the runbook used, timestamps, verification evidence and defects found.
  • 5.9 Annual tabletop. The Security Owner runs a scenario exercise with Executive Management, Engineering and People Operations, using scenarios such as loss of a hosting region at AWS and Vercel, ransomware affecting the primary database, or loss of a critical vendor such as one of Supabase, Stripe, Datadog and Slack, and records decisions, gaps and actions.
  • 5.10 Personnel continuity. People Operations maintains at least one trained alternate for each role named in a recovery procedure, and reassigns responsibilities within five business days when a person in a recovery role leaves or is on extended absence.

6. Exceptions

Exceptions, including a system that cannot meet its tier's objectives, require written approval from Executive Management on the Security Owner's recommendation, a compensating measure or accepted risk in the risk register, and an expiry date no more than 12 months away, and are reviewed at each annual policy review. Customer commitments exceeding the Section 2 objectives require Engineering confirmation before signature.

7. Enforcement

Failing to maintain recovery procedures, backups or infrastructure code for a system in one's care, bypassing security controls during recovery, or failing to take part in required tests is a violation of this policy and may result in disciplinary action up to and including termination of employment or contract. Vendors whose continuity commitments are not met are subject to the remedies in their contract.

8. Review Cadence

Engineering and the Security Owner review this policy on a annual basis and after every declared disaster, any test in which a Tier 1 objective was missed, significant changes to the architecture of Northwind or to Northwind Cloud Inc's footprint on AWS and Vercel, and changes to customer commitments. Each review is approved by Priya Natarajan, CEO and recorded in Section 9.

9. Revision History

VersionDateDescriptionApproved by
1.02026-09-02Initial releasePriya Natarajan, CEO

Frequently asked questions

Who should own the Business Continuity and Disaster Recovery Policy?
In the Policyseed template the Engineering Lead owns the Business Continuity and Disaster Recovery Policy: they maintain the text, run the procedures in section 5 and hold the evidence those procedures produce. The approver you name in the intake signs it, and section 8 sets the review cadence you choose (annual, semi-annual or quarterly).
Which SOC 2 criteria does the Business Continuity and Disaster Recovery Policy address?
5 criteria in the Policyseed crosswalk: CC7.5 (Recovering from incidents), CC9.1 (Mitigating business disruption risk), A1.1 (Capacity management), A1.2 (Environmental protections, backups and recovery infrastructure) and A1.3 (Recovery testing). Each mapping points at a numbered section of this policy, and the Audit Kit exports the same mapping as an Excel crosswalk with an evidence checklist.
Is the Business Continuity and Disaster Recovery Policy template free to use?
Yes. The template is Apache-2.0 licensed and the generator renders it in your browser with your company, stack and owner names filled in; nothing is stored server-side. The Audit Kit ($39 one-time) rewrites sections 4 and 5 for your named tools with Claude and adds Word documents, the crosswalk spreadsheet, acknowledgment forms and a review calendar. Refunds are available within 14 days on request. Policyseed provides governance policy templates, not legal advice; the CPA firm performs the examination.

Policyseed provides governance policy templates and AI tailoring. It is not legal advice and not a compliance guarantee. Management adopts the policies; the CPA firm performs the SOC 2 examination.