Managed IT
A backup is a copy of your data. Business continuity is a promise about how long you will be down and how much work you will lose. Those are different purchases, and plenty of businesses have bought the first while believing they own the second. The gap between them has two names, RTO and RPO, and the only way to find out whether yours are real is to test a restore on a day nothing is wrong.
Backup software is very good at telling you it succeeded. A green job status means data was written somewhere. It does not mean the data is complete, that the copy is readable, that anyone knows the restore procedure, or that the systems you would restore onto still exist. Those are separate claims, and the backup report makes none of them.
A backup is not a recovery plan
The Unitrends State of Backup Report puts 58% of SMBs as never having tested a restore. For most of that group the first real test happens during an outage, with people waiting, which is the worst possible set of conditions for discovering that a job has been silently writing an incomplete copy for eleven months.
The cost of finding out late is what makes this worth scheduling. Gartner puts the average cost of IT downtime at $5,600 per minute, before data loss and recovery labour. Longer outages get disproportionately worse: a widely repeated figure, credited to the University of Texas and carried by SCORE, holds that 93% of businesses that lose data access for ten days or more file for bankruptcy within a year. That number is old and worth treating as directional rather than precise, but the direction is not in dispute. Extended outages are not a scaled-up version of short ones.
Backup failures are silent by design
A failed backup produces no symptom. Nothing slows down, no user complains, and the only signal is a status field somebody has to go and read. Compare that to a failed server, which announces itself immediately. It is the same pattern as unread firewall logs or stale access point firmware: the systems most likely to be quietly broken are the ones nothing prompts you to check.
The practical consequence is that backup health cannot be event-driven. It needs a calendar, because nothing else is going to raise its hand.
Backup, continuity, disaster recovery
Three Different Purchases
These three get used interchangeably in sales conversations and mean quite different things in an outage. Each one answers a different question, and each costs more than the one before it.
What it gets you
Backup: the data still exists
Scheduled copies · Retention window · Separate from production · Restorable, eventually
Backup answers one question: does a copy exist. It says nothing about how long recovery takes, and for most businesses running file-level backup to cloud storage the honest answer is hours to days, assuming the restore works on the first attempt.
Two things determine whether a backup is worth anything in a ransomware event. Whether it is reachable from the environment being encrypted, and whether an administrator account that has been compromised can delete it. That is what immutable storage addresses, and it is covered properly in the immutable backup guide rather than repeated here. The short version is that a backup an attacker can reach is not a backup.
What it gets you
Business continuity: you keep operating
Defined RTO and RPO · Standby images · Prioritised systems · Tested procedure
Continuity is a commitment about time. Critical systems come back within an agreed window, using standby images that can be spun up locally or in a cloud environment rather than rebuilt from a file-level copy. The difference in an outage is minutes against days.
What makes this a different purchase rather than a better backup is the work around it. Somebody has to decide which systems are critical and in what order they come back, because restoring everything at once is neither possible nor useful. That prioritisation is the part most likely to be missing, and it is not something a vendor can do for you without asking which parts of the business stop when which system is down.
What it gets you
Disaster recovery: the site is gone
Geographic separation · Full environment rebuild · Documented runbook · Periodic failover test
Disaster recovery assumes the primary location is unavailable, whether through fire, flood, extended power loss or a compromise severe enough that the environment cannot be trusted. It requires the copy to be geographically separate, and it requires documentation good enough for somebody to follow under pressure.
The distinction from continuity is scope rather than speed. Continuity assumes you still have somewhere to restore to. Disaster recovery does not, which is why the runbook matters as much as the data and why a plan nobody has walked through is closer to a hope than a plan.
Which one you actually need
Most businesses need more than backup and less than full disaster recovery, and the way to find out where you sit is not to pick a tier from a list. It is to work out what an hour of downtime costs, and what a day of lost work costs, and then see which level those numbers justify. That calculation is the next section.
RTO and RPO, and how to set yours
Two numbers describe what a recovery plan actually promises. They are simple individually and get confused constantly, usually because both are quoted in hours.
RTO: how long until we are working again
Recovery Time Objective is the maximum time a system can be down before the consequences become unacceptable. It runs forward from the moment of failure and it is a business decision rather than a technical one. The technical work is making the number achievable; deciding what the number should be belongs to whoever owns the consequences of being down.
RPO: how much work we can afford to lose
Recovery Point Objective is the maximum amount of data you can afford to lose, measured backwards from the failure. It is set by backup frequency and nothing else. Nightly backups mean an RPO of up to 24 hours, which in practice means a failure at 4pm costs you the whole working day. If that is unacceptable, the answer is more frequent backups, not a faster restore.
Deriving your own numbers
There is no credible table of typical RTOs by industry, and any you find is somebody's guess formatted convincingly. The numbers come out of your own operations, and the method is straightforward enough to do in an afternoon.
Take your genuinely critical systems, which for most businesses is fewer than ten. For each, ask what stops when it is unavailable, how long that is survivable, and what the organization does in the meantime. Some systems have workable manual fallbacks that stretch an acceptable outage into days. Others stop revenue within an hour. Then ask what re-entering a day's lost work would take, since that is what sets the RPO. The output is a per-system pair of numbers and a recovery order, which is the thing most plans are missing.
Expect the exercise to reclassify things. Systems assumed critical turn out to have a workaround, and something nobody listed turns out to block everything else.
How to test a restore
A test is the only thing that converts a stated RTO into a real one. It does not need to be elaborate, but it does need to be honest about a few things.
Restore to somewhere other than production, so a failed test cannot become an incident. Restore something that matters rather than a token file, because small restores succeed in ways full ones do not. Time it from decision to usable, including the part where somebody finds the documentation, since that is real recovery time and it is where the surprises live. Have somebody who is not the usual administrator run it, which tests whether the procedure is written down or resident in one person's memory. Then check the restored data properly: open records, confirm the most recent entries are present, and verify the restore point is where you expected it to be.
Record what the test actually took, not what it should have taken. A test that produces a slower number than your stated RTO has not failed. It has told you the truth, which is the entire point of running it before you need it.
Backup Readiness Checklist
The Checklist
Eight questions across restore testing, ransomware protection, retention and compliance, and recovery objectives. Tick the ones you can answer without going and asking someone, and the score updates as you go. Use it yourself or share it with your team.
Tick what you can answer without checking. The blanks are the gaps.
Restore Testing
When did somebody last restore something that mattered, and how long did it take?
A date and a duration. A green job status is a claim about writing data, not about reading it back.
Could somebody other than your usual administrator run a recovery?
This tests whether the procedure is documented or living in one person's memory.
Ransomware Protection
Is your backup reachable from the environment it is protecting?
A copy an attacker can browse to is a copy an attacker can encrypt.
Could a compromised admin account delete your backups?
Same credentials for production and backup means one compromise takes both.
Retention & Compliance
How far back can you actually restore from, and who chose that window?
Often a default nobody selected. Intrusions frequently predate the oldest recovery point.
Can your provider produce current certifications and backup records on request?
Auditors, PCI assessors and underwriters ask for the same evidence. Worth having before they do.
Recovery Objectives
What is your RTO for each critical system, and in what order do they come back?
The recovery order is the part most plans are missing, and it cannot be decided during the outage.
How much work would you lose if everything failed right now?
Your RPO, set entirely by backup frequency. Nightly backups make a 4pm failure cost the whole day.
Frequently asked questions
Backup answers whether a copy of your data still exists. Business continuity answers whether you can keep operating, which is a commitment about time rather than about data. Continuity uses standby images that can be brought up locally or in a cloud environment, so critical systems return within an agreed window instead of being rebuilt from a file-level copy over hours or days. The other difference is the work around it: somebody has to decide which systems are critical and in what order they come back, because restoring everything at once is neither possible nor useful.
RTO, the Recovery Time Objective, is how long a system can be down before the consequences become unacceptable. It runs forward from the moment of failure. RPO, the Recovery Point Objective, is how much data you can afford to lose, measured backwards from the failure. RTO is time to get back; RPO is work you never get back. They are set by different mechanisms, since RPO is determined entirely by backup frequency while RTO depends on recovery method, so improving one does nothing for the other.
Derive it from your own operations rather than from a benchmark table, because no credible table of typical RTOs by industry exists. Take the systems that are genuinely critical, which for most businesses is fewer than ten. For each one, ask what stops when it is unavailable, how long that is survivable, and what the organization does in the meantime. Some systems have manual fallbacks that stretch an acceptable outage into days, and others stop revenue within an hour. The output is a per-system target and a recovery order, which is the part most plans are missing.
Restore to somewhere other than production so a failed test cannot become an incident, and restore something that matters rather than a token file, because small restores succeed in ways full ones do not. Time it from decision to usable, including the part where somebody locates the documentation, since that is real recovery time. Have somebody other than the usual administrator run it, which tests whether the procedure is written down. Then check the restored data properly by opening records and confirming the most recent entries are present.
They overlap but differ in scope. Business continuity assumes you still have somewhere to restore to and focuses on bringing critical systems back inside an agreed window. Disaster recovery assumes the primary site is unavailable altogether, through fire, flood, extended power loss or a compromise severe enough that the environment cannot be trusted. That requires geographic separation and documentation good enough for somebody to follow under pressure, which is why the runbook matters as much as the data.
Frequently enough to meet the RPO you have set, since backup frequency is the only thing that determines it. Nightly backups mean losing up to a day of work, so a failure at 4pm costs the whole working day. If that is unacceptable for a particular system, the answer is more frequent backups for that system rather than a faster restore, because restore speed affects RTO and does nothing for RPO. Different systems can reasonably run on different schedules based on how much work each accumulates.
Yes, and separation means more than a different folder or server. A backup in the same environment as production shares that environment's risks, including the ransomware that encrypted it and the flood or fire that took the building. Geographic separation is what makes recovery possible when the primary site is unavailable. The related question is whether the copy is reachable from the environment being protected, since a backup an attacker can browse to is one they can encrypt or delete.
Usually because nothing failed loudly beforehand. A backup problem produces no symptom: nothing slows down, no user complains, and the only signal is a status field somebody has to read. The common causes are scope drift, where new servers and migrated file shares quietly fall outside the job, incomplete copies that report success, retention windows shorter than the problem being recovered from, and procedures nobody has rehearsed. All of them are visible in a scheduled restore test and invisible in a backup report.
When did you last test a restore?
Get a free backup assessment. We'll review what's actually in scope, run a restore against something that matters, and give you a recovery time measured rather than estimated.



