Mon–Sat 10:00–18:00 London · UK
Remote & on-site ☎ 0207 096 0936
← Guides
Guide

How to Run a Disaster Recovery Test That Proves Something

A backup you have never restored is a hypothesis. Here is how to run both kinds of disaster recovery test — the tabletop and the live restore — without taking production down.

A backup that has never been restored is a hypothesis, not a backup. It is a claim your software makes about itself every night, in green text, with nobody checking. Until you have pulled real data back out and opened it, that claim is untested — and the morning you need it is a poor time to find out.

Businesses that get caught out here are rarely careless. They have backups running. What they have never done is close the loop. Writing the plan is the first job, and the business continuity plan guide covers what belongs in that document. This one is about the harder part: proving the plan works.

There are two kinds of test, and you need both. They find completely different faults.

The tabletop: an hour and an awkward scenario

Get the right people round a table for an hour. Nobody touches a keyboard. One person reads out a scenario and everyone talks through what they would do.

Make the scenario specific and slightly cruel. Ransomware hit at 9am on Monday. The file server is encrypted, and the finance manager is on holiday. Then work forward in real time. Who decides this is an incident rather than a glitch? Who phones the IT provider — and what is the number? Not “it is on the intranet”. The intranet is down. Say the number out loud, now.

That last exercise is the sharpest tool in the box. A plan that lives only on the system you have just lost is not a plan. Within about ten minutes a tabletop tells you whether your documented process is real or decorative, and it costs nothing but the hour.

The live restore test: prove the data opens

The tabletop tests your people. The restore test tests your technology, and it works best in three steps of increasing nerve.

One file. Pick a real document from a real folder and restore it. Time the whole thing, including finding it in the backup console.

One folder. Restore a shared folder with its structure intact. Check the permissions came back too — a restore that returns the files but flattens who can see them has created a data problem, not solved one.

One whole system. Restore a full server or a key workstation to isolated hardware or a cloud instance, and boot it.

Then do the step almost everyone skips: open the data. Log into the restored system and use it. Run a report. Open the accounts file. Mount the database. Files can restore perfectly and still be useless if the application they belong to will not start, or if corruption was quietly copied into the backup weeks ago. A file that exists is not the same as a file that works.

Scope it so the test is not the disaster

Testing badly is worse than not testing. A few rules keep it safe.

Restore to a different location, never over the top of live data. Put restored servers on an isolated network segment, so a recovered mail server or domain controller cannot start arguing with the production one. Take a fresh backup before you begin. Pick a window when losing an hour would not hurt. And tell people it is happening, so a helpful colleague does not report your test machine as an intrusion.

If your backup and continuity setup makes an isolated restore genuinely difficult, that is a finding in itself. Recovery you cannot rehearse is recovery you cannot rely on.

Choose the scenario you would rather avoid

The convenient scenario is the one you already know you can handle. Deleted file, single failed laptop, tick. Useful once, then worthless.

The scenarios worth testing carry a complication. The person who normally fixes it is unreachable. The building is shut, so nobody can touch the hardware. The backup itself is suspect, because the encryption sat dormant for a fortnight before it triggered. Ransomware deserves its own run-through, and the step-by-step ransomware recovery guide makes a reasonable script to test against.

Vary it each year. The same scenario twice teaches you nothing the second time.

Get non-IT people in the room

Recovery is not an IT event. It is a business event with an IT cause.

Bring whoever talks to customers, whoever can authorise emergency spend, whoever handles staff communication, and whoever owns the system in question. IT can restore a server. IT cannot decide whether to tell clients, whether to trade on paper for a day, or whether a faster recovery is worth what it costs. Those decisions get made badly under pressure unless someone has thought about them in advance.

Measure two numbers, and compare them honestly

Every test should produce two figures.

How far back did you lose data? If the last clean backup ran at 11pm and the incident hit at 3pm, you have lost a working day of changes.

How long were you down? Measured from the incident to the point people could actually work — not to the point the restore finished.

Then compare both against what the business said it could tolerate. That tolerance has to come from the business, not from IT. If finance says half a day of lost invoicing is survivable and your test shows a full day, you have a gap with a price attached, and now you can have a sensible conversation about closing it.

Write down what broke

The first test always finds something. Not usually — always.

It tends to be mundane. A licence key nobody kept. An administrator password that left with someone. A dependency nobody documented, so the application starts but cannot reach the licensing server it quietly needs. Restore speeds that were fine for a folder and hopeless for a terabyte.

Log every one with an owner and a date. A test that produces a list of fixes has done its job. A test that produces a clean sheet usually means you tested the easy thing.

How often to do this

Annually at the absolute minimum. Again after any significant change — a new server, a new line-of-business application, a move to different premises. And always after a migration, because the system most likely to be missing from a backup is the one you only started using last month.

If you would like someone to run the first test alongside you and be straight about what it turns up, that is what our disaster recovery work involves. No theatre. Just a restore that either works or tells you exactly what to fix.

Frequently asked questions

How long does a disaster recovery test take?

A tabletop walkthrough is genuinely an hour, including the argument it will start. A live restore of a single file or folder takes minutes once the tooling is set up. Restoring a whole system to isolated hardware is the long one — allow half a day the first time, because the first attempt is where you discover the missing licence key or the driver nobody thought about. It gets much quicker on the second run, which is rather the point of doing it before you need it.

Can we test recovery without disrupting staff?

Yes, if you scope it deliberately. Restore to a separate location rather than over live data, keep restored servers on an isolated network so they cannot talk to the production copies, and pick a quiet window. The one thing you should not do is keep it secret — tell the team a test is happening, or someone will see an unfamiliar machine appear and raise a genuine incident.

Our backup software reports success every night. Is that not enough?

No. A success report tells you the job ran and the files were copied. It says nothing about whether the data inside those files is usable, whether you have the credentials to restore it, or whether the restore would finish inside the window the business can survive. Plenty of backups report green for years and fail at the only moment that counts. The report is a smoke alarm, not a fire drill.

Related services

Free · no obligation

Want a hand with any of this?

Tell us what you're trying to sort out and we'll come back with a clear, no-obligation plan and price.