Backups Aren’t Backups Until You’ve Restored One

A backup that has never been restored isn’t a backup. It’s an unverified assumption sitting on a disk somewhere, and the number of organisations that discover the difference only during an actual outage remains remarkable given how well-known this failure mode is.

This isn’t a novel observation — every experienced technologist has heard some version of it. What’s less discussed is why, given how well-known it is, restore testing still doesn’t happen as routinely as backup jobs themselves, and what a genuine restore-testing discipline actually requires beyond the platitude.

Why the backup runs and the restore never does

Backup jobs are automated, scheduled, and monitored for the narrow thing they’re built to report: did the job complete without error. That’s a meaningfully different question from “is this backup actually restorable, to a working state, within the time the business needs.” A backup job can complete successfully every single night for years while quietly backing up a corrupted database, an incomplete snapshot, or a configuration that’s silently drifted out of sync with what a real restore would need — and the monitoring that watches the job succeed has no visibility into any of that, because it was never built to check it.

Free · 4 minutes

Do you actually know what you are running — and what it is about to cost you?

Fourteen questions on the systems you depend on, the ones nobody owns, and the support dates that turn a routine upgrade into a forced re-platform. Banded finding on screen, full sheet by email.

Restore testing doesn’t get the same automated attention because it’s operationally disruptive to do properly, genuinely time-consuming to do well, and — critically — it never produces a visible problem until the day a real restore is actually needed and fails. A backup job’s success is confirmed every night. A restore’s success is confirmed only when tested, and testing competes for time against everything else that feels more urgent in the moment, which it reliably loses to, repeatedly, until the day it can no longer afford to.

What a genuine restore test actually verifies

A real restore test checks four things a completed backup job says nothing about. Completeness — does the restored system actually contain everything it should, or only the subset the backup process happened to capture cleanly. Integrity — is the restored data actually usable, or does it merely exist, technically present but corrupted in ways that only surface once something tries to use it. Time — how long does a genuine restore actually take, measured against the recovery time objective the business believes it has, not the theoretical time a vendor’s documentation claims. Dependencies — does the restored system actually work once live, or does it fail because of configuration, credentials, or external connections that existed in the original environment and were never captured by the backup in the first place.

Each of those is a genuine failure mode, independently, and a backup job’s green checkmark verifies none of them. Restore-tested backups are the only ones that have actually confirmed all four; everything else is an assumption resting on the narrower, easier question of whether last night’s job completed without throwing an error.

Making restore testing routine rather than heroic

The fix isn’t a single annual disaster-recovery exercise, valuable as that is — it’s treating restore testing as a scheduled, boring, recurring operational task, the same way the backup job itself is scheduled and monitored. A rotating sample of systems, restored on a defined cadence, to an isolated environment, checked against the four criteria above, with results tracked over time rather than treated as a pass-fail event that gets forgotten once complete. This is the same discipline I write about across regulated sectors elsewhere on this site — an RTO or RPO that has never been tested is a hope stated with the confidence of a fact, and the only way to know which one it actually is, is to test it.

What a first real test usually finds

A mid-sized firm running its first genuine restore test in over two years — beyond the routine “job completed successfully” monitoring — discovered that a critical application’s backup had been silently excluding a configuration directory for the better part of a year, following an infrastructure change nobody had thought to check against the backup scope. Every nightly job had completed without error throughout that entire period; the backup software had no way to know it was backing up an incomplete picture of what a real restore would actually need. The gap was found in a scheduled test, in business hours, with no pressure and no consequence beyond the finding itself. Found instead during an actual incident, the same gap would have meant a restored system that came up broken, at the exact moment nobody had time left to diagnose why. The scheduled test cost an afternoon. The alternative would have cost the incident.

The same tested-versus-assumed distinction applies at the infrastructure level — see multi-cloud: when it’s insurance and when it’s tax.

Establishing a genuine restore-testing cadence — and finding out, deliberately, whether current recovery objectives are real capability or documented assumption — is core technology control assessment work, and it’s considerably cheaper to discover the gap in a scheduled test than during a live incident.

Free interactive tool

Website compliance checklist

What your site has to do, based on what it actually does

Answer as much or as little as you like — the list builds as you go. Nothing is stored against your name and no email is required.

Free interactive tool

Interactive deadline calculator

Check which regulations apply to you and when

Regulation across the EU, UK, US and Asia-Pacific has moved considerably in the past eighteen months, and several headline dates have shifted more than once. Twelve questions, about three minutes.

Results are shown on screen — no email required. A dated summary is available to download, and can be sent on if that's more useful. What we do with your answers.

Most technology problems are not technology problems. They are control problems.

The systems exist. The investment has been made. The question is whether leadership can understand, direct, evidence, and sustain what those systems produce. Find out where control exists — and where it only appears to.

Full Governance by Sixteen Pillars

Govern your business. Prove your compliance.

A board assurance cockpit for EU-regulated financial firms — tamper-evident, hash-chained proof of governance across DORA, GDPR, NIS2, ISO 27001, the EU AI Act and MiCA. In development.

See what's coming