Skip to main content
OT/ICS Cybersecurity

The Restore Drill: Prove Your OT Backup Works Before You Need It

Maersk's whole recovery from NotPetya hung on one domain controller in Ghana that a blackout had switched off. Luck is not a backup design. A bench drill for one controller project, one HMI and one historian, timed lap by lap, with a strict pass line.

A papercraft test bench with a controller rack, HMI panel and server tower joined by paper cables, beside a large red stopwatch. A port with cranes and container ships is visible behind it.

Item: one domain controller. Location: a Maersk office in Ghana. Condition: switched off by a blackout. Estimated value in the last week of June 2017: most of a global shipping company.

That server is the most famous backup in industrial security, and nobody chose it. When NotPetya wiped Maersk's network on 27 June 2017, the company's recovery team in Maidenhead found something they could work with and something they could not. Wired's reconstruction of the attack describes backups of almost all of Maersk's individual servers, between three and seven days old. What nobody could find was a backup of the domain controllers. There were 150 or so of them, set to sync with each other so that any one could stand in for the rest. That design had not planned for every one of them being wiped at once.

After calling hundreds of IT administrators around the world, the team found one survivor. A power cut had knocked the Ghana machine offline before the wiper arrived, and it stayed off. In Wired's words:

"It thus contained the singular known copy of the company's domain controller data left untouched by the malware."

The magazine's next six words do the job of a whole incident report: "all thanks to a power outage."

The rest was just as improvised. The Ghana office's bandwidth was too thin to send several hundred gigabytes to the UK. Nobody there held a British visa. So one employee flew the drive to Nigeria and handed it to a colleague at the airport, who carried it to Heathrow. Accounts of the route differ, but agree on the shape: the keystone of the recovery travelled as hand luggage.

Maersk's chair later told Davos the company rebuilt 4,000 servers and 45,000 PCs in ten days, at a cost of $250 million to $300 million. Wired reports that staff privately thought the figure low.

So the recovery plan that worked went roughly like this. Wait for the grid to fail in the right office in the right week. Find someone with the right visa. Book a flight. Luck is an excellent recovery strategy with one flaw: it never says in advance whether it is coming.

This article is for whoever would be handed the backup media on the bad morning and asked to bring a controller, an HMI and a historian back. If nobody at your site has that job yet, you have already found your first gap, and you found it for free.

A backup is a claim until it has a lap time

Maersk is a shipping company, not a plant, but the lesson travels. Replication felt like a backup until it faithfully copied the damage everywhere. A plant has its own Ghana question: which copy survives, and does it bring the process back? You cannot answer the second half by looking at a folder. You answer it on a bench, with a stopwatch. Three assets, one morning, two people, no connection to anything running.

The pass line is strict on purpose. A drill that passes cleanly the first time has either an excellent procedure or a lenient referee, and the second is more common.

Why these three? Each fails differently. The controller trips on firmware and versions. The HMI trips on a license bound to dead hardware, or on the screens the night shift added last spring. The historian is the one everybody forgets until someone asks for last month's records.

This is not an exotic standard. NIST's OT security guide, SP 800-82 Revision 3, points to bench tests and other offline testing for control system backups, never a download to a running controller. Version 1.0.1 of CISA's Cross-Sector Cybersecurity Performance Goals asks for OT backups stored separately and tested at least once a year. Neither tells you how long your own restore takes.

A quick check. Could you start lap 1 tomorrow with only what is in your isolated copy, installer and license included? If not, you have a finding without touching a controller.

"Our backups are fine"

"Our backups are fine. Every project gets saved to the engineering server after every change. I can show you the folder."

They probably are fine, as files. Maersk had files too: backups of almost every server, a few days old. The gap sat in the layer that everyone assumed was covered because it replicated. In a plant the equivalent gap usually sits right next to the file. The installer for the exact release the project was saved in. The license tied to a laptop that is now a paperweight. The spare controller in stores running firmware two versions behind.

Lap 1 tests exactly that. It is no criticism of whoever saved the projects. It tests everything the project needs that is not the project.

"We test restores every quarter"

"We test restores every quarter, and the backup console has been green for two years. What exactly does a stopwatch in a workshop add?"

A green console says the job completed. A quarterly restore test says the server boots. Neither says the HMI talks to a controller, the historian collects a tag, or the engineering software starts without phoning its vendor. Those are laps 3, 4 and 1, where OT restores stall.

The second half of the answer is about where the console lives. Maersk's domain controllers did exactly what they were built to do, which is how they all died together. CISA's ransomware guide warns that attackers look for reachable backups and delete or encrypt them before the main event. So the question for the backup admin is not whether the jobs succeed. It is this: if an attacker held every credential in the production domain, domain administrator included, which copy of these three projects could they still not touch?

The backup admin is the ally here, not the opposition. They own the isolated copy; the automation engineer owns the toolchain. The drill puts both at one bench, which in many plants would be a first.

"Not near my production line"

"I'm not having anyone poke at production to prove a point. If a test trips the line, it's my phone that rings at three in the morning, not yours."

Agreed, and the drill is built so that it cannot. No cable runs to production, the controller is a spare from stores, and the switch has no uplink. Declaring the live machines gone means nobody touches them at all, a stricter rule than most maintenance work follows. The cost is two people for a morning. If the site has no spare controller, that is worth knowing too, and the plant manager can fix it.

What the plant manager gets back is a number. Add the lap times for the systems you need to restart the process, one after another, because on the real day they queue for the same engineers. Then ask operations how long the process can run on manual fallback. If the restore adds up to twenty hours and manual holds for eight, the twelve-hour gap is the clearest spending case your recovery budget will ever see. The figure belongs to operations, not IT, because the process sets the limit.

Does that comparison exist at your site? If yes, the drill checks it. If no, the drill produces half of it.

Stop waiting for a blackout

The Ghana server saved Maersk because the power failed at the right moment and a colleague in Nigeria held the right visa. Nobody can write that into a procedure. What you can write is a lap sheet, a hunt log and a pass line, and run them on a quiet Tuesday when nothing is on fire.

A backup becomes a restore only when someone has timed it. Until then it is a hopeful folder. Time yours before an attacker or an unlucky week does it for you.

Sources: Wired, Andy Greenberg, The Untold Story of NotPetya, the Most Devastating Cyberattack in History, 22 August 2018 | NIST SP 800-82 Revision 3, Guide to Operational Technology Security | CISA Cross-Sector Cybersecurity Performance Goals | CISA #StopRansomware Guide

ICS/OT Incident Response Drills and Exercises

An incident response plan is only a theory until the people who must use it have run it under pressure. This course helps OT engineers, operations leaders, IT responders, and security teams turn a plan on paper into a tested capability.

Define safe states and manual fallback before an incident forces the decision. Build escalation paths, tabletop exercises, and scenarios that expose the gaps a document review cannot find.

Then turn the results into evidence: an after-action report, owned improvements, recovery checks, and a drill cadence that keeps the program alive after the exercise ends.

Explore the Course


ICS/OT Incident Response Drills and Exercises course preview

Run the drill before the incident runs you

Build an OT incident response program around safe state, manual fallback, escalation, realistic exercises, recovery checks, and audit-ready evidence. ICS/OT Incident Response Drills and Exercises gives operations, engineering, IT, and security teams a practical way to test the plan before the plant is under pressure. Academy purchases include a 30-day money-back guarantee when 25% or less of the course has been completed.

Explore the Course