Backups fail during a ransomware attack for reasons that have little to do with whether the backup job ran. They fail when the backup target can be reached from the compromised network, when no copy is immutable or offline, and when nobody has ever tested a restore. In this case all three met on the same Monday morning, two months after the server room had overheated.
This is an anonymised, composite case from my own client work; figures are rounded.
This continues Part 1, where a manufacturing client declined roughly €58,000 a year in security recommendations for three years running. This is what happened when leadership changed and the remaining safeguards started coming apart.
New Leadership, Old Blind Spots
Early 2023. The founder retired; his son took over as CEO. The old CFO left for another company, replaced by an aggressive cost-cutter who described himself as a "digital native" who "knew IT."
In his first sixty days, the new CFO cancelled the managed antivirus subscription (Windows Defender is free, isn't it, same thing) and let go of the one internal IT person, a 55-year-old who'd been there twenty years and cost too much for what he apparently did. The handover was informal, which is a polite way of saying there wasn't one.
Three weeks later, payroll went down. Nobody knew the database password; it lived in the departed employee's head. The company called him back for a three-hour emergency consultation at €2,400. The CFO paid it and was furious about paying it.
What nobody found until later: the departed admin had left a scheduled task, never documented, that renewed the company's critical certificates in the background. It ran under his own account, which nobody disabled. Ninety days after he left, that account's password expired and the task quietly stopped. The certificates it had last renewed were still valid for months, so nothing broke right away.
Does Turning Off the Server Room AC for a Weekend Really Save Money?
January 2024. Outside temperature: minus 8°C. The new CFO walked past the server room, heard the AC running, and asked the facilities manager why they were cooling servers in the middle of winter. He had it switched off for the weekend to save electricity. The facilities manager, who wasn't IT and had no reason to argue, complied.
The same wall unit this company had refused to replace for €3,500 back in Part 1 had been running on a failing compressor for months. Cycling it off and back on finished it.
Saturday, 2 a.m.: a monitoring alert. Server room temperature at 48°C and climbing. I drove in at 2:30, propped the door with a chair, grabbed every fan in the building, and spent the night watching the temperature come down a degree at a time. By morning, three servers had permanent damage. Two of them ran the ERP system. Production stopped Monday morning for a hundred and twenty people.
Emergency hardware replacement plus two days of lost production ran close to €228,000. The AC replacement that started this would have cost €3,500.
The 47 Minutes
Late February 2024. A salesperson clicked a link in a fake DHL delivery notice, on the Windows 7 laptop nobody had upgraded because it "still worked." The link dropped a loader, and for the next 23 days nothing visible happened. The attacker was inside the network, reading email, learning the org chart, picking targets.
Then, in the early hours of a Monday in March, the encryption started. It spread over SMB, and forty-seven minutes later every reachable file share, server, and workstation on the network was encrypted, including the backup server, because the backup server was domain-joined like everything else.
The Friday USB backup hadn't been plugged in for three weeks, and the replacement drive they'd bought after losing the first one had never been configured correctly. Its job ran and finished every time. It had been backing up an empty folder for weeks without anyone checking.
5:47 a.m., my phone rang: nobody could log in. I got to the office by 6:30 to a hundred and twenty people standing around a screen demanding a ransom of roughly €780,000, payable in bitcoin within seven days. Phobos ransomware.
Incident response cost €45,000 to bring in. The company's cyber insurance, downgraded during the cost-cutting the previous year, carried a €150,000 deductible. Forensics confirmed the 23 days the attacker had spent inside before triggering the encryption, and found one more thing: the certificate the fired admin's script used to renew had expired ten days before the attack, breaking the VPN's monitoring in a way nobody had caught.
They didn't pay the ransom. Decryption isn't guaranteed even if you do, the attacker had already taken data out, and paying just funds whoever hits the next company.
So Why Do Backups Fail Exactly When You Need Them?
In this case the backups still existed. What failed were two assumptions: that the backups were separate from what they were protecting, and that a backup job that runs is the same thing as a restore that works. Three properties decide whether a backup survives an attack like this one. They are separate controls, and having one of them gives you none of the others.
1. Network reachability of the backup target
Ransomware running with domain credentials can encrypt or delete anything it can reach over the network and authenticate to. Here the backup server was domain-joined and sat in the same network segment as production, so it was just one more machine on the list. Taking the backup server out of the production domain, giving it separate credentials and restricting network access to the backup traffic it actually needs shrinks that exposure. It does not make the backup copies themselves tamper-proof.
2. An immutable or air-gapped copy
That is the job of at least one copy that cannot be changed or deleted during its retention period, either because the storage enforces immutability or because the media is physically disconnected. The Friday USB drive was meant to be that offline copy. It hadn't been connected for three weeks, and its replacement was faithfully copying an empty folder.
3. A tested restore
A successful backup job tells you that the selected data was copied without errors. It does not tell you whether the right data was selected, whether the copy can be read back, or how long it takes to bring a server up from it. Only a restore answers that: restoring real systems or files into an isolated environment on a schedule and checking the result. Nobody here had ever run one, which is how an empty folder went unnoticed for weeks.
The single points of failure in this story (one aging compressor, one expiring certificate, one departed admin's undocumented script) failed in parallel and unnoticed, until the day something forced all of them to matter at once.
None of these failures happened in isolation. The cancelled antivirus, the AC nobody wanted to fix, the backup server sitting on the domain, the certificate script only one person understood: each was a separate decision, made by a different person, for a reason that made sense to them at the time. Part 3 covers what it cost to add all of them up.
Zero Trust. Zero Drama. Zero Bullshit.
When did your team last restore a complete server from backup into an isolated network, and how long did it take?
If you want to see what a restore test looks like in your own environment: drop me a note. Thirty minutes, I'll show you.
Talk it through with meFrequently Asked Questions
Why did the backup server get encrypted during the ransomware attack?
It was domain-joined, on the same network segment as production, trusted by the same directory. When the ransomware spread through the domain, the backup server was just another machine on it.
What is the difference between a successful backup job and a verified restore?
A successful backup job means the software copied the selected data and logged no errors. A verified restore means someone brought that data back into an isolated environment and confirmed it is complete and usable within an acceptable time. In this case the replacement USB drive completed its jobs for weeks while backing up an empty folder, which only a restore test would have exposed.
What is a dead man's switch in IT, and why do departing admins leave them?
It's a script or scheduled task that keeps running, or stops something, after an admin leaves. Some are deliberate; more often, like here, it's just an undocumented task nobody thought to hand over. The second kind is more dangerous, because nobody knows to look for it.
Should backup servers be domain-joined?
Not to the production domain. Common hardening guidance is to run the backup server in a workgroup or a separate management domain with its own credentials, so a compromised production domain doesn't hand the attacker the backups as well. Separately, keep at least one copy that is immutable or offline; the 3-2-1-1-0 variant of the 3-2-1 rule makes that explicit, together with zero errors in restore verification.
How fast can ransomware encrypt a network?
In this case, 47 minutes from the moment the encryption was triggered to a fully encrypted network, after the attacker had already spent 23 days inside. The exact time varies by ransomware family and network segmentation, but “we'll have hours to react” is not a safe assumption for any of them.