A plant manager I have known for fifteen years called me on a Monday, and for the first ten minutes he sounded like a man who had won something. Ransomware had hit his tier-2 stamping house the week before – Medusa, by the note left on every desktop – and his IT provider restored everything from backup inside 72 hours. The lines never stopped. The ERP came back on Tuesday. Then his largest customer's auditor arrived for the surveillance visit and asked for last quarter's CMM results and the heat-treat certificates behind two part numbers. The files were not corrupted. They were gone. "Everything", it turned out, meant the systems – not the evidence that a year of shipments had been any good.
His week is about to become everyone's. The FBI and CISA, with health authorities, have updated their Medusa advisory: ransomware-as-a-service, rented out to affiliate crews who hunt volume and never learn what their victims make. A second number belongs beside it on the boardroom wall – roughly three in four ransomware victims are mid-market firms. Tier-2 and tier-3 manufacturers sit in the sweet spot. Enough revenue to pay, too little security to keep the crews out.
The real hostage is the evidence, not the machines
Plants survive ransomware better than offices assume. Machine tools keep cutting, inspection runs on paper for a shift or two, and the ERP is usually the first system recovered. What stays encrypted is the plain Windows share in the quality department – CMM exports, calibration certificates, control charts, PPAP packages, signed 8Ds. That share is the most reachable data on the network precisely because thirty people need write access to it. The affiliate who encrypted it neither knows nor cares what a PPAP is. They encrypt what they can reach. What they can reach is your proof.
Now count what that share underwrites. Every date code shipped, every escape closed with data, every record a customer-specific requirement obliges you to keep – fifteen-year retention clauses are routine in automotive, and aerospace is stricter still.
Production restarted. The proof never came back.
You can restore a server in 72 hours. You cannot restore proof.
Restoring servers is not restoring assurance
Disaster-recovery plans are written around two numbers: recovery time and recovery point. Both measure systems. Assurance runs on a different metric – when a customer names a lot number, can you produce the objective evidence within a day? A restore that brings back the QMS application but not its attachments, not the shares behind it, restores nothing a customer will accept. IT met every metric it was given. Nobody had defined quality's metric. Records, not servers.
Then there is the recovery-point gap. If the last clean backup is 24 hours old, 24 hours of shipped parts sit in what I call the window of unknowns – product in the field whose conformance you can no longer demonstrate. That is no longer an IT incident. It is one giant retrospective nonconformance, and exactly one function in the company is trained and authorised to disposition it.
The tools are the ones quality already owns. QRQC to establish facts fast: which date codes, which shifts, which record locations failed. 8D for the customer, with containment defined by lot and date, an escape boundary drawn honestly and notification triggered by evidence rather than embarrassment. Where retained samples exist, re-measure them. Where they do not, you are negotiating, not proving. I once closed a quarter with zero critical customer escalations, and the mechanism was unglamorous: data arrived before the customer asked for it. Delete the data and trust is the only currency left. A ransomware week makes you spend it fast.
Treat objective evidence like a production-critical asset
I built the quality function of a 900-plus-employee greenfield plant from an empty building. Traceability architecture. Retention matrix. Where CMM results land, how a calibration certificate binds to the gauge that released the part. I designed that record system deliberately, and here is the honest gap: I treated records as architecture, never as an asset with a threat model. No line in any budget protected them. The IT plan covered servers; evidence continuity appeared nowhere, because no standard asks the question and no customer audits it. Until the day it is all that matters.
My other hat is a certified ethical hacker's – T-Mobile's public bug-bounty hall of fame carries my name for a clickjacking disclosure – and from that side of the fence the attack needs no guessing. Automated enumeration finds the unpatched Windows box running CMM software the vendor abandoned in 2014. It finds the shared local administrator password and the share whose logging someone switched off for convenience. Affiliate crews do not hunt crown jewels in a vault. They take the filing cabinet in the corridor with the door propped open. In a mid-market manufacturer, that cabinet is the quality evidence share.
So the fix belongs in the quality system, not the IT budget. Quality owns the records; quality defines their criticality. That means an evidence recovery point measured in hours and immutable offline copies of calibration certificates and CMM raw data. It means customer-signed certificates kept in a form ransomware cannot reach, and an evidence-restore test built into the internal audit programme. You already test whether servers come back. Test whether proof does.
Key takeaways
- Build one register of every record type a customer can demand – CMM data, certificates, control charts, 8D, PPAP – and map where each physically lives. Any record whose address is "a Windows share" is a single point of failure.
- Set an evidence recovery point, not just a system one: decide how many hours of unprovable shipments the business tolerates, then protect records to that interval with immutable, offline copies.
- Audit the restore of proof, not servers: pull a shipped date code at random and demand the complete evidence package within 24 hours.
- Draft the disposition plan now – a QRQC/8D template for the window of unknowns, with containment rules and customer-notification triggers agreed before an incident, not during one.
Take one question to your next management review and insist the answer carries part numbers: if the QMS were encrypted tonight, which shipped parts could you still prove? If the honest answer is "whatever is on the paper traveller cards in a drawer", your real recovery time is not 72 hours. It is whatever it costs to re-earn, shipment by shipment, what the evidence used to establish for free. Evidence continuity is a quality responsibility. When the servers come back and the proof does not, the auditor will not be calling IT. They will be calling you.