The ransomware crew that hit a hospital network in Canada did not spend its opening move on patient records. It stopped the doors. Elevators froze, ventilation and air-conditioning controllers went dark, and staff moved patients down stairwells because the building itself stopped cooperating. I read that and thought of the plants I have run quality in: paint shops where booth interlocks and fire suppression share a network with the MES; heat treatment lines where atmosphere setpoints live on a SCADA server nobody has rebooted since the last surveillance audit; pressurised processes where the relief logic is now a PLC output. 60% of manufacturers are racing toward full digitalisation, 30% of UK manufacturers have already taken a hit, and the average PFMEA was last revised before the OT network looked anything like it does now.
The occurrence rating you cannot estimate
Across two decades of PFMEA and APQP work under IATF 16949 and AS9100, I have scored more failure modes than I can count, and every occurrence rating draws from the same well: warranty data, process history, capability studies. Occurrence is history's opinion of a failure. That holds up as long as the failure has no opinion of you.
An adversary does. Adversarial failure is chosen – the attacker picks the furnace controller rather than a random actuator. It is adaptive: the first attempt fails, the second one doesn't. And it is correlated, because FMEA quietly assumes failure modes are independent while an attacker chains them deliberately. Random faults scatter. Intent compounds.
I also hold a Certified Ethical Hacker certification and a place on T-Mobile's public bug-bounty Hall of Fame for a clickjacking disclosure, so more than once I have put a line item reading safety interlock intentionally defeated in front of a PFMEA team. Severity takes thirty seconds: a 10, no debate. Occurrence stalls the room. Someone reaches for the warranty data and finds nothing, because warranty data records what broke, not what chose to break. Detection is worse. The scale assumes failures announce themselves, and the attacker's entire craft is announcing nothing. I have watched a team offer occurrence 1 on the grounds that nobody would bother with them. That is not an occurrence rating. That is hope, typed into a cell.
Failure is physics without intent. Attack is intent wearing the uniform of physics.
Fail-safe assumes the failure is honest
Everything elegant in our safety architecture – dual-channel controllers, dissimilar sensors, 1oo2 voting, defined fail-safe states – was engineered against a random, indifferent fault. Redundancy defends against coincidence. An attacker hunts the common mode: the shared engineering workstation, the flat VLAN, the integrator's remote-access account whose default password has survived three contract renewals.
The interlock is not collateral damage on the way to production. Often it is the objective. Defeat it and the process still runs – precisely the state you spent years designing to be impossible. Meanwhile your detection ratings assume a failed sensor throws a fault code; a manipulated sensor sits inside tolerance, patiently. The ransomware families named in the current advisories do not encrypt on arrival. They sit, exfiltrate, then extort, with dwell times measured in weeks. Our detection logic expects a fire alarm. They disable the alarm before lighting the fire.
Threat modelling belongs in the control plan
Here is the practical part: none of this requires new theory. It requires the discipline we already run, applied to a cause that plans ahead.
- Run an attack tree against your highest-severity process step. Start from booth interlock defeated or furnace overtemperature unrevealed and work backwards through remote access, the engineering workstation, the integrator VPN. Every node becomes either a control-plan line or a segmentation ticket with an owner and a date.
- Put OT security events into the nonconformance and 8D escalation path. A security event on a process controller is a product event until proven otherwise, and containment must fence shipped goods, not just restored servers.
- Segment safety networks from the MES before, not after. The architecture you need on day zero has to be bought, cabled and validated on day minus one.
I ran a tabletop exercise for a heat treatment supplier where IT declared recovery complete at hour 30, and quality asked which furnace lots from the dwell window were already on trucks. Nobody could answer. Years of customer de-escalation taught me the sequence never changes: what shipped, to whom, fence it, notify – then restore. The 8D containment discipline that carried me through critical escalations transfers to OT byte for byte, because customers and regulators do not care whether the nonconformity arrived as a bad batch or a bad packet.
Key takeaways
- When occurrence cannot be estimated from history, estimate exposure instead – remote-access paths, flat networks, integrator accounts – and treat OT-reachable safety functions as severity 9–10 by default.
- Attack trees on the highest-severity process steps should drive both the control plan and the network segmentation backlog, with owners and dates rather than a slide deck.
- Containment covers product: security events on process controllers enter the nonconformance system and 8D, with lot traceability and shipment fences, before anyone celebrates restored servers.
- Detection ratings built on honest failures need an adversarial twin – for every safeguard, ask what an attacker would do to keep it quiet, and score that.
Quality management and security engineering are the same discipline with different badges. Both study how complex systems fail and get paid to build them so they don't. The difference is that a growing subset of our failure modes now thinks back. The plants that fuse FMEA discipline with threat modelling will own the next decade of manufacturing. The rest will meet the problem anyway, the way you meet an auditor who arrives unannounced – except this one brings no checklist, rates nothing, and has already read your control plan.