Twelve US states. Water utilities. One week. The post-mortem distilled to five words: they could not see their assets. Not a zero-day. Not an advanced persistent threat with a memorable name. Configuration drift — settings that wandered from their known-good state until someone's exposed automatic tank gauge became an open door. I've walked enough shop floors to recognise this pattern in a different jacket. A press operator changes a parameter at 2 AM because the previous shift said it helped. Nobody logs it. The control plan says one thing; reality says another. By the time the customer finds the nonconformance at receiving inspection, you've shipped 4,000 bad parts. Same failure mode. The methodology to catch it was built in 1924.

Configuration drift is process drift wearing a network jacket

Walter Shewhart drew the first control chart at Western Electric to solve exactly this: gradual deviation from a known-good state that goes unnoticed until it produces a defect. Every process carries two kinds of variation — common cause and special cause — and you cannot separate them without a baseline. The control chart made drift visible. Nearly a century later, the OT security market is arriving at the same place with new vocabulary and a much higher price tag.

In automotive and aerospace, drift detection is woven into the standards. IATF 16949, VDA 6.3, AS9100 — each assumes you are continuously measuring against a baseline and reacting when the signal tells you something moved. You do not ship flight-critical hardware or safety-relevant automotive components on the assumption that nothing has changed since your last audit.

In operational technology security, drift detection is being treated like a product category launched last Thursday. Viakoo just announced a Device Configuration Manager for OT and IoT. Armis and Dispel are integrating around exposure analytics. New York State is committing $9 million to hardening municipal water systems. The market has discovered that devices on networks change over time and nobody is watching — a realisation quality engineering internalised a century ago.

No baseline means no detection

When I built the greenfield QA/QC department at SNOP for 900+ employees, the plant was a concrete shell. No desks, no instruments, no systems. I could have started with inspection gauges and calipers. I didn't. Step one was establishing the baseline: control plans, PPAP submissions, SPC charts on critical characteristics — before a single part shipped. If you don't know what "correct" looks like with numbers attached, you cannot detect "incorrect." You can only detect "different," and by the time you notice "different," you're already in containment mode with a customer on the phone.

The water utilities that got hit skipped this step. Their OT environments had grown organically over years — PLCs added, firewalls configured by whoever was on shift, vendor remote access left open because an integrator needed it six months ago and nobody revoked it. Their PFMEA equivalent didn't exist. Nobody had systematically asked what their failure modes were, what the severity looked like, what detection rating applied, what control prevented each one.

CEH-certified, listed on T-Mobile's bug-bounty Hall of Fame for a clickjacking disclosure. I've sat on both sides of this table — the quality floor where we chase 0.01 mm tolerance deviations on stamped metal, and the penetration-test debrief where we show a client their Modbus gateway has been internet-facing for eleven months. Different vocabulary. Identical failure pattern. The response discipline that quality engineering drilled into me is exactly what's missing in OT security.

Your quality manual already has the incident response plan

Here is what frustrates me watching the OT security market reinvent SPC: the methodology already exists, documented, proven, sitting in the quality manual of every Tier 1 supplier I've ever audited.

QRQC — Quick Response Quality Control — is a drift-response protocol. When a signal fires, whether a control chart breach, a customer complaint, or a spike in scrap rate, QRQC dictates: go to the gemba, contain within the hour, apply structured root-cause analysis, verify the countermeasure against the standard. I used this at Witte Automotive to drive substantial failure-cost reduction across multi-site operations. It maps directly onto what OT incident response should be — detect the anomaly, isolate the asset, root-cause the drift, remediate the configuration, verify the fix against baseline.

A3 problem-solving operates on the same logic under a Toyota template. 8D wraps it in Ford's eight-step format. PFMEA goes further — it is threat modelling done by people who treat failure as a professional obligation. You list failure modes, rate severity, occurrence, and detection, calculate risk priority numbers, act on the highest ones. Most cybersecurity threat-modelling frameworks are still catching up to that, and PFMEA has been at it since the 1990s.

The configuration-drift tools now reaching the market are reinventing control charts for networks. Automated baseline capture, continuous comparison, alerting on deviation. Useful technology. I would deploy it. But tools without the methodology are dashboards. What makes SPC effective is not the chart. It's the organisational discipline to act when the chart tells you something you'd rather not hear.

If you can't describe your known-good state with numbers, you don't have a security baseline — you have an opinion.

Key takeaways

  • Configuration drift in OT and process drift in manufacturing share the same failure mode — both demand a numerical baseline before detection is possible
  • SPC methodology — control charts, common-cause versus special-cause detection — ports directly to network and device configuration monitoring
  • QRQC, A3, and 8D map cleanly onto OT incident response: detect, contain, root-cause, remediate, verify against baseline
  • PFMEA is threat modelling that predates every cybersecurity framework — apply it to OT assets before deployment, not after the breach

Those water utilities got hacked because nobody could see their assets. Not because the attacker was brilliant. Not because the firewall failed. The baseline drifted and there was no control chart to catch it. Quality engineering solved this a hundred years ago, refined it through decades of automotive and aerospace standards, and embedded it in every supplier quality manual on the planet. The OT security world is welcome to the methodology. It's proven, documented, and sitting unused while vendors rebuild it from scratch — without the discipline, and with a much higher body count when it fails.