A survey crossed my desk this week: most UK manufacturers feel "ready" for AI-led change. I have never seen a readiness score drop after a good demo. The instrument measures appetite, and appetite has never contained a single part. The same week's trade press carries a vendor essay promising a "missing intelligence layer" that MOM platforms will drape over our shop floors. Layers are cheap. What is expensive is 06:10 on a Tuesday, when the layer starts telling the truth faster than your organisation can absorb it.
I have lived the alternative. At SNOP, a greenfield automotive plant built for more than 900 people, I created the QA/QC department from bare concrete up. The first thing I wrote, before the first production part and years before anyone said "AI" in that hall, was the defect taxonomy and the reaction-plan architecture. Codes, owners, clocks, escalation ladders. QRQC and A3 programmes later ran on that spine and cut defect cost by 70%. No neural network anywhere. Detection was never the constraint. Deciding what counted as a defect, and who reacted within the hour, was.
Readiness is a feeling until the first shift it isn't
There are two kinds of readiness and only one of them ships product. Deployment readiness is models, GPUs, dashboards, MES integration, a demo that lands well in a boardroom. Absorption readiness is what happens to the findings: who triages them, how fast, into what quarantine, with what authority to stop a line. Surveys measure the first because the first is easy to feel. The second shows itself only under load, on the first shift when detection outruns reaction.
I hit the same wall building MultiPS, my own orchestration platform that runs 63-plus AI models in parallel with consensus synthesis. The bottleneck is never inference. Compute is a purchasing decision. Routing the disagreements to a named owner with a clock running — that is organisational design, and nobody demos it.
AI will not raise your standards. It will publish them.
The four prerequisites no survey asks about
Four things must exist before a detection system earns the right to alarm. I have audited plants that had all four on paper and none of them on the floor. Start with a defect taxonomy a model can actually learn: mutually exclusive codes, operational definitions, stable across shifts. If two inspectors code the same flaw differently on the same part, the model does not learn your process. It learns your disagreement, then amplifies it at machine speed. Then comes master data you would stake a recall on — images and signals tied to part number, batch, cavity, machine and timestamp. Not to a shared folder named "misc". A finding you cannot trace to a lot is an opinion.
Third, reaction plans with surge capacity: QRQC inside the hour, 8D for whatever leaks past it, quarantine space sized for three times today's volume, an MRB that can run hot without borrowing the quality engineer off a launch. And fourth, a triage cadence that survives a bad week — named deciders, defined clocks, an escalation ladder that still works when three alarms fire in the same week as a customer audit.
None of it ships with the platform.
When detection triples, containment becomes the constraint
Every plant whose quality system I have built or audited carries a defect base its current standards were quietly calibrated to tolerate. Visual standards, gauge thresholds, sampling plans — all of them encode the escape rate the customer will absorb, whether anyone admits it or not. A working detection model does not tolerate. It surfaces the base. Detection triples and the constraint moves overnight from the model to your containment fence, your MRB throughput and your supplier quality inbox. Every reopened PPAP drags capacity studies and re-qualification behind it like an anchor chain. When machine-vision AI performs brilliantly in one plant and flops in the next, that is not a technology gap. It is an absorption gap.
This is why you stage an AI rollout exactly the way you stage a launch. Shadow mode first: one shift, thirty days, no enforcement. Measure quarantine volume, MRB hours, reaction-time adherence. Two decades of scars say you never put a new part on three shifts at once. A system that triples your daily nonconformance count is a launch. The containment fence, not the licence fee, is where the euros go in the first year.
Run the test before you sign the purchase order
Pull thirty days of your own nonconformances — not the vendor's case study — and check them against your own reaction-plan clocks. What share was reacted to within the hour, within the shift? Count how many sat open past a week, and who owned them. Then double the count and ask whether Tuesday survives. If the organisation cannot absorb today's truth at today's volume, more detection is not readiness. It is a confession.
Key takeaways
- Demand shadow mode, not a demo: one shift, thirty days, no enforcement — and measure quarantine volume, MRB hours and reaction-time adherence before scaling a single line further.
- Audit the defect taxonomy first. If two inspectors code the same flaw differently, the model will not learn your process — it will learn your disagreement.
- Your real readiness score sits in your own data: reaction-plan clocks checked against the last thirty days of nonconformances, not a survey of how people feel about change.
- Stage the rollout like a launch: shift by shift, with quarantine space and MRB surge capacity agreed in writing before the first alert fires.
So the observation stands. The headlines measure ambition while absorption goes unmeasured, and that gap is where deployments fail. The ready plant is not the one with the best model. It is the one whose Tuesday can absorb what the model found on Monday.