Here is an uncomfortable observation. The business press is celebrating Samsung's reported fifteen-fold jump in validation efficiency on AI-run production lines, and SK Hynix rebuilding its KPIs around AI deployment, and almost nobody is asking the only question that matters on a shop floor: faster than what, measured by whom, on whose sample. I have seen this film before. I once audited a supplier whose scrap rate had halved in nine months. Immaculate trend. Full bins. The scrap had been reclassified as rework trials by a supervisor whose bonus was welded to that curve. Praise the engineering. Interrogate the metric. The moment a company attaches careers to a tool rather than an outcome, it has done something more consequential than deploying AI. It has given the data a motive.

Fifteen times faster than what

Take the claim apart the way you would a supplier's Cpk chart. Baseline first. Fifteen times faster than a manual validation process defined when, by whom, under what workload? If the reference is a badly documented process from five years ago, the honest multiplier is three, not fifteen. Then scope. Fifteen times on one high-runner product family is not fifteen times across the line, and a wafer fab enjoys process discipline most automotive plants can only envy – the variation a vision model sees on silicon bears little resemblance to a stamped bracket three weeks after tool maintenance.

Then the distinction no press release makes: validation efficiency is not validation sufficiency. In the AS9100 world, validation carries regulatory teeth; under IATF 16949, contractual weight. If an AI compresses a validation cycle from six weeks to two days, the honest question is what stopped being checked – how many boundary conditions, how many worst-case units, how much of the sample plan got optimised away. A velocity claim with no measurement system analysis behind it is not engineering data. It is marketing with decimal places.

Goodhart arrives on the shop floor

When the KPI is deployment, people protect the deployment. The failure modes are predictable, because they are the old ones. Anomalies the AI flags get recoded as operator error, since operator error does not threaten the programme. Escalation channels go quiet. QRQC sessions get shorter – not because problems disappeared, but because nobody brings the defect that has no category. Your nonconformance data does not get hacked. It gets negotiated.

I say this with scars, not theory. At SNOP, the 70% defect-cost reduction in a 900-strong plant never came from a smarter classifier. It came from QRQC data that operators, team leaders and shift managers trusted enough to report fast and honestly. The quarter we logged zero critical customer escalations was not the quarter with the best technology. It was the quarter in which bad news travelled faster than good, which is the entire discipline of de-escalation. That culture took two years to build. One badly wired KPI can unwind it in a quarter.

Building a greenfield QA/QC department for 900+ people from zero taught me the anatomy behind this. A KPI tree is the nervous system of a quality organisation; wire one nerve to a bonus and the whole organism learns to flinch around it. Samsung will manage – fabs have measurement science to spare. The exposure is downstream: the automotive tier-two and the aerospace supplier that copy the semiconductor scorecard without the semiconductor metrology.

A metric with a career attached stops being a measurement and becomes a negotiation.

Governance that keeps the data honest

The answer is not less AI. I design and deploy these systems myself. The answer is to treat an AI programme the way a competent quality organisation treats any new gauge on the line: with suspicion, calibration and an independent check.

Split the KPI in two. Deployment – lines covered, uptime, cycle-time gain – belongs to the programme. Detection – true anomalies caught, escape rate, time-to-escalation – belongs to quality, with a different owner and a different bonus pool, so the first cannot quietly feed the second. Audit the classifier like an instrument: repeatability, reproducibility, blind golden samples every month, an escape-rate trend reviewed like any other control chart. You would not accept a coordinate-measuring machine without an MSA. Do not accept a neural network without one.

Never let a single model's verdict gate the line. I built MultiPS – an orchestration platform running 63+ models in parallel with consensus synthesis – precisely because I refuse to let one model pass unchallenged. It is the same instinct behind the clickjacking disclosure that put my name on T-Mobile's bug-bounty Hall of Fame: complex systems fail at the seams where trust is assumed. On a production line that means independent consensus checks on every boundary case and human disposition whenever the models disagree. And keep one channel the programme cannot own: any operator or engineer can raise a nonconformance straight to quality, outside the programme's chain of command, and the reward system must make that bad news profitable to deliver.

Key takeaways

  • Split the scorecard: a deployment KPI owned by the AI programme, a detection KPI owned by quality – different owners, different bonus pools.
  • Audit the classifier like a gauge: repeatability, reproducibility, blind golden samples, an escape-rate trend reviewed monthly.
  • Never let one model's verdict gate the line – verify by consensus and route every disagreement to human disposition.
  • Keep a nonconformance channel that runs outside the programme's chain of command, and reward the bad news that arrives through it.

None of this is an argument against AI in manufacturing. Within two budget cycles, boards across automotive and aerospace will inherit the semiconductor scorecard, and some of the gains will be genuine. What fails silently is the measurement system wrapped around the tools: the baselines, the incentive wiring, the honesty of the nonconformance record. Say it now, while the programme is still a slide deck rather than a €4 million line item. Measure detection, not deployment. Reward bad news. Verify by consensus. Keep one reporting line the programme cannot touch. The AI was never the hard part. Keeping your defect data honest once careers hang on the tool – that was always the hard part, and it still is.