A vendor demo finds the defect in four seconds — on a dataset of 400 rows an intern labelled one Tuesday afternoon. I have since watched that same model meet a live nonconformance log where "scratch?" was a defect code, in three languages, across two shifts. It died within the hour. Nobody in the room was surprised except the vendor.

Rockwell Automation is touring the message that AI will remake manufacturing, and the trade press has performed the usual ritual — listicles comparing twenty-one platforms, feature matrices, ranking grids. Those lists tell me something their authors did not intend: factory AI has entered procurement season. Plants are ranking tools before anyone has audited the data those tools will consume. The shortlist is not the decision. The data underneath it is.

What the demo dataset has that your line doesn't

Demo datasets are engineered artefacts. Class-balanced, so the model meets defects at a comfortable rate instead of the one-in-four-hundred reality of a capable line. Cleanly labelled — one taxonomy, one language, one interpretation of "scratched". Pre-joined, MES and quality tables already reconciled on keys that line up.

Week one at a real plant looks different. Twelve defect codes with four meanings each, depending on which shift supervisor wrote them. Free-text root causes in shorthand only the author could decode. MES exports that will never reconcile with the QMS, because part numbers lost their leading zeros in a spreadsheet somewhere and the timestamps live in two time zones. Rework orders that never closed the nonconformance they came from.

When I built the QA/QC department at SNOP's greenfield plant — 900+ people, nothing inherited, nothing to unlearn — the first months of nonconformance records were honest and unusable in equal measure. Operators wrote what they saw, in Polish, French and English, sometimes inside one sentence. Shifts invented codes when the dropdown did not fit. I could not audit my own data, and I had signed it. The taxonomy work that followed — one defect language, one owner, closed loops — was the least glamorous project of my career, and it sat underneath everything after it, including the 70% defect-cost reduction the dashboards later took credit for.

Audit the data before you shortlist the tool

You already know how to do this. You do it to suppliers every quarter. Treat your own records like an incoming lot from a vendor you do not trust, sample them, and run three checks that need no software at all.

  • Taxonomy check. Pull 100 closed nonconformances at random and hand the codes to two quality engineers, separately. Where their readings diverge, the taxonomy is fiction.
  • Master data check. Count how many part numbers describe the same physical part. If the answer is not one, nothing downstream survives the join.
  • Join test. Trace 50 real orders end to end across MES, QMS and ERP. The share that reconciles is your true data-readiness score — not the number in last year's maturity assessment.

Three checks. One afternoon. Cheaper than any pilot.

Building MultiPS taught me the same lesson at machine scale. The platform runs 63+ models in parallel with consensus synthesis, and the ensemble reaches useful agreement only because the inputs are structured. Model count was never the variable. Feed five architectures a mislabelled, contradictory dataset and all five will nod at the same lie. Consensus on garbage is confident garbage — and confidence scales faster than accuracy.

A model inherits the plant it was fed: every shortcut in the record becomes a blind spot in the prediction.

Make the pilot run on your worst six months

If a pilot is worth running, run it on evidence. Write a proof-of-value clause into the contract: the vendor trains and validates on your ugliest records — the summer agency staff filled the log, the quarter a supplier change scrambled part numbering, the month the second shift booked everything as "other". Not their showcase dataset. Yours, at its worst.

If the model dies in week one of your free text, the ranking of twenty-one platforms collapses to a single number: zero.

And contract for data remediation, not licences. Dropdown discipline at the point of entry. One part master. Closed-loop nonconformances. These are old disciplines — standard work and poka-yoke applied to data instead of metal — and vendors will scope them as professional services if you insist, before the licence line item. I would rather spend €60,000 on a data audit and a taxonomy workshop than €600,000 on a platform that performs beautifully on someone else's records.

Key takeaways

  • Audit before you demo — a taxonomy check, a master data check and a 50-order join test across MES, QMS and ERP predict pilot success better than any feature matrix.
  • Make proof-of-value clauses bite: the vendor trains and validates on your worst six months of records, not a curated showcase set.
  • Put data remediation in the statement of work — dropdown discipline, one part master, closed-loop NCRs — and fund it before licences.
  • Rank your data readiness, not their feature list. Two plants running the same tool get different results because of what sits underneath the licence.

The winners of procurement season will not be the plants with the best tool. Tools converge; give the market eighteen months and every feature matrix will look the same. The winners will be the plants whose records could have survived an audit before the vendor arrived — running their new model on signal from day one while their competitors run it on confident noise. I watched a greenfield department learn that with paper and pencils, before the AI vendors ever came calling. Fix the record first. Then buy the robot.