A defect root cause identification system will not make an AI reveal one “true cause.” That expectation is a common reason projects stall in Thai factories. Quality-data correlation can show which hypothesis deserves investigation next; it does not establish causation. A useful system creates a closed loop from customer protection and containment to product genealogy, process conditions, measurement context, hypothesis testing, correction, corrective action, and effectiveness review. This guide turns that loop into a practical RFP, 90-day proof of concept, FAT/SAT plan, backup design, TCO model, and handover package. Every duration, threshold, and cost marked as an illustrative project example is a design assumption—not a standard, law, BOI requirement, market price, quotation, customer result, or guarantee.
The output of a root-cause system is an evidence chain, not an answer
Immediately after a defect is detected, the priority is not to declare a cause. The plant must protect the customer and downstream processes, limit exposure, and preserve evidence. Divide the workflow into seven controlled stages:
- Detection: register the inspection result, alarm, or customer complaint with product ID, time, and location.
- Containment: hold suspected WIP, finished goods, shipped goods, and units exposed to comparable conditions.
- Hypotheses: create multiple candidates across material, machine, method, people, measurement, and environment.
- Evidence: compare genealogy, conditions, change points, and measurement context for conforming and nonconforming populations.
- Test: separate confounders and use an approved, technically safe test or repeatable observation.
- Correction and corrective action: distinguish fixing the detected nonconformity from removing a cause to prevent recurrence.
- Effectiveness: use the same scope and definitions after the change to monitor recurrence, related defects, and side effects.
The system should not merely display “temperature was the root cause.” It should show which population was compared under which definition version, which evidence supported and contradicted each hypothesis, who approved the test, what changed, and how effectiveness was reviewed. “Unable to confirm” must be a legitimate state. A mandatory drop-down that forces a cause creates organizational certainty before technical certainty.
ISO’s public status page identifies ISO 9001:2015 as Edition 5, published in 2015, confirmed in 2021, and accompanied by Amendment 1:2024. It remains current on 3 September 2026 but is expected to be replaced; this article does not mislabel it as a new 2026 edition. A 2016 ISO 9001 Auditing Practices Group paper separates correction, cause analysis, and corrective action, warns against calling the first factor found the root cause, and discusses objective evidence of implementation and effectiveness. The paper also says it is educational guidance without formal endorsement by ISO/TC 176 or IAF. Certification decisions must rely on the applicable normative text and competent advice.
Quality data correlation analysis ranks hypotheses—it does not prove causes
Suppose defective units had a higher mold temperature than conforming units. That observation is not equivalent to proving that temperature caused the defect. During the same window, the material lot, machine, shift, gage, ambient humidity, maintenance state, and setup method may also have changed. A calibration error or different sampling position can create an apparent difference with no process mechanism behind it.
Five questions that move an investigation from correlation toward causation
- Temporal order: did the candidate condition change before the defect, or was a post-inspection value mixed into the predictors?
- Comparability: are conforming and nonconforming groups comparable by product, tool, material, shift, and measurement method?
- Measurement reliability: are sensor, gage, master, calibration status, units, and missing-data treatment consistent?
- Alternative explanations: can another simultaneous change explain the same observation, including maintenance or recipe revision?
- Reproduction and refutation: does an approved change reproduce the phenomenon, and does reversing it remove the phenomenon? What contrary evidence exists?
Correlation coefficients, clusters, tree-based feature importance, AI summaries, and anomaly scores are all useful aids. None proves causality by itself. If a model ranks “material lot A” first, the team should connect receiving inspection, storage, drying, issue-to-line, machine settings, and conforming controls, then approve a test with Quality and Process Engineering. If data are sparse, missingness is biased, reproduction is unsafe, or interactions are likely, the system should preserve “possible” or “unconfirmed” rather than force a false conclusion.
Fishbone and 5-Why records also need evidence links
Sticky notes from a meeting are not reproducible records. For every hypothesis, retain:
| Field | Minimum record | Misuse prevented |
|---|---|---|
| Hypothesis ID | Unique ID and author | Silent overwriting |
| Candidate factor | Material, equipment, condition, measurement | Premature convergence |
| Supporting evidence | Events, trends, inspection records | Screenshot-only conclusions |
| Contrary evidence | Conforming or defective exceptions | Cherry-picking |
| Test method | Comparison, sample, approval, constraints | Unsafe or undefined testing |
| Status | Unconfirmed, possible, confirmed, rejected | Mixing probability with certainty |
| Next action | Containment, correction, corrective action | Missing owner and due date |

The minimum data model for linking manufacturing conditions to defects
A data lake full of signals is not useful if records cannot be returned to the affected product. A serial-number ledger without conditions is only genealogy. The minimum model links six objects: subject, event, condition, measurement, change, and action.
Do not confuse unit, lot, and container identity
Use unit IDs where individual characteristics matter, lot IDs for bulk material where lot resolution is adequate, and container IDs for logistics assets. Re-entry, split, merge, blend, and rework cannot be represented as a simple one-to-one line. Record parent-child relationships and quantities as events so the plant can answer both “what went into this finished unit?” and “where did this incoming lot go?”
GS1 EPCIS 2.0, ratified in June 2022, enables different applications to create and share visibility-event data within and across enterprises. GS1’s Global Traceability Standard 2.0 provides an interoperable traceability framework. They can help standardize identifiers and event exchange, especially across partners. They are not mandatory for every factory, and adoption does not prove that a process condition caused a defect. The RFP should separate externally shared events, confidential internal process parameters, and the identifier model.
Events require sequence and context, not just a timestamp
“13:05” could mean machining started, a setting changed, a measurement occurred, a disposition was made, or a unit was reintroduced. Store event ID, event type, product or lot ID, equipment ID, process step, event time, record time, source, recipe/version, operator or shift, reason code, and quality status. Record time helps expose late manual entry and communication delay. If device clocks drift, conditions from the preceding unit may be attached to the defective one.
In the illustrative project example, the selected line targets device-clock alignment within ±2 seconds, valid identity and genealogy for at least 98% of in-scope units, and completeness of at least 95% for defined critical conditions. These are not standard values. A ten-minute batch and a 200-millisecond cycle need different tolerances. During the PoC, measure cycle time, buffering, device clocks, and network delay before approving a tolerance.
Separate setpoints from actual values
A recipe setpoint of 200°C does not prove an actual temperature of 200°C. Treat setpoint, measured value, aggregation, alarm, and manual override as different fields. Depending on the mechanism, retain maximum, minimum, slope, time outside a range, and position within the process—not only an average. Unlimited signal collection also increases storage, engineering, and analysis cost. Begin with conditions linked to plausible defect mechanisms and enough controls to refute them.
For the collection architecture, see our localized guide to a manufacturing data collection system in Thailand. For bidirectional genealogy, use the trace-forward and trace-backward system guide. Approval and retention controls are covered in the quality-assurance record system guide.
Use ISA-95 and OPC UA to define boundaries, not as product labels
ISA’s public overview describes ISA-95, ANSI/ISA-95, and IEC 62264 as a standards family for enterprise-control integration. Its layers and interfaces can help assign the system of record for ERP material and shipment data, MES/MOM work and quality data, SCADA state, and PLC or sensor measurements. The objective is not to force everything into one table, but to define ownership and exchanges.
The OPC Foundation explains that OPC UA Companion Specifications publish industry-, device-, or use-case-specific information models. OPC UA supports domains from field devices to enterprise management and can expose real-time and historical variables, alarms, and object-based information. A suitable companion model, such as one for machine tools, can reduce one-off tag translation when both machine and vendor implement it. An “OPC UA capable” statement alone does not confirm the required nodes, units, timestamp quality, history, alarm semantics, or security configuration. Require an actual data list and an acceptance test.
In-process defect tracking begins with containment
An accurate investigation that arrives after more product ships is still too late. When a defect case opens, the system should search products sharing equipment, material, recipe, time window, tool, or gage and propose a suspect scope. Automatic holds must be risk-approved because an overbroad hold can stop production unnecessarily.
Do not put containment, correction, and corrective action behind one button
- Containment protects customers and downstream operations during investigation—for example segregation, added inspection, or shipment hold.
- Correction removes a detected nonconformity—for example rework, replacement, or sorting.
- Corrective action removes a cause to prevent recurrence—for example condition control, fixture modification, procedure, training, or maintenance change.
Successful sorting does not mean the cause disappeared. Conversely, changing settings too early can destroy reproducibility or damage another characteristic. Snapshot programs, recipes, parameters, calibration status, and work standards before a change. Separate emergency correction from permanent change control.
The suspect scope must be explainable, not merely maximal
Holding all inventory may appear conservative but can create avoidable sorting cost and delay. A narrow range may permit escape. Show search criteria, data completeness, missing records, time boundaries, and re-entry. Record why the scope was selected and which population was excluded. Products with missing evidence require an approved conservative rule; they must not be treated as conforming for analytical convenience.
A 90-day PoC from defect hypothesis to corrective action
The following is an illustrative project example covering one high-loss defect family, one product family, two process steps, one inspection step, and one line. Ninety days, 20 baseline days, 98%, 95%, three repetitions, and 30 post-change days are not universal acceptance criteria. Approve values based on risk, cycle, statistical plan, and customer requirements.
Days 0–30—freeze the boundary, baseline, and measurement context
- Select one historic defect and unify its definition, image examples, unit, and severity.
- Use an illustrative 20 operating days to baseline defect, scrap, rework, sorting, and investigation effort.
- Inventory product/lot, machine, tool, material, recipe, gage, and shift identifiers.
- Measure clock synchronization, network delay, missingness, late entry, and re-entry.
- List hypotheses and the evidence required; explicitly decide which signals not to collect.
- Set containment authority, change approval, and the cybersecurity boundary.
Deliverables are a data dictionary, genealogy map, event definitions, baseline, missing-data map, test plan, and risk register. Do not freeze dashboard cosmetics before these artifacts.
Days 31–60—validate data connections and hypothesis ranking
- Check the illustrative gates of at least 98% valid identity/genealogy and 95% critical-condition completeness.
- Stratify conforming and defective populations by product, equipment, material, and relevant context.
- Use distributions, correlations, change points, and time sequence to rank multiple hypotheses.
- Review measurement-system status, calibration, gage version, and decision-threshold changes.
- Traverse from source record to display and from display back to source.
- Run reproduction only after safety, quality, and production approval.
The example suggests three repeats only where technically and ethically appropriate. Three is not statistical proof. Destructive tests, high-risk processes, and customer product may require historical evidence, simulated material, a test rig, or expert review instead of deliberate reproduction.
Days 61–90—operate corrective action and effectiveness review
- Separate confirmed causes from unconfirmed contributors and issue a controlled change request.
- Record pre-change backup, rollback, approval, implementer, and version.
- Apply the change to a limited line or product, then monitor quality and side effects.
- Approve release from containment separately from root-cause confirmation.
- Use an illustrative 30 operating days or equivalent approved volume for effectiveness review.
- Decide continue, modify, stop, or scale at a gate review.
“Zero defects” is not enough when production volume is low. Include population, opportunities, defect-mode counts, process capability, inspection sensitivity, rework, downtime, and missing data. ISO 22400-1:2014 provides an industry-neutral framework for manufacturing-operations KPIs and remains current after confirmation in 2025. It does not prescribe the illustrative gates in this article.

Fifteen requirements for a defect root-cause analysis RFP
“Use AI to analyze causes” lets bidders assume different inputs, granularity, and responsibilities. Compare evidence, boundaries, and acceptance methods instead.
- Defect scope: definition, products, processes, lines, severity, exclusions.
- Identity and genealogy: unit, lot, container, split, blend, re-entry, rework.
- Sources: PLC, inspection equipment, machine, MES, ERP, LIMS, manual entry.
- Time quality: synchronization, tolerance, record time, delay, invalid clocks.
- Conditions: setpoint, actual, unit, sampling, aggregation, missingness, version.
- Measurement context: gage, fixture, calibration, inspector, decision threshold, retest.
- Hypothesis management: multiple candidates, support, contradictions, state, version, approval.
- Analytics: stratification, comparison, correlation, change point, explanation, recalculation.
- Causation boundary: no automatic root-cause confirmation; controlled verification workflow.
- Containment: scope search, hold, release, missing-evidence rule, customer-notification interface.
- CAPA: correction, corrective action, owner, due date, change control, effectiveness.
- Integration and ownership: API, OPC UA, event exchange, source of truth, export.
- OT security: accounts, roles, remote support, logs, segmentation, patching.
- Backup: objects, frequency, storage, encryption, restore test, RTO/RPO method.
- Acceptance and handover: FAT/SAT, performance, training, documents, source, settings, support, exit migration.
Require bidders to separate standard function, configuration, custom work, exclusion, and customer work, then map each to initial and annual cost. If a model changes, require training-data scope, feature list, version, approval, reproducibility, and impact on prior results. A black-box score belongs in hypothesis screening, not as the final quality disposition.
FAT/SAT should test round-trip evidence, not just screens
At FAT, use a known-answer data set to test genealogy, time drift, missing records, re-entry, rework, multiple hypotheses, access rights, audit logs, and backup. Freeze expected results. Include cases where an ID is unreadable, a machine clock is in the future, a lot splits, or an inspection result is corrected.
At SAT, use the Thai plant’s PLCs, inspection machines, network, terminals, Thai-language work practice, shift handover, and recovery from power or communications loss. A trend chart is not sufficient. Starting from a defect case, traverse to products, materials, conditions, measurements, hypotheses, and actions, then identify each source system and version. In the opposite direction, start from a material lot and find impacted products and their hold status.
Each acceptance record should contain requirement ID, test ID, input data, executor, date/time, software version, expected result, actual result, logs, evidence image, deviation, punch item, and retest. FAT acceptance does not prove site connectivity; SAT acceptance does not prove corrective-action effectiveness. Keep system acceptance, controlled production, and quality-effect review as separate gates.

Treat OT security and backup as quality functions
Evidence that can be altered, lost, or not restored cannot support a quality decision. NIST SP 800-82 Rev. 3, published in September 2023, guides OT security while considering performance, reliability, and safety requirements. Do not copy an office-IT restart or patch routine into a PLC network without examining assets, communications, availability, service access, and physical consequences.
NIST SP 1339, the OT Backup Quick Start Guide published in June 2026, says OT backups should be integrated with change management, created regularly, tested, and reviewed in recovery exercises. That is more actionable than “perform daily backup.” Define who captures PLC/HMI programs, recipes, gage configuration, gateways, certificates, clocks, tag dictionaries, models, roles, and audit settings; link each backup to a change ID; and state where restoration is possible.
A successful backup job does not prove recoverability. The illustrative design calls for a simulated restoration during FAT, an approved restoration target at SAT, and recurring exercises after go-live. Derive RTO and RPO from outage impact, manual-operation duration, tolerable genealogy loss, and replay capability—not a vendor default.
TCO includes connectivity and operating work
Initial quotations often show server and license while omitting tag engineering, machine modifications, clock synchronization, network segregation, data cleansing, master ownership, model revalidation, recovery exercises, training, and support. Divide TCO into eight boxes: software; connectivity; machine/network; data preparation; verification; operations; recurring maintenance; and exit/migration.
Illustrative project economics
The following numbers are an illustrative project example only. They are not market prices, a quotation, customer performance, guaranteed savings, or BOI-eligible costs.
| Illustrative item | Assumption |
|---|---|
| Avoided scrap and rework | THB 760,000/year |
| Avoided sorting and expedite | THB 360,000/year |
| Value of investigation hours | THB 180,000/year |
| Gross annual benefit | THB 1,300,000/year |
| Annual run cost | THB 290,000/year |
| Net annual benefit | THB 1,010,000/year |
| Initial implementation | THB 1,450,000 |
| Simple payback | About 17.2 months |
The calculation is 1,450,000 ÷ 1,010,000 × 12 ≈ 17.2 months. An avoided cost becomes financial value only when finance records confirm it. Freeze definitions for production quantity, defect opportunities, material value, sorting invoices, expedited freight, and investigation labor; avoid double counting. Suggested sensitivity cases are 70%, 100%, and 130% benefit, implementation cost +15%, and a three-month delay. They are not forecasts. Finding a cause creates no benefit if the corrective change is not implemented.
Thailand BOI’s current Smart and Sustainable Industry page describes support for efficiency enhancement and operational upgrades. The official guide linked from the current site is titled Investment Promotion Guide 2025 and includes categories for systematic information links and data analytics. This system is not automatically eligible. Confirm activity, existing/new project status, application timing, Thai-development conditions, eligible expenditure, and pre-approval purchasing against current announcements and BOI advice. The base business case should work without an assumed incentive.
Handover is complete when the Thai plant can reproduce the investigation
Handover is incomplete if the plant cannot open a new case, search genealogy, add a hypothesis, inspect source data, and restore configuration after the vendor leaves. Require the following:
- data dictionary and equipment/process/tag mapping with units and quality flags;
- genealogy rules for re-entry, split, blend, and rework;
- connectivity, certificates, accounts, roles, and remote-support approval;
- versioned export of software, configuration, models, and recipes;
- procedures for normal operation, missing data, hypothesis approval, CAPA, and abnormal recovery;
- backup scope, storage, restoration procedure, and restore evidence;
- FAT/SAT evidence, open punch list, warranty, SLA, and local contacts;
- training material and hands-on exercises understood by Thai users.
Training should use a simulated defect case, not only a screen tour. Operators capture identity and anomalies; Quality controls containment and hypotheses; Process Engineering compares conditions and tests; Maintenance owns machine data and restore; IT/OT owns accounts, communications, and logs. The group closes the same case ID end to end.
Common failures and design corrections
- AI score displayed as the cause: label it hypothesis ranking and require supporting/contrary evidence and approved testing.
- Different populations: stratify by product, equipment, material, shift, and measurement system.
- Only setpoints collected: separate actual values, time quality, alarms, and manual overrides.
- Clock drift ignored: measure synchronization, record time, delay, and tolerance during the PoC.
- Missing values converted to zero: retain a quality flag and use a conservative containment rule.
- Sorting treated as closure: separate containment, correction, analysis, corrective action, and effectiveness.
- FAT is a screen demo: use known answers and verify round trips to source.
- Backup is never restored: connect it to change control and exercise restoration.
- BOI benefit assumed: exclude it from the base case until project eligibility is confirmed.
- Vendor-only analytics: make data, settings, model, procedures, and training deliverables.
FAQ about selecting a root-cause system
Can a defect root-cause identification system work without AI?
Yes. With sound identity, genealogy, timestamps, conditions, measurement context, and change history, stratification, comparisons, trends, and SQL can narrow many hypotheses. AI can support screening and summaries; it cannot repair missing evidence or clock drift.
Can quality data correlation analysis confirm the cause?
No. Check temporal order, comparison groups, measurement reliability, confounders, reproduction, and refutation. Use correlation to decide the next investigation—not to auto-fill the CAPA root-cause field.
Where should linking manufacturing conditions to defects begin?
Start with one high-loss defect and one product family. Connect product/lot identity, process events, machine and recipe, critical conditions, and inspection results. Do not begin by collecting every tag from every machine.
Is MES mandatory for in-process defect tracking?
Not always. A PoC may integrate PLC, inspection equipment, databases, ERP, and a simple workflow. Still define the source of truth, ID creation, process order, re-entry, roles, and audit logs. Include export and API requirements for a future MES.
Will a 90-day PoC always reach a confirmed root cause?
No. Ninety days is an illustrative period to prove that the loop can operate. If the defect does not recur, reproduction is unsafe, or material cycles are long, use data completeness, suspect-scope retrieval, hypothesis quality, and test planning as gates.
What should FAT and SAT verify separately?
FAT should verify known-answer data, exceptions, permissions, backup, and expected results in the supplier environment. SAT should verify the actual machines, communications, clocks, terminals, local-language operation, and recovery from power or network loss at the Thai factory. Neither one replaces corrective-action effectiveness review.
What is the most important sentence in the RFP?
“An analytical result is a cause hypothesis and shall not become a confirmed root cause until linked objective evidence and an approved verification are complete.” Attach source, model version, contrary evidence, approver, and audit log.
How much does the system cost?
It depends on scope, machine modification, identity, existing systems, history, integration, validation, and security. THB 1,450,000 in this article is an illustrative assumption only. Compare initial, recurring, change, outage, and exit costs under one RFP.
Is database backup sufficient?
No. PLC/HMI programs, recipes, gage settings, gateways, certificates, clocks, tag dictionaries, models, roles, and audit configuration may also be required. Prove recovery with an approved restoration test.
Conclusion—buy a verifiable process, not a promised cause
The value of a defect root cause identification system is not a persuasive AI label. It is the evidence chain connecting containment, conforming and defective genealogy, process conditions, measurement, and change; the ability to test multiple hypotheses; the separation of correction from corrective action; and effectiveness review. Correlation is a powerful entrance to the investigation, not its causal exit.
Start with one high-loss defect and use a 90-day PoC to complete one cycle of data quality, suspect-scope search, hypothesis control, verification, CAPA, and restore. Put exceptions, timing, missing data, refutation, backup, ownership, Thai-language operation, and handover into the RFP and FAT/SAT. Replace example economics with plant data and treat BOI as a separate, eligibility-confirmed scenario.
TOMAS TECH can help before a product is selected—scoping the defect, assessing available evidence, designing the 90-day PoC, and preparing RFP, FAT/SAT, operating, and handover requirements. If you want to determine what your current data can genuinely prove, contact TOMAS TECH at the evaluation stage.
Sources
- ISO 9001:2015 status: https://www.iso.org/standard/62085.html
- ISO 9001 Auditing Practices Group guidance: https://committee.iso.org/files/live/sites/tc176/files/documents/ISO%209001%20Auditing%20Practices%20Group%20docs/Auditing%20General/APG-ReviewNonconformity2015.pdf
- GS1 EPCIS 2.0: https://ref.gs1.org/standards/epcis/
- GS1 Global Traceability Standard 2.0: https://ref.gs1.org/standards/global-traceability/2.0.0/
- ISA-95 overview: https://www.isa.org/standards-and-publications/isa-standards/isa-95-standard
- OPC UA Companion Specifications: https://opcfoundation.org/about/opc-technologies/opc-ua/ua-companion-specifications/
- OPC UA for Machine Tools: https://reference.opcfoundation.org/specs/OPC-40501-1/full
- ISO 22400-1:2014 status: https://www.iso.org/standard/56847.html
- NIST SP 800-82 Rev. 3: https://csrc.nist.gov/pubs/sp/800/82/r3/final
- NIST SP 1339: https://csrc.nist.gov/pubs/sp/1339/final
- Thailand BOI Smart and Sustainable Industry: https://www.boi.go.th/index.php?language=en&page=smart_sustainable
- Thailand BOI official guide: https://osos.boi.go.th/images/BOI_NEWS/2026/BOI_A_Guide_EN.pdf
*Sources checked 3 September 2026. This is general information, not quality-certification, statistical, legal, tax, investment-promotion, or cybersecurity advice. Confirm applicable standards, customer requirements, Thai law, and BOI conditions at decision time with official sources and qualified specialists.*