Blog

2026.09.27

Manufacturing AI Data Readiness: A Pre-Deployment Guide for Thai Factories

Manufacturing AI Data Readiness: A Pre-Deployment Guide for Thai Factories

A factory may have years of ERP transactions, MES records, PLC signals, inspection results and maintenance notes, yet still be unready for a useful AI evaluation. The issue is rarely just “not enough data.” Records may use different equipment codes, timestamps may not line up, a quality label may mean different things across shifts, or a vendor may be unable to explain which source supported an answer. Manufacturing AI data readiness means checking whether the information needed for one operational decision is defined, linkable, dependable, authorized and auditable—and deciding what people must still verify.

This guide is for Japanese manufacturers operating in Thailand and ASEAN that are moving from an AI idea toward vendor selection, an RFP or a pilot. It focuses on the work before model selection: defining the decision, inspecting the data, setting access boundaries, and designing a 90-day evaluation that produces evidence for an investment decision. The proposed readiness framework and schedule are TOMAS TECH practical proposals, not an official standard, certification or guaranteed result.

Manufacturing AI Data Readiness: A Pre-Deployment Guide for Thai Factories - figure 1

Start with a decision, not a model

“Use AI to reduce defects” is not yet a testable project. Which inspection step is in scope: incoming material, in-process checks or final inspection? Who uses the result? What action follows a flag—reinspection, machine check, lot hold or a process adjustment? How quickly must the information arrive? What is the consequence of a false alarm or a missed defect, and who can override the system?

Write one target decision in a sentence. Then map how it is made today: who collects which records, how exceptions are checked, what is approved, and where the decision is recorded. Experienced operators often rely on context that is missing from formal systems. Ask them what they look for and when they distrust a report. Do not translate their answer into database columns too early; first capture the situation, evidence and consequence.

A useful scope statement names the plant, line, product family, time period, users and exclusions. It might say that a pilot will help a production supervisor prepare a morning review of delayed orders on one line, using approved MES and ERP records, while keeping the final schedule decision with the supervisor. It should also say whether the evaluation is retrospective or connected to live data. A narrow scope is not a commitment to stay narrow; it makes the first test interpretable.

Six checks for manufacturing AI data readiness

The following six dimensions are a TOMAS TECH working framework. They are prompts for a joint review by operations, IT, equipment engineering and security—not a compliance checklist that replaces laws, customer requirements or company policies.

1. Meaning and ownership

A field name does not guarantee a shared definition. Confirm whether “quantity produced” means good pieces or total output, whether downtime includes planned stops, and whether a defect date is the event time or the entry time. Document units, time zone, rounding, code values and what an empty field means. Assign an owner who can approve definition changes. If the meaning changes, record when it changed and which reports or integrations are affected.

Review master data for products, equipment, lines, process steps, lots and shifts. Names are poor join keys: spelling differs, the same machine may have multiple aliases, and codes can be reused after a system migration. If no stable identifier exists, define a controlled cross-reference with an owner, effective dates and change history. For a deeper primer, see master data to organize before improving a factory with AI.

2. Coverage and missingness

A blank is not a single condition. The value may be unknown, not measured, not applicable or forgotten. Missingness may cluster during a night shift, a particular product changeover or equipment downtime. Measure coverage by line, product, shift and relevant operating state, then ask why records are absent. A high overall completion rate can conceal a gap exactly where the use case needs evidence.

Decide whether the evaluation period covers normal production as well as important exceptions: changeovers, maintenance, restart after a stop or a new product introduction. More history is not always better. Old data may refer to a different sensor, process recipe or quality rule. State which periods are representative and which should be isolated. If coverage is limited, constrain the use case or collect more data; do not quietly label the dataset “ready.”

3. Accuracy, consistency and time

Check plausible ranges, units, duplicate records, state transitions and sudden jumps. If a machine is marked stopped while its production counter rises, possible causes include delayed signals, counter resets or duplicate processing. Ask equipment and operations staff what the anomaly means before automatically deleting or correcting it. The unusual values may be the event the AI is supposed to detect.

Compare a representative sample with source records and shop-floor knowledge. Determine which system is authoritative for each business fact. Align timestamps across PLC, MES, ERP and local files, and document the time zone. Measure delay against the decision window: a delay of several minutes may be acceptable for a daily review but unsuitable for a fast response. Retain correction history when the source process supports it.

4. Identity and integration

An AI answer about a particular machine or lot depends on reliable joins. Map identifiers across ERP order, MES operation, PLC tag, inspection record and maintenance log. Specify how the mapping handles equipment moves, product-code changes, split lots and rework. When an identifier is incomplete, label that limitation instead of silently matching by a display name.

An API connection is not the whole integration plan. Define read frequency, retries after network loss, duplicate handling, late data, sequence changes, monitoring and manual fallback. Establish whether the pilot reads data or writes back. Any write path to production equipment needs a separate safety and change review; a readiness assessment alone does not authorize control actions.

5. Access and protection

Document the data flow: where information starts, which service receives it, where it is stored, who can access it, how long logs remain, and how data is deleted. Check the company’s privacy, cybersecurity, customer-contract and cross-border requirements with the responsible specialists. Do not assume a dataset is anonymous simply because names were removed; combinations of shift, role, timestamp and event can still identify people.

Carry the source system’s access boundaries into AI retrieval and output. A user who cannot open a restricted quality record should not receive its details through a generated summary. Review role changes, vendor accounts, authentication, audit logs, prompt logging, retention and model-provider reuse. Confirm in writing whether submitted data may be used to train or improve a provider’s models and how an opt-out is enforced.

6. Provenance and auditability

For a result to be reviewable, retain suitable evidence of its source, extraction time, transformations, model or configuration version, references returned, and human review. This does not mean keeping every intermediate file forever. Agree on evidence and retention periods based on business need and policy. If a record is corrected, know whether an earlier result can still be explained.

For natural-language search or retrieval-augmented generation, test whether an answer points to the source record, reflects freshness, and admits when evidence is insufficient. For detection or classification, pair predictions with a human disposition and, when available, later-confirmed cause. Record who created the evaluation labels and how disagreements were resolved. For an adjacent example of AI supporting management work in Thailand, see AI assistants for managers in Thai manufacturing.

A practical data inventory for one use case

A data inventory should connect an operational event to the dataset used for evaluation. Draw the path from a sensor, inspection form or production transaction through PLC/SCADA, MES, ERP, a database or approved file, and into the evaluation environment. Label each connection with its owner, method, refresh frequency, delay, access condition and behavior during an outage. Include manual spreadsheets when they function as the current record; do not hide them because they are inconvenient.

For each required field, record its business definition, type, unit, sample value, source, expected coverage, permissible delay, missing-value treatment, owner, retention and access condition. Separate must-have evidence from helpful context. If a field is missing, decide whether a safe manual collection can temporarily support a limited test, whether a different decision can be evaluated, or whether a system change is essential. These are different remedies and should not be collapsed into a single “data quality” score.

Review samples with the people who create and use the records. Include more than a clean production day: select relevant shifts, product changes, equipment stops and quality exceptions. There is no universal sample count that proves readiness. The period and volume should reflect the decision, event frequency and risk. Record how samples were selected, which were excluded, and any manual cleanup so that a vendor demo cannot be mistaken for live operating performance.

Classify each item as usable now, usable with a stated limitation, blocked pending remediation, or not yet assessed. Attach evidence, business impact, owner and a review date. A readiness score may help organize discussion, but it should never let several low-risk items mathematically offset a critical gap in traceability, access control or safety. Display blockers separately.

What to put in an AI RFP

Ask suppliers to respond to the same data flow, sample, use case and acceptance conditions. Request a clear boundary between customer, integrator and provider responsibilities.

Data and connectivity: Which sources are in scope? Is access read-only? What protocol, polling interval, network change, plant downtime or local component is needed? How are missing, duplicated, late or out-of-order records detected? Who owns transformation rules and cross-reference tables? What happens during an outage, and how does the operator return to the existing process?

Evaluation and evidence: What is the test set, who defines correct answers, and how are tuning and final evaluation separated? Which metrics fit the business impact—precision, recall, missed events, false alarms, response time or review workload? Can results be inspected by product, line, shift and period? How are unsupported questions, stale inputs and uncertainty presented? Are source records visible to the user?

A headline such as “99% accuracy” is not comparable without the sample, class balance, period, threshold and exclusions. A classifier can appear highly accurate by calling every rare defect normal. Agree on the relative cost of missed events and false alarms with the process owner, and preserve the evaluation conditions alongside the number.

Security and operations: Where is data processed and stored? What encryption, access logs, backup, deletion and subcontractor controls apply? Are prompts or records used for provider model improvement? How are identities, roles, account expiry, incidents, recovery and version changes managed? Who handles a data-quality incident versus a model issue? What support hours and recovery expectations are included?

Handover and contract: Who owns the data dictionary, mappings, integration code, evaluation set, operating procedure and configuration created during the pilot? What are the exit, deletion confirmation and migration terms? What changes in scope or fees when adding a line, site or system? Clarify confidentiality, incident notification, subcontracting and reuse rights with the relevant legal team. Do not assume the AI supplier automatically owns the customer’s data governance responsibilities.

For Thai operations connected to a Japanese headquarters or a regional cloud, ask the local plant and corporate security teams to review the same data-flow diagram. Whether a transfer is allowed depends on the data, contract, service and applicable requirements. “Cloud is safe” and “on-premises is safe” are not useful substitutes for a specific control review.

A 90-day evaluation plan—with decision gates

Ninety days is a planning example from TOMAS TECH, not a standard duration or promise. If access approval, representative labels or connectivity take longer, extend or split the evaluation. The objective is a defensible decision, not a launch date at any cost.

Weeks 1–2: agree on the decision and test. Define users, boundaries, baseline, exclusions, owners and approval roles. Decide whether the evaluation is retrospective or live. Set acceptance conditions for source traceability, access control, freshness, human review and fallback, as well as appropriate model measures. If a reliable baseline is unavailable, document the limitation instead of inventing a savings figure.

Weeks 3–4: inventory and sample. Build the flow diagram and field inventory, request least-privilege access, and compare selected records with the source and shop-floor understanding. Log missing links, inconsistent codes, delay, privacy constraints and unresolved meanings. If essential data cannot be approved or no one can establish correct labels, pause or change scope before building a demo.

Weeks 5–8: test in a bounded environment. Connect only approved data and run agreed scenarios. Technical owners check stability, freshness, retries and logs. Operators compare outputs with source evidence and record false alarms, misses, exceptions and effort to review. Keep tuning and final evaluation separate where possible; check that the test does not leak future information or put related lots in both sets. A favorable result does not prove performance outside the tested conditions.

Weeks 9–12: rehearse operations and decide. Practice reporting an error, correcting a source, suspending the AI and returning to the approved manual workflow. Present verified results, limits, open risks, required remediation and next-stage costs. Options include a controlled expansion, better master data, a different use case, a longer collection period or stopping. Name the decision owner, next review date and evidence required for expansion.

Manufacturing AI Data Readiness: A Pre-Deployment Guide for Thai Factories - figure 2

Thailand and ASEAN context: use the numbers carefully

Thailand’s investment environment is relevant context, but public investment announcements do not establish the business case for a particular factory. In 2026, Thailand’s Board of Investment reported approval of nine projects worth USD 1.99 billion across high-value sectors including AI and advanced electronics. The announcement included GPU-server infrastructure and energy planning. It describes investment approvals, not measured AI productivity at manufacturing plants.

An NSTDA and BOI update dated 31 July 2026 reported 2,062 Industry 4.0-related applications worth THB 206.054 billion through May, and 17 projects worth THB 1.033 billion approved on 18 June. These are program-related figures, not an AI-only adoption rate. BOI’s Smart and Sustainable Industry measure specifies conditions for eligible investment promotion; eligibility and tax treatment depend on the current official rules and project facts. Confirm those details directly with BOI or qualified advisers.

A 2026 American Economic Association study draws on about 28,500 U.S. establishments and reports AI use as of 2021. Its findings can prompt questions about infrastructure, processes and expertise, but they should not be presented as a Thailand adoption estimate. NIST’s Manufacturing Extension Partnership lists data availability and quality, integration, privacy and cybersecurity, and skills among issues in U.S. manufacturing support. Its examples, including a U.S. company case, are not Thai prevalence data. AWS describes data fragmentation across systems such as PLM, ERP, MES and supply chain applications; this is a vendor technical perspective, not an independent survey.

For Japanese-owned ASEAN operations, agree which definitions, controls and security rules are corporate-wide and which need a local mapping. Thai, Japanese and English forms may use different terms for the same event—or the same term for different events. Keep the stable key separate from display labels, record the responsible owner, and test how shift, plant and time-zone context travels through the system. Bring operations, local IT, equipment engineering, headquarters security and the supplier into the same design review.

Incentives should not drive an AI scope beyond what the business can operate. Confirm thresholds, eligible activities, implementation deadlines and other conditions in the current official notice. Investment promotion and data readiness are distinct questions: tax eligibility does not prove that a dataset is reliable, that access is appropriate or that an AI workflow creates value.

Manufacturing AI Data Readiness: A Pre-Deployment Guide for Thai Factories - figure 3

Common mistakes to avoid

Building a lake before choosing a decision. Broad ingestion increases cost and governance work while leaving unclear who will use the data. Start from one decision and expand the reusable parts only after ownership and access are clear.

Treating a cleaned demo file as production-ready. A CSV prepared by an analyst may omit outages, delays, access rules and operational exceptions. Document every transformation and test the approved source path, fallback and freshness.

Contracting on one accuracy number. Use case-specific measures and retain the test conditions, class mix, period and exclusions. Include workload and consequences, not just model metrics.

Letting the AI replace undocumented judgment. Ask experienced staff to explain the checks they perform. Keep a human confirmation path, an override, and a way to record corrections and later findings.

Using a dataset without an owner. Assign ownership for definitions, mapping tables, access review, change notices and retirement. A file that continues to work is not necessarily governed.

Expanding the purpose silently. A pilot for daily summaries is not automatically approved for quality release, personnel evaluation or customer decisions. A changed purpose requires a new risk and data review.

Pre-deployment checklist

For each item, mark yes, partial, no or unknown, then assign an owner and due date to gaps. This is a practical self-review, not a certification or legal determination.

  • The operational decision, user, timing and action after an output are defined.
  • Product, line, equipment, period, exclusions and human approver are named.
  • Required fields have a definition, unit, time basis, owner and missing-value meaning.
  • Cross-system identifiers and mapping owners are documented.
  • Samples include relevant shifts and exceptions, with selection and exclusions recorded.
  • Missingness, duplication, delay, inconsistent codes and corrections are understood.
  • Purpose, location, users, retention and deletion have been reviewed.
  • AI output permissions preserve the source system’s access boundaries.
  • There is a plan to follow a result back to its source record.
  • False alarms, missed events and unanswerable cases have a human process.
  • Network or AI outage has an approved manual fallback.
  • Incidents, model changes, data changes and support ownership are assigned.
  • Pilot assets, data deletion and exit or migration terms are clear.
  • The expansion decision, participants and evidence are scheduled.

A set of “yes” answers is not a substitute for resolving a safety, legal or security blocker. Readiness is about matching the use case to known evidence and controlled limitations, not getting a high average score.

Conclusion: make data and responsibility visible before AI

Manufacturing AI data readiness is the work of connecting an operational decision to reliable evidence, accountable owners and an appropriate human review. Define one use case, inspect its data meaning and lineage, test the joins and access boundaries, then put the limits into the RFP and evaluation plan. A small, evidence-led test can reveal what is usable now and what needs work before expansion.

Thailand’s investment activity and overseas research help explain the environment, but neither substitutes for evidence from your own process. Use them as context, keep U.S. findings labeled as U.S. findings, and confirm incentives from current official sources. If your Thailand or ASEAN team is deciding which data to inventory, how to set ERP/MES/equipment boundaries, or how to shape an RFP and pilot, contact TOMAS TECH to discuss the current stage and target process.

Sources

Related articles