When planning an Industrial Data Fabric for manufacturing, do not define the first deliverable as “put all enterprise data in one place.” The operational problem is more specific. When equipment stops or a quality abnormality appears, SCADA tags, MES production records, ERP materials and orders, maintenance history, and quality decisions sit in separate systems. Engineers cannot reconstruct them as one episode. The first useful outcome is a governed, repeatable answer to one question: what changed before this stop, which lots may be affected, and have the same conditions occurred before?
On 10 September 2026, AWS published an Industrial Data Fabric example from an automotive paint shop at Mahindra & Mahindra. The architecture brings together SCADA, MES, downtime logs, energy management and engineering documents, contextualises sensor tags with assets, process stages and functional locations, and represents relationships among equipment, processes, failure modes and operating conditions in a graph. It is a timely implementation example, not a neutral standard or a guarantee of results. This guide turns the underlying lessons into vendor-neutral requirements for factories in Thailand and Southeast Asia, with particular attention to the tag dictionary, clock alignment, ownership, proof of concept, RFP and acceptance evidence.
Executive answer: procure a reproducible investigation path, not another data repository
The value of an Industrial Data Fabric is not measured by storage capacity or connector count. Make the following six outputs contractual:
- Episode spine: one stop or abnormality can be traced from onset and detection through response, recovery and quality impact under a stable ID.
- Canonical IDs and aliases: assets, tags, materials, work orders, lots and failure codes are mapped with versioned validity periods.
- Tag dictionary: meaning, unit, range, quality, sampling, time source and owner are explicit.
- Time policy: UTC, local display time, device time, event time, ingest time, corrections and late arrival are kept distinct.
- Ownership and change control: the people who approve meaning, protect connectivity and authorise use are named.
- Evidence pack: every result can be traced back through source records, transformations, versions, missing-data warnings and rerun output.
A proposal that promises only “cloud integration,” “AI root-cause analysis” or “a unified dashboard” cannot be accepted objectively. The supplier must prove how it handles a five-minute clock discrepancy, three different asset names, second-level SCADA events, lot-level quality records and day-level ERP transactions.
What an Industrial Data Fabric is—and what it is not
Industrial Data Fabric is neither one product nor an international standard. In this article it means an operating system of governed connections, identities, semantics, relationships, quality, ownership and delivery paths across distributed industrial data. A consumer should be able to use operational context without rebuilding bespoke joins for every investigation.
Separating adjacent concepts makes an RFP much clearer.
| Concept | Primary role | What remains missing on its own |
|---|---|---|
| Connector or ETL | Acquire, transform and move data | Asset/process meaning, ownership and continuing change control |
| Data lake or lakehouse | Store and analyse large volumes | Tag-to-lot relationships, event ordering and plant vocabulary |
| Historian or time-series store | Retain high-frequency equipment values and events | ERP orders, maintenance history and quality disposition |
| Unified Namespace | Publish plant events under a consistent namespace | All historical/enterprise data and contractual accountability |
| Knowledge graph | Represent relationships among assets, processes and events | Reliable source data, connectivity and clock discipline |
| Industrial Data Fabric | Govern connection, context, quality, ownership and delivery across stores | It still fails when no use case or acceptance criterion exists |
AWS, Ignition, HighByte, SiteWise and Neptune are possible components, not mandatory architecture. The same principles can be implemented with another cloud, on-premises services, an existing historian, a message broker and relational databases. Start with the information contracts and acceptance evidence before fixing the product list.
Start with one equipment-stop or quality-abnormality episode

“Visibility across every asset” is too broad for a first PoC. Select one line, one stop family and one quality-abnormality family. A paint process might choose oven-temperature deviation, conveyor stop, low pump pressure and coating-thickness failure. An assembly line might choose fastening-torque deviation, vision-inspection failure, conveyor jam and component shortage.
Define the facts required from each system to investigate one episode.
| System | Facts required | Typical join keys | Important traps |
|---|---|---|---|
| SCADA/historian | State, alarm, temperature, pressure, speed, current | asset ID, tag ID, event time | tag rename, unit, quality flag, gaps, sampling differences |
| MES | Order, operation, start/end, good/reject, lot, operator role | work order, operation, lot, asset | rework, manual entry, backdating, skipped operations |
| ERP | Material, BOM version, plan, supplier lot, summarised actual | material, order, batch | daily granularity versus plant events, master version |
| CMMS/maintenance | Failure code, work request, replacement, return to service | asset, maintenance order, failure mode | free text, aliases, delayed completion time |
| QMS/inspection | Specification, reading, disposition, hold, deviation, action | lot, serial, specification version | sampling, approval time, retest and revised result |
The goal is not to place five panels on one dashboard. Selecting a stop ID should reveal the asset, preceding condition changes, running product/order/lot, related quality disposition, maintenance action and post-recovery confirmation on one time axis, with a path back to each source record.
Make the event episode the spine of factory data integration
Connecting tables does not automatically accelerate investigation. Model the operational episode explicitly. A useful minimum set of objects includes:
Asset: line, cell, machine, instrument and fixtureProcessStep: operation, recipe phase and inspection pointMaterialLotorSerial: incoming lot, work in progress and finished unitProductionOrder: ERP/MES order and versionObservation: sensor value, quality flag and time sourceAlarmorStopEvent: onset, acknowledgement, clearance and recoveryMaintenanceAction: diagnosis, replacement, adjustment and trial runQualityResultorDeviation: specification version, measurement, hold and dispositionPersonOrRole: use a crew or role when individual identity is unnecessaryEvidence: raw record, transformation, rule version and query version
Every event should carry at least event_id, event_type, asset_id, event_time, ingest_time, source_system, source_record_id, quality_code and schema_version. If it contains process conditions, retain the engineering unit and applicable specification version. Minimise personal data and sensitive recipes and enforce purpose-based retention and access.
Store relationships instead of repeatedly guessing them
If every analyst infers “which tag belongs to which machine,” “which lot was present” and “which specification applied” in a new SQL query, answers will diverge. Treat the relationship itself as versioned data.
For example:
- tag
PT-204.PVrepresents discharge pressure for pump P-204 on Paint Line 2 during its stated validity period; - P-204 belongs to process step
TOPCOAT_SUPPLYand is associated with stop modeLOW_PRESSURE; - a specified MES order and lot passed during that interval;
- a coating-thickness result belongs to that lot and specification version; and
- a CMMS order records replacement of the P-204 seal.
A manufacturing knowledge graph helps explore these many-to-many relationships. Merely buying a graph database does not create context. Without canonical IDs, provenance, effective dates and accountable owners, it only preserves bad links more elegantly.
Turn the tag list into an enforceable data contract
A CSV containing tag names is not a sufficient tag dictionary. Require at least the following fields.
| Field | Required definition | Acceptance check |
|---|---|---|
| Canonical ID | Stable plant-wide identifier | History survives rename or relocation |
| Source alias | PLC, SCADA, historian and MES names | Searchable both ways with validity dates |
| Meaning | Operator-understood process meaning | No unexplained names such as TEMP1 |
| Data type and unit | Type, engineering unit, precision, conversion | °C/°F or bar/kPa cannot silently mix |
| Valid range | Physical, normal and alarm range | Sensor failure is distinct from process abnormality |
| Sampling and deadband | Frequency, change threshold, aggregation | Aggregation does not erase important peaks |
| Timestamp | Event/source/ingest time and clock source | Ordering policy is explicit |
| Quality code | good/bad/uncertain and missing reason | Missing values are never silently filled with zero |
| Owner and steward | Semantic approver, technical owner, support route | Every change has a destination |
| Version and effective date | Schema version, start and end | Historical analysis can use historical meaning |
AI can propose candidate meanings, mappings and duplicate aliases, but it cannot replace approval by the equipment and process owner. A wrong unit or asset mapping inside root-cause analysis produces a fluent but dangerous story. Automate candidate generation and difference detection; keep semantic approval accountable.
Clock synchronisation is a core acceptance item

Cross-system investigation does not require pretending every source owns a perfect clock. It requires an explainable meaning, accuracy, source, correction history and arrival behaviour for each timestamp.
OPC UA Part 6 states that clocks on communicating machines need reasonable synchronisation for certificate and certificate-revocation checks and that incorrect data and event timestamps may cause interoperability problems. It cites NTP as a standard approach and recommends logging possible time-synchronisation errors. Extend that guidance into practical plant rules:
- Store the common reference in UTC and convert to the site time zone for display.
- Keep
event_time,source_time,ingest_timeand, where necessary,corrected_timeseparate. - Inventory the clock sources for PLCs, IPCs, SCADA servers, historians, MES, databases and edge gateways.
- Monitor synchronisation status, offset, clock jumps, restart and daylight-saving configuration.
- During store-and-forward, preserve the original event time and sequence rather than only receipt order.
- Mark late, duplicate, reversed and future-dated events with quality information.
- Set tolerances by use case. Safety control, stop investigation and daily costing cannot share one generic number.
Thailand normally operates at UTC+7 without daylight saving, but headquarters, regional sites and cloud logs may use different zones. Saving display time alone creates ambiguity during incident review. Test device local time, database storage and browser/export rendering separately.
Example time acceptance tests
The thresholds are project examples; the plant must approve values based on process hazard, sampling, network and device capability.
| Test | Injection | Expected behaviour | Evidence |
|---|---|---|---|
| Clock offset | Move one source beyond the approved tolerance | Monitoring detects it and affected data is flagged | offset history, alert and affected records |
| Link outage | Disconnect for a defined interval, then reconnect | Original event time and sequence are replayed | sent/received count, sequence, gap report |
| Duplicate delivery | Resend one event ID | No double count; duplicate is auditable | deduplication log and final count |
| Out-of-order arrival | Deliver later event first | Episode uses event time and identifies late arrival | event/ingest time and rebuilt output |
| Time-zone mix | Mix UTC, ICT and another site time | Storage and display conversion remain consistent | input, stored value, UI and export |
| Restart | Restart PLC, IPC or gateway | Gap and recovery are explicit; no silent zero fill | uptime, quality code and gap report |
Assign ownership across OT and IT
An Industrial Data Fabric is neither an IT-only platform nor an OT-only connectivity project. Divide responsibility for meaning, availability, security and use.
| Role | Main responsibility | Approval scope |
|---|---|---|
| Process owner | Process/quality meaning and abnormality definition | KPI, event, specification and purpose |
| Asset owner/maintenance | Asset hierarchy, tags and failure modes | asset ID, alias and maintenance code |
| OT engineer | PLC/SCADA connection, load, network and safety boundary | collection method, rate and outage behaviour |
| MES/QMS owner | Meaning of orders, lots, inspections and rework | transactions and versions |
| ERP/master owner | System of record for material, BOM, orders and supplier | Level 4 ID and effective period |
| Data product owner | SLA, priority and users for the cross-system product | schema, quality and roadmap |
| Data steward | Vocabulary, aliases, quality issue and change history | dictionary and exception resolution |
| Platform/security | Platform, identity, access, monitoring and backup | non-functional controls |
| Consumer owner | Validity of BI, analytics, AI and operational use | decision process and model use |
“The business owns data” is too vague. Name the decision owner. A data steward may correct an English label, but the process/asset owner approves what a pressure tag means. Production planning or master-data ownership approves whether an ERP material equals an MES part. A new downtime classification changes historical KPI comparability, so its impact must be reviewed by the data product owner.
Treat change management as normal operation
Tag addition, equipment modification, PLC upgrade, material creation, MES change and maintenance-code consolidation happen continually. A clean migration drifts within months unless the change path is designed. Each request should record target ID, before/after value, reason, effective time, affected data products, backward compatibility, reprocessing need, tests and approvers.
Where a manufacturing knowledge graph adds value
A graph is useful for questions that are awkward in repeated key joins: machines using the same pump model; operating conditions historically associated with a failure mode; assets traversed by a nonconforming lot; maintenance immediately preceding an abnormality; or combinations of supplier lot and recipe version.
Do not force all data into the graph. High-frequency waveforms can stay in a time-series store, transactions and quality results in relational/lakehouse storage, documents in object storage, and navigational relationships in the graph. The graph becomes a context layer connecting place, process, asset, lot, failure, specification and document—not a replacement for source systems.
Graph acceptance criteria
- Every relationship has provenance, generation rule, version and effective period.
- Inferred relationships are distinguishable from approved relationships.
- Users can navigate back to the source record.
- Negative tests prove that unsupported relationships are not generated.
- Historical relationships remain reproducible after relocation, rename or material consolidation.
- Relationship exploration cannot reveal restricted recipes or personal data to unauthorised users.
A 12-week PoC that completes one investigation
Twelve weeks is an illustrative planning frame, not a delivery promise. Legacy PLC constraints, network approval, shutdown windows and regulatory review can extend it. The gates matter more than the calendar.
Weeks 1–2: fix the question, scope and baseline
- Select one line, one stop family and one quality-abnormality family.
- Choose a representative set of historical cases available at the plant.
- Measure current investigation time, systems visited, manual steps and missing data.
- Put decision owners and source owners into a RACI.
- Confirm separation from safety control and classify sensitive data.
Do not impose one sample count on every process. Rare high-consequence failures may require a small case study; frequent micro-stops need enough episodes to expose bias.
Weeks 3–4: build source contracts and the tag dictionary
Define owner, method, load limit, collection rate, replay, retention, ID, schema, quality and test data for every source. Walk down representative tags against physical equipment instead of trusting screen labels. Resolve ERP-to-MES order mapping, lot split/merge, rework and maintenance free text.
Weeks 5–7: implement the contextualisation pipeline
Separate raw, standardised, contextualised and served zones. Version every transformation. Preserve event and ingest time, deduplicate, and process late arrivals. Use canonical IDs and an alias table rather than making a spreadsheet mapping an invisible production dependency.
Weeks 8–9: deliver the investigation workbench
Provide a UI or API that moves from stop ID to timeline, order, lot, quality and maintenance evidence. Provenance, quality, timestamp and version matter more than visual polish. If AI summaries are added, require source links, time range and missing-data warnings; never let the model declare a root cause without accountable review.
Weeks 10–11: run negative and reproducibility tests
Inject source loss, stale aliases, clock offset, wrong unit, missing data, duplicate, reversed order, late MES entry, revised QMS disposition and unauthorised access. Verify prevention of false conclusions, not only correct happy-path output. Prove that the same snapshot and rule version reproduce the same episode result.
Week 12: make an evidence-based scale decision
The gate should allow continue, conditional continue, remediate and retest, reduce scope, or stop. Do not scale merely because the demonstration worked. Missing ownership, unsustainable dictionary maintenance and poor source quality are valid reasons to pause.
What the RFP must require
1. Use case and exclusions
State which stop and quality episode must be reconstructed. Explicitly exclude, where appropriate, real-time control, safety-PLC changes, ERP replacement and enterprise-wide master consolidation. Without exclusions, proposals differ mainly in hidden assumptions.
2. Source-by-source connection design
For SCADA, historian, MES, ERP, CMMS, QMS, spreadsheets and documents, require protocol, direction, read/write, rate, bandwidth, outage/replay, authentication, audit, environment and responsibility. Start PoC access read-only; never grant generic OT write access for convenience.
3. Canonical model and alias service
Require canonical IDs, aliases, validity periods, split/merge behaviour, version and owner for assets, materials, lots, orders, process steps, tags, failures and quality specifications. ISA-95 Part 7 offers a useful alias-service concept, but the plant must define its own identity governance.
4. Data quality and time
Require measured rules for completeness, duplicate, uniqueness, range, unit, freshness, late arrival, order, time synchronisation and quality code. The customer approves thresholds by use case.
5. Lineage and replay
Demand a path from source record through ingestion, transformation, dictionary version, query, API, screen and exported result. Corrections must show impact and permit controlled reprocessing. Manual corrections retain actor, reason and approval.
6. Security and the OT boundary
Specify network zones, direction, service identities, least privilege, secrets, encryption, patching, vulnerability response, logs, backup and incident handling. State that the data fabric never replaces PLC, SIS or deterministic machine interlocks.
7. Availability and degraded mode
Define local buffering, capacity, replay order, deduplication, recovery targets and loss notification when cloud or WAN service is unavailable. Production control must remain independent, and unavailable analytics functions must be visible.
8. Operations and handover
Contract source onboarding, tag change, dictionary approval, user management, monitoring, incident response, restore, cost monitoring and supplier-exit export. Deliver code, configuration, models, mappings, tests, runbooks and known limitations.
Acceptance criteria: evaluate evidence, not feature checkboxes

Agree acceptance criteria before the PoC starts. Replace the sample thresholds below with plant-approved values.
| Acceptance item | Test | Required evidence | Failure example |
|---|---|---|---|
| Source completeness | Reconcile a period against each source | source-level count and gap table | only total count matches; gap location unknown |
| ID and alias | Exercise rename, relocation, split and merge | time-bounded mapping and approval | historical data is overwritten with the current name |
| Unit and schema | Inject unit, type and version mismatch | conversion, rejection or quarantine log | silent coercion |
| Event ordering | Inject late, duplicate, reversed and future events | event/ingest time and reconstructed result | causality shown by receipt order only |
| Cross-system trace | Follow a sample stop across every source | episode report with evidence links | manual spreadsheet required |
| Quality impact | Match lots and inspections around the stop | reasoned included/excluded set | every lot produced that day is marked affected |
| Lineage | Navigate a displayed value to source | source ID, rule version and query | calculation cannot be explained |
| Access control | Attempt out-of-role access | denial log and access matrix | URL knowledge bypasses control |
| Recovery | Stop and restore gateway, WAN and processing | gap, buffer, replay and observed recovery | duplicate or silent loss after recovery |
| Reproducibility | Rerun one snapshot/rule version | checksum and result comparison | unexplained result drift |
Do not accept only an average accuracy rate. Joining the wrong asset, mixing units, reversing event order or exposing restricted data is a critical defect even when most records pass. Define severity-based exit criteria, corrective action, retest and residual-risk approval.
How to compare vendor proposals
| Evaluation axis | Question | Evidence of a strong proposal |
|---|---|---|
| Use-case fit | Which decision becomes faster or more reliable? | concrete episode and current baseline |
| Brownfield connectivity | How are old assets and shutdown limits handled? | walkdown, load test, read-only path and buffer |
| Context model | Who maintains IDs and relationships? | owner, version and effective date |
| Data quality | How are gaps, units and late data exposed? | quality flags and quarantine, not concealment |
| Openness | Can data and mappings move to another platform? | documented schemas, APIs and export |
| Security | What reaches OT and under which identity? | clear zone, direction, identity and audit |
| Operations | Who handles change and failure? | specific runbook, SLA, RACI and training |
| Acceptance | What proves success? | contracted negative tests and evidence pack |
Compare total cost across connectors, site survey, mapping, data-quality remediation, network, platform usage, monitoring, training, change and exit—not only licence cost. Prefer sustainably operated data products and reproducible investigations over a large connector count.
Preserve the roles of ERP, MES and control systems
The fabric does not replace MES or ERP. ERP remains the record for enterprise planning, transactions, cost and inventory; MES for execution, orders, production results and traceability; SCADA/PLC for supervision and control; QMS for specifications and disposition; CMMS for maintenance work. The fabric connects the required context and delivers it to investigation, analytics and applications under governance.
For the wider organisational and architecture boundary, see OT/IT Convergence in Manufacturing: Three Barriers and a Practical Path. To define MES scope, machine connectivity, ERP boundaries and acceptance, see MES Implementation for Thailand Factories: Go/No-Go. This article is deliberately narrower: it specifies how one stop or quality abnormality is contextualised, procured and accepted across systems.
Common failure patterns
- Collect everything first: no decision owner means quality remains invisible while storage cost grows.
- Join asset names as strings: abbreviations, languages, rename and relocation break history.
- Overwrite time into one column: source, event, receipt and correction can no longer be distinguished.
- Fill missing data with zero: missing values look normal and corrupt analytics.
- Make graph technology the objective: ownerless relationships become a new source of error.
- Use administrator access for the PoC: production permissions, audit and network limits remain untested.
- Lead with AI summarisation: weak context generates a fluent false causal story.
- Postpone operating ownership: one tag change requires the supplier and the dictionary becomes stale.
- Accept on an average score: critical mismapping or leakage disappears inside the average.
FAQ: implementing an Industrial Data Fabric
How is an Industrial Data Fabric different from a data lake?
A data lake mainly provides storage and analytics. A fabric adds governed connection, identity, semantics, quality, ownership, access, lineage and delivery across multiple stores and systems. A lake may be one component, but it is not a fabric by itself.
Is a manufacturing knowledge graph mandatory?
No. It is valuable when relationships among assets, processes, lots, failures and specifications are complex. Relational models are sufficient for simpler domains. Canonical identity, provenance, effective dates, ownership and source lineage matter more than graph technology.
Which line should be selected for the PoC?
Choose a line with real stop or quality-investigation work, multiple systems to reconcile and cooperative owners. The newest line is not automatically best. Select a scope with measurable value where read-only access keeps production and safety risk controlled.
Is NTP installation enough for time synchronisation?
No. Also design clock sources, offset monitoring, event/source/ingest time, late arrival, ordering, time zones and correction history. Set tolerances by use case and test them.
Which master data should be fixed first?
Start with the assets, tags, process steps, materials, orders, lots, failures and quality specifications required by the selected episode. Define the system of record, owner and validity period. Do not turn the PoC into an enterprise master-data replacement.
When should natural-language search or AI be added?
After contextualised evidence, access control, quality warnings and a validated question set exist. Begin with search and summarisation, require citations and abstention, and do not let AI autonomously certify root cause or control equipment.
Must the architecture be cloud based?
No. Use edge, on-premises and cloud according to network, residence, latency, availability, skills and cost. Production control must stay independent during communication loss, and replay, gaps and access must remain explainable.
Which deliverables should be fixed in the RFP?
Fix the boundary diagram, source inventory, canonical model, alias table, tag dictionary, time policy, quality rules, lineage, access matrix, RACI, test specification, evidence pack, runbook and exit/export procedure—not only a product name and feature list.
Summary: begin by explaining one episode end to end
A successful Industrial Data Fabric for manufacturing is not an attempt to finish collecting all enterprise data. It is the ability to explain one equipment stop or quality abnormality by connecting SCADA conditions, MES orders and lots, ERP materials and plans, CMMS maintenance, and QMS disposition without losing the meaning of time and identity. A user can return to evidence and reproduce the result.
Build the episode spine, canonical ID and aliases, tag dictionary, time policy, ownership and lineage first. Limit the PoC to one line, one stop family and one quality family. Test not only the successful demonstration but also missing data, clock offset, duplicates, late arrival, wrong aliases and unauthorised access. In the RFP, procure information contracts, change operations, negative tests, handover and evidence rather than a fashionable platform label.
TOMAS TECH supports manufacturers in Thailand and Southeast Asia with site walkdowns, SCADA/MES/ERP/maintenance/quality boundaries, tag dictionaries, PoCs, supplier comparison, RFPs and acceptance tests. You can contact us while evaluating architecture or when you want to prove feasibility with one stop episode before selecting a platform.
Primary references
- AWS for Industries, “Reducing paint shop downtime with Industrial Data Fabric on AWS,” 10 Sep 2026: https://aws.amazon.com/blogs/industries/reducing-paint-shop-downtime-with-industrial-data-fabric-on-aws/
- AWS Solutions Guidance, “Industrial Data Fabric with HighByte Intelligence Hub on AWS”: https://docs.aws.amazon.com/solutions/industrial-data-fabric-with-highbyte-intelligence-hub-on-aws/
- AWS Solutions Guidance, “Industrial Data Fabric with Ignition on AWS”: https://docs.aws.amazon.com/solutions/industrial-data-fabric-with-ignition-on-aws/
- AWS IoT SiteWise User Guide: https://docs.aws.amazon.com/iot-sitewise/latest/userguide/
- ISA, “ISA-95 Standard: Enterprise-Control System Integration”: https://www.isa.org/standards-and-publications/isa-standards/isa-95-standard
- OPC Foundation, OPC UA Part 1, Overview and Concepts: https://reference.opcfoundation.org/specs/OPC-10000-1/4
- OPC Foundation, OPC UA Part 6, Time synchronization: https://reference.opcfoundation.org/specs/OPC-10000-6/6.3
- AWS Well-Architected Framework, “Modern Industrial Data Technology Lens”: https://docs.aws.amazon.com/pdfs/wellarchitected/latest/modern-industrial-data-technology-lens/modern-industrial-data-technology-lens.pdf
This article is general implementation and procurement guidance based on public primary sources checked through 19 September 2026. It does not replace equipment-specific safety assessment, legal compliance, quality assurance or cybersecurity certification.