Blog

2026.09.03

Anomaly Detection Sensor Selection for Thai Factories: PoC to FAT/SAT

Anomaly Detection Sensor Selection for Thai Factories: PoC to FAT/SAT

When a Thai factory starts an anomaly-detection sensor project, discussion often jumps to vibration versus temperature, wireless hardware, or a target “AI accuracy.” The first procurement decision should instead be which failure mode must be observed, through which physical change, under which operating conditions, and what a named person will do after detection. This guide turns temperature sensor monitoring, vibration-based equipment diagnostics, false-alert control, a 30–90 day PoC, an RFP, and FAT/SAT into practical acceptance requirements. It is deliberately a sensor-selection and acceptance guide—not a generic predictive-maintenance overview.

Conclusion first: procure a failure-mode-to-action measurement system

Success is not determined by a sensor datasheet or one model score. The RFP should procure an end-to-end measurement and response system in which:

  1. Assets and failure modes are prioritized and bounded.
  2. Observable signals that lead or accompany each failure are selected.
  3. Range, bandwidth, mounting location and direction, protection, and calibration are reproducible.
  4. Speed, load, process step, recipe, ambient condition, startup and shutdown are recorded on the same time base.
  5. missing data, drift, clipping, clock error, and communications loss are not mistaken for equipment anomalies.
  6. Advisory notifications, operator alarms, and safety interlocks have explicit boundaries and owners.
  7. A truth ledger supports continuing review of false positives and false negatives.
  8. FAT and SAT can replay the same inputs, expected outcomes, and evidence.

“Predict equipment failure” must never become a promise to forecast every failure. A system can detect only changes observable with the selected sensors and acquisition settings. Sudden fracture, a fault outside the instrumented area, or a change hidden by normal variability can be missed. Evaluate proposals on named failure modes, observable scope, exclusions, and the downstream work process—not merely the presence of AI.

1. Define failure modes before selecting sensors

The opening workshop should not attempt to instrument the whole asset register. Narrow the scope using production impact, quality impact, safety and environmental consequence, redundancy, and field observability. “Monitor the pump” is not a testable requirement. Bearing degradation, imbalance, misalignment, cavitation, seal leakage, blockage, and motor overload produce different signals and require different actions.

For each failure mode, state what changes physically, when it tends to appear, whether PLC or inspection data already observes it, whether action can wait for planned downtime, and which corroborating signal is needed. The earliest signal and the most diagnostic signal may differ. Broadband vibration can expose change early but is sensitive to load and mounting. Temperature is intuitive but can lag because of thermal mass and must be interpreted with ambient and process heat.

Failure mode to signal and sensor selection table

Failure mode or stateObservable changePrimary sensor candidateSampling and installation caveatRequired context
Rotating imbalanceRunning-speed component; amplitude and phase changeAcceleration or velocity vibrationFix direction on a rigid bearing location; match bandwidth to speedSpeed, load, operating mode
Misalignment or loosenessAxial, harmonic, or impact changeTriaxial or directional vibrationDocument point and orientation; do not equate temporary magnetic and permanent mountingTime after startup, coupling state, maintenance history
Bearing degradationHigh-frequency impacts, envelope trend, heatingHigh-bandwidth accelerometer plus temperatureAllow enough record length at low speed; test mounting resonance and saturationSpeed, lubrication, load, replacement date
Cavitation or fluid upsetBroadband vibration or sound; unstable pressure and flowVibration/acoustic plus pressure and flowDo not decide from ambient sound alone; synchronize suction and valve conditionsFlow, pressure, fluid temperature, valve position
Motor overload or electrical faultCurrent, power, power factor, and temperature changeCurrent/power plus temperatureVerify VFD waveform suitability and electrical-work safetyFrequency, torque command, load, phase values
Poor lubricationFriction-related high frequency and temperature riseVibration/ultrasound plus temperatureTreat the temporary change after greasing as a separate modeGrease amount/type/time, runtime
Gear damageMesh-frequency and sideband changeHigh-bandwidth vibrationVariable speed may require order-related analysis; verify transmission pathShaft speeds, load, gear ratio
Oven/dryer thermal upsetDeviation, heat-up rate, or spatial distribution changeThermocouple, RTD, or infraredVerify emissivity, field of view, thermal contact, and open-circuit diagnosisProduct, setpoint, door state, ambient temperature
Cooling degradationInlet/outlet temperature difference, pressure, or flow changeTemperature plus pressure/flowUse differential measures, and align sensor response timesLoad, coolant temperature, filter state
Compressed-air leakageUltrasound, pressure decline, increased compressor operationUltrasound plus pressure/powerSeparate location finding from system-total monitoring; segment production noiseProduction state, pressure setpoint, running units
Conveyor jam or rising frictionMotor current, speed difference, temperature, vibrationCurrent plus speed/temperatureModel normal variation from product mass and product typeProduct, throughput, speed, jam history
Sensor or wiring faultFlat line, spikes, noise, missing data, driftSelf-diagnostics and redundant referenceUse a data-quality flag separate from asset anomalyCalibration, communications, power, environment

This is the skeleton of a requirement, not a catalogue of models. For every candidate, also state when it cannot represent the target. A surface temperature probe exposed to airflow and poorly coupled to the target may be stable yet not represent internal temperature. An accelerometer on a thin cover may mostly measure the cover’s resonance rather than bearing condition.

Anomaly Detection Sensor Selection for Thai Factories: PoC to FAT/SAT - figure 1

2. Temperature, vibration, electrical, acoustic, and process signals have different jobs

Where temperature sensor monitoring works—and where it does not

Temperature is readily understood for overheating, cooling capacity, heat balance, oven distribution, bearings, and electrical cabinets. Existing PLC-connected RTDs or thermocouples may be historized without a new sensing network. A single absolute threshold, however, can be distorted by ambient temperature, product recipe, load, and warm-up time.

Consider absolute level together with deviation from baseline, inlet-outlet difference, peer-machine difference, rate of rise, or residual at a comparable load. The RFP should identify sensor type, thermowell, insertion depth, response time, lead compensation, open-circuit behavior, and calibration. Infrared measurement adds emissivity, reflection, field of view, viewing-window, and contamination requirements. Measurements produced by dissimilar installations must not be treated as interchangeable “temperature.”

Where vibration-based equipment diagnostics works—and where it does not

Vibration can reveal imbalance, misalignment, looseness, and bearing or gear changes, but “vibration” is not one number. Acceleration, velocity, displacement, waveform, spectrum, envelope, and phase serve different purposes. ISO 20816-1:2016 remains the current published general guideline for measuring and evaluating machine vibration. ISO also marks it for revision and says replacement by ISO/FDIS 20816-1 is expected. The draft must not be treated as a final published standard. FAT/SAT criteria therefore need a versioned published reference plus the applicable machine class, measurement point, quantity, bandwidth, and operating condition.

Low-speed, variable-speed, and intermittent machines may not suit the same fixed time window or frequency bins as a constant-speed machine. Synchronize a speed tag and consider mode segmentation or order-related analysis. For wireless sensors, verify whether the device sends only onboard summary features, whether raw waveforms can be retrieved around an event, and the actual bandwidth, duration, and repeatability.

Electrical, acoustic, and process context

Current and power can monitor load change, jamming, or no-load operation from the cabinet, but an identical current rise might come from product mass, speed setting, or mechanical friction. Acoustic and ultrasound sensing can help find leaks and impacts but requires baselines for production noise, adjacent machinery, and air blow. Pressure, flow, speed, and quality values provide evidence that links equipment condition to process consequence.

A small combination of slower temperature/process values, fit-for-purpose vibration, and speed/load often explains an event better than one premium sensor. The decision criterion is not the channel count. It is whether the chosen measurement can observe the failure and provide evidence that field staff can verify.

For safe signal acquisition on legacy machinery, see IoT retrofit for aging equipment. For edge protocols, buffering, and clock behavior, use the industrial IoT gateway selection guide.

3. Build baselines by operating mode, not merely by asset

More normal data does not create a better baseline if unlike conditions are mixed. Stopped, startup, warm-up, steady operation, changeover, cleaning, setup, low load, high load, manual operation, and post-maintenance states have different distributions. If the detector learns them as one normal population, thresholds become too wide and miss degradation, or every transition becomes a false alarm.

At minimum, synchronize these tags with sensor time series:

  • Asset, measurement-point, sensor, mounting direction, and configuration-version IDs.
  • PLC state, running/stopped, process step, recipe, and product.
  • Speed, load, flow, pressure, and setpoint.
  • Ambient temperature, shift, weekday, and elapsed warm-up time.
  • Maintenance start/end, part change, lubrication, calibration, and remounting.
  • Alert, field inspection, work order, confirmed fault, false-alert decision, and closure.

Do not prescribe that “30 days is enough.” A machine that runs at high load once a week may still have few high-load examples after 30 days. A repetitive machine may deliver useful comparisons sooner. Define PoC coverage in terms of modes, repetitions, load range, products, and maintenance events as well as calendar days.

Time synchronization is measurement quality

Without aligned clocks, investigators cannot tell whether load changed before vibration rose or output declined after temperature increased. Specify the time source, time zone, allowed offset, and correction history for sensors, gateways, PLC, SCADA, MES, and CMMS. If an offline device buffers locally, retain original event time when reconnecting rather than overwriting it with upload time.

“OPC UA compatible” is also insufficient. OPC Foundation’s public specification describes authentication of client/server applications and users, communication integrity and confidentiality, and selectable profiles for installation needs. Request supported profiles, certificate operations, signing/encryption, clock behavior, reconnect and subscription-loss behavior, and audit logs. In line with NIST SP 800-82 Rev. 3, SAT must also verify that monitoring additions respect OT performance, reliability, and safety and do not create an unnecessary control-network route or load.

4. Manage false positives and false negatives in one truth ledger

Raising a threshold to reduce false positives can increase false negatives; increasing sensitivity can burden the field with noise. A one-sided KPI can reward a detector that produces nothing or a notification stream that staff ignore. Evaluate both alert-level and fault/event-level outcomes.

The truth ledger should record alert ID, asset, time, operating mode, signal, rule/model version, severity, evidence plot, owner, field inspection, work order, component condition, cause, disposition, and closure time. Misses cannot be found by reviewing alerts alone. Reverse-check breakdown, maintenance, quality, and downtime records to determine whether earlier detection existed.

Anomaly Detection Sensor Selection for Thai Factories: PoC to FAT/SAT - figure 2

Practical controls before sophisticated modeling

  1. Segment stopped, startup, cleaning, and setup modes; suppress only where action is not intended.
  2. Use speed-, load-, or ambient-adjusted baselines and residuals.
  3. Define persistence, repeat count, trend, and hysteresis rather than a single crossing.
  4. Attach corroborating evidence instead of deciding from temperature or vibration alone.
  5. Route missing data, low battery, saturation, and flat line as data-quality notifications.
  6. Manage post-maintenance and remounting as baseline-change periods.
  7. Group multiple tags caused by one condition at asset/cause level.

Do not accept a vendor default threshold as truth. Base thresholds on OEM guidance, an applicable published standard, historic normal and fault evidence, asset criticality, and feasible field action. Asset owner, maintenance, and production should approve them. When a rule or model changes, preserve the old-versus-new result on a frozen validation set and an audit trail of activation.

5. Separate advisory, alarm, and safety interlock

An anomaly score must not silently become a trip command. The three layers have different purpose, recipient, response time, and validation burden.

ClassPrimary purposeExpected actionImplementation and acceptance boundary
Advisory/notificationReview trend, propose inspection, prioritize planned workMaintenance reviews by a due date and opens work if justifiedCMMS, email, or dashboard; track receipt, assignment, and closure
Operator alarmA timely operator response is required for an abnormal conditionAcknowledge, diagnose, execute defined action, escalateGovern priority, cause, consequence, response, suppression, and audit under the alarm philosophy
Safety interlock/tripIndependent protection prevents an intolerable hazardValidated logic performs the defined safe actionSubject to safety requirements, independence, verification, and change control; never replaced by an analytics notification

ISA’s public ISA-18 series page treats alarm management as a lifecycle and emphasizes meaningful, prioritized, actionable alarms, while distinguishing non-alarm notifications. Sending every anomaly to DCS/SCADA can flood operators with messages that have no operator action. Conversely, a maintenance advisory still needs an owner and due date.

A PoC analytics score should not be wired directly into a safety stop. A change to a protective function belongs to the site’s safety lifecycle, risk assessment, verification, authorization, and change-management process. Start anomaly detection as decision support and a maintenance workflow. Treat any later move into automatic control as a separate safety and control design.

6. Structure a 30–90 day PoC around gates, not elapsed time

Thirty to ninety days is a planning envelope, not a performance guarantee. No failure during the period does not automatically mean failure of the PoC, and a working dashboard does not mean success. A short PoC should establish measurement quality, explanatory operating context, replay of known events, closed-loop action, and a defensible estimate for scaling.

Phase 0: pre-start definition

  • Approve assets, failure modes, exclusions, and criticality.
  • Confirm measurement-point photographs, mounting drawings, and safe power, terminal, and network access.
  • Name owners of PLC, SCADA, MES, and CMMS context.
  • Build the initial truth ledger from anonymized maintenance, failure, and quality records.
  • Define recipient, response due date, escalation, and stop authority.

Phase 1: installation and data-quality acceptance

Do not proceed because a value appears on screen. Test range, noise floor, saturation, missing data, direction, sensor ID, clock offset, restart, communications loss and recovery, buffering, and battery/power state. Introduce a known operating change and confirm that the sensor and PLC mode preserve the same sequence.

Phase 2: mode-aware baseline and rule hypothesis

Collect planned steady, startup, load, and product states and review their distributions. Maintenance and data personnel should jointly explain normal variability, replay known anomalies, and choose corroborating signals. Compare simple differences, trends, and persistence logic before assuming that complexity is superior.

Phase 3: shadow operation

Keep candidate alerts away from automatic shutdown. Review them with named staff and update the truth ledger. Determine whether mode, mounting, data quality, or notification wording caused an incorrect decision. Test Thai, English, and Japanese views as needed, field evidence capture, and work-order handoff.

Phase 4: acceptance and scaling decision

Decide pass, conditional pass, or retest. Assess installation standard, tag dictionary, data quality, cybersecurity, workload, maintenance feedback, change control, export, removal, and restoration—not detection alone. Separate reusable design from asset-specific work before estimating scale-up.

Anomaly Detection Sensor Selection for Thai Factories: PoC to FAT/SAT - figure 3

PoC acceptance metrics: agree values from site evidence

Acceptance metricMethod and evidenceHow the target is setOwner
Operating-mode coverageActual acquisition against an approved mode listAgree modes required by asset cycle and PoC objectiveProduction and asset owner
Data availabilityExpected versus received samples; missing-interval listSeparate planned stops and communications tests; agree by use caseOT/IT
Clock alignmentDifference from PLC events; correction and restart logAgree tolerance needed to resolve event sequenceOT/control
Measurement repeatabilityComparable-condition trend and remounting differenceBase on sensor, mounting method, and intended useMaintenance/vendor
Known-event detectionResult against approved fault/work eventsDefine event window and expected action by failure modeReliability/maintenance
False-alert burdenUnnecessary reviews, time spent, and cause classAgree an operable load by recipient and priorityMaintenance manager
Miss reviewReverse-check fault, quality, and downtime ledgerJudge by observable scope and consequence classAsset/quality
Response completionEvidence from notification to inspection, work, and closureAgree owner and due date by notification classMaintenance management
Recovery/idempotencyCommunications loss, resend, duplicate, and sequence testsAccept scenarios without data loss or duplicate workOT/IT
SecurityAuthentication, authorization, certificate, log, and remediation evidenceConform to site policy and risk assessmentInformation security

If precision or recall is used, freeze the denominator and review window. One fault producing many alerts gives different alert-level and event-level results. “Normal” should be corroborated by inspection, work, and quality records rather than defined merely as no alert. Where fault examples are scarce, report exactly which events and failure modes were assessed instead of manufacturing a generalized accuracy claim.

7. The RFP must connect hardware, data, operations, and evidence

Comparing sensor quantity and cloud subscription alone hides mounting, PLC tags, networks, time sync, CMMS integration, tuning, training, and FAT/SAT. Lock scope, exclusions, responsibilities, deliverables, and acceptance evidence first.

RFP, FAT, and SAT checklist

TopicRFP requirementFAT evidenceSAT evidence
Target faultsAsset, failure mode, signal, observable scope, exclusionsRequirement traceability and simulated/historic resultsField asset, point, and operating-mode confirmation
Sensor specificationPrinciple, range, bandwidth, accuracy, environment, calibrationDatasheet, calibration evidence, configuration versionInstalled wiring, protection, tag, orientation photo
SamplingRaw waveform/features, rate, window, storage, event captureKnown-input bandwidth, clipping, missing-data testActual noise floor, waveform, communications load
Mounting standardPosition, direction, fastening, surface, remounting, cableInstallation procedure and change controlPoint-level photo, torque/orientation, field approval
ContextSpeed, load, mode, product, ambient conditionTag dictionary and time-stamped test dataSynchronization and semantic check against live PLC/MES
Data qualityMissing, flat, spike, clipping, drift, batteryFault injection and expected quality flagsCommunications loss, restart, recovery, field notification
Analytics/versionFeature, rule/model, threshold, mode, approvalReplay on frozen data and old/new comparisonProduction setting, access, change and rollback procedure
Notification designAdvisory/alarm/interlock class, priority, wording, ownerRouting, acknowledgement, escalation scenariosReal device, language, shift handover, CMMS work order
SecurityArchitecture, direction, authentication, encryption, certificates, logsPositive and negative authorization testsFirewall, DNS, certificate renewal, audit logs
Continuing operationCalibration, battery, replacement, re-baselining, supportRunbook, service terms, backup and restoreField-user demonstration and sign-off
Exit/migrationData, settings, model export, removal, account deletionExport format and completeness testSite restoration, access removal, handover confirmation

Questions that make proposals comparable

  • Do hardware prices include mounting, cabinet work, cabling, installation, calibration, and spares?
  • Who owns wireless battery replacement, site radio survey, and repeaters?
  • Can the factory export raw waveforms? Are feature definitions and versions disclosed?
  • Can buffering meet the required shutdown and recovery scenarios, not merely a quoted number of hours?
  • What PLC/SCADA/MES/CMMS tags, protocols, test environments, and change assumptions are included?
  • How much initial tuning, re-baselining after remounting, version update, and false-alert review is included?
  • What Thai field training, English engineering support, and Japanese head-office reporting is in scope?
  • Who pays for correction and retest after FAT/SAT failure, and how are conditional acceptance and open items handled?

Cost, savings, payback, and ROI depend on machinery, installation, existing infrastructure, failure history, and operating organization. This article supplies no universal price or guaranteed return. Require every bidder to quote the same assets, failure modes, retention, integrations, and acceptance scenarios, then compare differences in assumptions.

8. FAT accepts reproducibility; SAT accepts field viability

FAT should replay the path from sensor input through notification and evidence in the vendor-controlled environment. Include missing data, flat line, saturation, reversed time order, duplicates, communications loss, unknown modes, and threshold-version change. Repeating the same versioned input should give the same result; any difference must be explainable.

SAT verifies actual machinery, mounting, cabinets, networks, clocks, PLC tags, devices, recipients, and shifts. A model that passes simulated data may change under mounting resonance, electrical noise, radio shadow, or product mix. Reconcile the measurement-point photo and sensor ID with the tag dictionary, asset register, and dashboard. Have field users demonstrate recovery from communications loss, receipt, work-order creation, and closure from the runbook.

Acceptance must record each open item’s owner, due date, compensating control, and retest condition. Do not close on “no critical defects” alone. Track data quality, operating-mode coverage, false-alert workload, miss review, response ownership, cybersecurity, and exportability. Preserve FAT/SAT inputs, configuration versions, expected results, and signatures as the baseline for future assets.

9. Thai-factory implementation considerations

Heat, humidity, dust, washing, oil, cabinet temperature, long cable runs, and radio shielding affect sensing and communications. Verify actual chemicals, wash direction, cable glands, grounding, protective enclosures, and maintenance removal rather than relying on an IP rating alone. In a multilingual plant, avoid an alert that shows only an “anomaly score.” Show asset, point, current value, baseline, duration, operating mode, recommended check, and a concise warning against unsafe action.

According to the cited Thailand BOI release, from the Smart & Sustainable upgrade measure’s start in 2023 through H1 2026 it received 1,397 applications/projects representing more than THB 146 billion. This is context on programme uptake. It does not mean this sensor PoC is eligible, will be approved, or will generate a similar result. Confirm current eligibility and filing conditions with BOI materials and qualified advisers.

For examples of how to bound an asset and maintenance use case, see predictive-maintenance case studies for Thai factories. Do not copy a case-study model or number unless measurement point, mode, fault definition, and evaluation window match your plant.

10. Common failure patterns and controls

“Install sensors everywhere first”

Scaling channel count before defining faults and action overwhelms data-quality review and response. Complete one closed loop—from measurement to decision, work, and feedback—for a critical failure family first.

“Learn normal and anomalies will be obvious”

Mixed operating modes create nuisance alerts. Departure from normal is also not automatically a fault. Use mode-aware baselines and field confirmation; treat unfamiliar change as “review required.”

Accepting “99% accuracy” at face value

It is not comparable without population balance, alert unit, review window, labels, and target faults. Review event-level detection, lead time, unnecessary investigation effort, and miss consequence alongside a confusion matrix.

“Data reached the cloud, so quality is good”

Poor mounting, clock error, clipping, hidden interpolation, or an undocumented sensor swap can all travel successfully. Store quality flags and configuration version with each value and maintain traceability to raw evidence.

“The notification completed maintenance”

Without recipient, due date, inspection method, work order, parts, closure, and result feedback, the system does not improve. Connect the PoC to CMMS or a controlled ledger and reject ownerless notifications.

FAQ: anomaly-detection sensor selection and acceptance

Should a factory choose vibration or temperature sensors?

Choose from the failure mode and observable physical change. Vibration is a candidate for early mechanical change in rotating assets; temperature is suitable for heating, cooling performance, and thermal processes. Speed, load, pressure, and flow may be needed for confirmation. Do not select either modality before the failure-mode table.

Is one fixed threshold enough for temperature monitoring?

A fixed safety limit may be required, but condition monitoring often changes with ambient, load, product, and warm-up. Combine absolute values, differential temperature, rate of rise, peer comparison, or mode-adjusted residuals, and accept sensor response and installation as part of the system.

How should vibration sampling frequency be selected?

Base it on the frequencies associated with the target fault, speed, sensor bandwidth, mounting, and analysis method. There is no universal value. State waveform versus feature, bandwidth, record length, window, event waveform, clipping, and time sync in the RFP and verify with known or historic signals at FAT.

Can an equipment-failure PoC finish in 30 days?

Thirty days may fit the plan, but sufficiency depends on the asset cycle and mode coverage. Extend or change scope if required load, products, starts/stops, or maintenance events are absent. Accept measurement quality, event replay, notification operation, and recovery rather than waiting only for failures.

What false-alert percentage is acceptable?

There is no universal percentage. Agree it from asset criticality, recipient workload, review effort, miss consequence, and the evaluation window. Separate alert-level from event-level results and use the truth ledger to decide whether work is operable and serious misses remain.

Can AI anomaly detection drive a safety interlock?

Do not connect a PoC score directly to a safety trip. Advisory, operator alarm, and safety interlock have different purposes and validation. Modifying protection requires a separate safety design under the site’s risk assessment, independence, verification, and change-management rules.

What evidence should an RFP require?

Require failure-mode traceability, point and mounting drawings, calibration, tag dictionary, clock synchronization, data-quality tests, model/threshold versions, FAT inputs and expected results, SAT photos and logs, response workflow, security tests, runbook, and settings/data export. A dashboard demonstration alone is not acceptance.

Conclusion: design measurement, decision, action, and learning as one loop

Anomaly-detection sensor selection is not a component comparison between vibration and temperature. Start with failure modes and observable signals, then control mounting, bandwidth, sampling, time, operating context, and data quality as a measurement system. Evaluate false positives and false negatives together, and preserve the boundary among advisory, operator alarm, and safety interlock.

A 30–90 day PoC can verify measurement quality, known-event replay, mode-aware baselines, the path from notification to work, recovery, and change control without promising that every fault will occur. Specify deliverables and ownership in the RFP, accept reproducibility at FAT, and accept field viability at SAT. That is the practical route from an equipment-failure prediction idea to sustained operation.

TOMAS TECH helps Thai factories define target assets, sensor and gateway architecture, PLC/MES/CMMS integration, a 30–90 day PoC, and evidence-based RFP, FAT, and SAT criteria. You can contact us while still defining failure modes or planning a small validation on one or a few machines.

Primary references