An equipment anomaly detection PoC can run for 90 days and still leave a Thai factory with nothing more than a dashboard and a queue of unexplained alarms. To prevent that outcome, acceptance criteria must be agreed before technology tuning begins. This article does not repeat a general introduction to anomaly-detection AI or sensor selection. It focuses on one practical question: how should a factory design a 90-day PoC for rotating equipment so that the result can be classified as accepted, conditionally accepted, or rejected? The framework covers detection lead time, false alarms and missed detections, four-level alarms, communication outages and missing sensor data, connection to maintenance work orders, responsibility boundaries, FAT/SAT, and exit criteria.
Unless a published source is explicitly cited, every duration, count, ratio, and effort estimate in this article is an illustrative model assumption, not a benchmark or a customer result. Actual criteria must reflect each machine’s failure modes, operating patterns, safety requirements, maintenance organization, and network quality.
Why “it detected something” is not enough to accept an equipment anomaly detection PoC
A condition monitoring system can usually start collecting values and drawing charts relatively quickly. The factory is not buying charts, however. It needs a maintenance capability: when equipment condition changes, someone must know when to decide, what to inspect, and which maintenance job to initiate. The accepted object is therefore not the algorithm alone. It is the entire chain below.
- The physical condition of the asset changes.
- Sensors acquire a valid signal.
- The edge unit processes it while distinguishing missing data from a healthy condition.
- The equipment anomaly detection system assigns an actionable condition level.
- The right person is notified and validates the event.
- A work order is created in the CMMS or maintenance register.
- Inspection and repair results are written back to the anomaly event.
- False alarms, missed detections, and detection lead time are recalculated.
ISO 17359:2018 provides general procedures for setting up a machine condition monitoring programme, including equipment audit, criticality, failure modes, and monitoring methods. The ISO 13374 series addresses processing, communication, and presentation of condition monitoring information. These standards do not prescribe a universal pass mark for a vendor product. They do, however, support a systems view in which the factory specifies the complete path from measurement to decision instead of accepting only sensor accuracy.
On 2 September 2026, NTN announced an industrial condition monitoring system that, according to the company, collects and analyses vibration data and displays four stages: normal, initial damage, caution, and warning. NTN says the system supports bearings from other manufacturers and that its service can cover equipment investigation, proof testing, implementation, and operation. In the same company announcement, NTN reports that a six-month evaluation at a resource-recycling company detected signs including cracks on a belt-conveyor rotating shaft and led to formal adoption. This is a vendor product announcement, not a guarantee that every plant will achieve the same result. It nevertheless illustrates why multi-stage status and the transition from proof testing to operation belong in an acceptance plan.
Emerson’s product announcement of 8 September 2026 similarly describes continuous data transmission from a turbomachinery protection system to an analytics platform, including analysis of startup and shutdown cycles that can last hours or days. A PoC assessed only during steady production cannot expose the false alarms caused by startup, shutdown, changeover, cleaning, or speed changes. The 90-day test must therefore prove that its logic survives different operating states, not merely that a chart stays stable during normal running.
Freeze the PoC acceptance specification in the first five days
If pass conditions are changed after results become visible, the evaluation becomes biased. During Day 1–5, the factory, maintenance team, IT/OT team, and solution provider should approve an acceptance specification. Every later change should have a version, a reason, and an approver. At minimum, map the following items in one traceability table.
| Requirement ID | Acceptance object | Metric | Example of a complete pass statement | Evidence | Owner |
|---|---|---|---|---|---|
| ACC-01 | Detection | Detection lead time | Define minutes from reference-event onset to Watch by failure mode | Raw waveform, event timestamps, action log | Maintenance owner |
| ACC-02 | Quality | False-alarm rate | Freeze eligible alarm count, denominator, and exclusions | Alarm register, operating state | Data owner |
| ACC-03 | Quality | Miss rate | Freeze definition and review method for ground-truth events | Inspection and failure records | Reliability owner |
| ACC-04 | Operations | Four-level alarm | Define recipient, response time, and action for every level | Notification and acknowledgement logs | Maintenance team |
| ACC-05 | Availability | Outage and missing data | Define detection, storage, resend, recovery, and gap display | Gateway and edge logs | OT/IT owner |
| ACC-06 | Workflow | Work order | Create a CMMS record at the target level and carry asset ID and evidence | CMMS audit trail | Maintenance planner |
| ACC-07 | Installation | FAT/SAT | Test I/O, time, tags, access, and recovery | Signed test sheet | Factory + provider |
| ACC-08 | Exit | PoC decision | Define Accepted, Conditional, and Rejected rules | Decision minutes | Sponsor |
“Accuracy above 90%” is not an acceptance requirement on its own. It becomes testable only after the team states what counts as ground truth, who confirms the label, what enters the denominator, how missing periods are treated, and whether repeated notifications from the same physical event count once or many times.
Limit the scope to three to five assets and one to three failure modes per asset
The following is an illustrative model assumption. A Thai factory selects four assets: a motor, pump, exhaust fan, and conveyor drive. Each asset is linked to a small number of observable failure modes rather than a promise to detect “any anomaly.”
| Asset | Monitored failure modes (assumed) | Primary signals | Operating context | Example outside the PoC scope |
|---|---|---|---|---|
| Conveyor drive shaft | Bearing degradation, imbalance | Vibration, temperature | Speed, load, material present | Belt mistracking, foreign-object jam |
| Cooling-water pump | Bearing degradation, cavitation indication | Vibration, pressure, current | Flow, valve position | Detecting the pipe leak itself |
| Exhaust fan | Imbalance, looseness | Vibration, RPM | Damper position | Filter differential-pressure fault |
| Spindle motor | Bearing degradation, overload trend | Vibration, temperature, current | Product, speed | Internal drive fault |
For the measurement limits of vibration sensors, see our practical guide to vibration-sensor machine diagnostics. For sensor, collection, diagnostic, and cost layers, see the predictive maintenance system implementation guide. This article starts at the next question: how to accept or reject a candidate system.
Divide the 90 days into six gated phases
If the plan merely says “collect data for 90 days,” critical problems can remain hidden until the final week. Divide the PoC into six phases that correspond to acceptance requirements.
| Period | Phase | Main work | Gate |
|---|---|---|---|
| Day 1–5 | Definition | Freeze assets, failure modes, metrics, roles, and safety limits | Acceptance specification approved |
| Day 6–15 | FAT and installation preparation | Simulate tags, time, alarms, buffer, and CMMS interface | FAT passed |
| Day 16–30 | SAT and baseline | Install on site, identify operating states, establish normal ranges | SAT passed |
| Day 31–60 | Evaluation run | Test scenarios, notifications, data gaps, outages, and work orders | Interim review |
| Day 61–80 | Controlled tuning and retest | Adjust only pre-approved thresholds and rules | Frozen release |
| Day 81–90 | Acceptance decision | Blind label comparison, KPI calculation, open items, handover | Go/Conditional/No-Go |
Unlimited tuning after Day 61 is unacceptable. If thresholds are continually adjusted against the evaluation data, the PoC can overfit its own 90-day history. Specify what may be tuned, how many times, and who approves it. Freeze the configuration on Day 80 and use Day 81–90 as a practical holdout period.

Make detection lead time the central KPI
Detection lead time is not simply “days before failure.” It is impossible to measure when the reference event has no agreed timestamp. Define three times for every failure mode:
T_signal: the timestamp at which an expert review concludes that a persistent signal change began.T_alert: the timestamp of the first valid Watch-or-higher alarm from the system.T_action: the timestamp at which maintenance acknowledged the notification and accepted an inspection or work order.
Technical detection delay is T_alert − T_signal. Operational response delay is T_action − T_alert. If repair or shutdown occurs at T_event, actionable margin is T_event − T_action. An early alarm is not automatically useful; it may simply be an early false alarm. The factory should ask whether it creates enough time to arrange inspection, obtain parts, or include the task in a planned stop.
A natural failure is not guaranteed during a 90-day window. Separate the evidence into four classes.
- Confirmed physical event: degradation or failure verified by inspection or the removed part. This is the strongest evidence.
- Historical replay: past data replayed in chronological order without using future information.
- Safe test scenario: sensor removal, simulated input, communication loss, or speed-state change that cannot damage the asset.
- Analytical indication: a specialist confirms a signal change, but the physical component cannot be inspected.
Do not combine all four classes into one “successful detection” count. Report each separately. If the PoC observes no confirmed physical event, the final report must not claim that failure-prediction performance was proven. It can state that signal acquisition, state detection, and the operating workflow were validated.
Derive lead-time acceptance from the maintenance process
Consider an illustrative model where arranging inspection takes four hours, checking part availability takes one business day, and negotiating a planned stop takes two business days. A Critical alarm immediately before shutdown adds little value. Watch or Caution must provide enough time for secondary diagnosis; Critical remains a final risk decision governed by the plant’s safety procedure.
| Level | Purpose | Illustrative response | Acceptance evidence |
|---|---|---|---|
| Normal | Confirm the baseline remains valid | Scheduled review | Baseline and operating state |
| Watch | Observe an early change | Review trend within one business day | Reviewer and comment |
| Caution | Decide on planned inspection | Secondary diagnosis within four hours; create WO if required | Waveform, diagnosis, WO number |
| Critical | Make an immediate risk decision | Follow plant safety and shutdown procedure | Notification, acknowledgement, action log |
These level names are an article design example. They resemble the four stages in NTN’s announcement but do not describe NTN’s proprietary decision logic or performance. If the plant already uses different andon or alarm names, prioritize a clear owner, deadline, and action over introducing new colours.
Manage false alarms and missed detections in the same scorecard
Reducing false alarms can increase missed detections, and reducing misses can increase false alarms. A criterion for only one side can be gamed by extreme tuning. Report at least the following together:
- True positive (TP): a valid alarm occurred within the permitted time window for a ground-truth event.
- False positive (FP): an actionable alarm occurred without a corresponding ground-truth event.
- False negative (FN): a ground-truth event occurred without a valid alarm in the permitted time window.
- True negative (TN): no ground-truth event and no actionable alarm occurred in the evaluation window.
For continuous time series, counting every second makes TN enormous and can produce an apparently perfect accuracy. Event-level or asset-day denominators are more useful for a factory PoC. Define a suppression rule in advance—for example, notifications from the same asset and same failure mode within 30 minutes count as one event.
| Metric | Example definition | Interpretation warning |
|---|---|---|
| Event recall | TP ÷ (TP + FN) | Very uncertain when confirmed events are few |
| Precision | TP ÷ (TP + FP) | Approximates how trustworthy an actionable notification feels |
| False-alarm burden | FP ÷ monitored asset-days | Easy to translate into labour |
| Missed events | Absolute FN count and severity | Do not dilute a critical miss into a percentage |
| Response compliance | Responses inside deadline ÷ actionable events | Measures the operating process as well as the system |
The solution provider must not define ground truth alone
Ground-truth review should involve plant maintenance, equipment engineering, operations, and where useful the provider’s diagnostic specialist. Hide the system score during label review and use waveforms, temperatures, loads, inspection results, removed parts, and production records. Events on which the panel cannot agree should be marked “uncertain,” removed from the primary KPI, and included in sensitivity analysis.
Exclusions must also be frozen in advance: post-installation settling, deliberate hammering during maintenance, construction while stopped, or operation outside the contracted speed range. Removing an inconvenient alarm after the fact as “special operation” biases the result. Routine startup, shutdown, and product change are not special. Emerson’s 2026 announcement highlights long-duration startup and shutdown data; the broader design lesson is to include transients in the monitoring plan.
Turn a four-level alarm into action, not colour
A green-yellow-red screen has no value if it does not change maintenance work. Every level needs criteria, recipient, deadline, authority, and clearance rules.
| Item | Normal | Watch | Caution | Critical |
|---|---|---|---|---|
| Condition | Inside valid baseline | Persistent early change | Multiple indicators or diagnosis require inspection | Severe condition or rapid change |
| Notification | None or daily report | Dashboard + assigned maintainer | Maintenance team + supervisor | Plant emergency route |
| Deadline | Scheduled | One business day (example) | Four hours (example) | Immediate procedure |
| Action | Continue monitoring | Check trend and operating state | Secondary measurement, create CMMS WO | Safety check; authorized person decides shutdown |
| Clearance | Continue | Stable normal period or approval | Inspection result and approval | Return approval after corrective action |
An anomaly detection system must not silently become an automatic shutdown system. Protection and condition monitoring have different roles. Shutdown decisions must follow the existing safety design and authority matrix; a condition-monitoring alarm is not a substitute for a protection relay or safety instrumented function.
Alarm escalation and de-escalation should include hysteresis, persistence, and suppression conditions. A single threshold crossing should not necessarily produce Caution. Escalation may require persistence in a defined operating state. Clearance should use a lower return boundary and a confirmation interval to avoid chatter. The chosen settings are plant-specific assumptions and should be injected during FAT and confirmed against real signals during SAT.

Never display a communication outage or missing sensor data as Normal
No data is not evidence of healthy equipment. Communication loss, sensor failure, battery depletion, gateway failure, tag errors, and time-sync errors must have a data-quality status separate from equipment condition.
| Data quality | Display | Alarm evaluation | Required action |
|---|---|---|---|
| Good | Timestamp and refresh interval valid | Normal evaluation | Continue |
| Delayed | Delay exceeds the defined interval | Do not treat the last value as current Normal | Delay notification; inspect buffer |
| Missing | Missing ratio exceeds tolerance | Set equipment state to Unknown | Check sensor, power, and path |
| Invalid | Out of range, stuck value, reversed time | Exclude from decision | Check calibration, tag, and clock |
| Recovering | Resending and reconciling after reconnect | Distinguish old from live data | De-duplicate, reorder, confirm completion |
Accept the edge buffer and resend behaviour
Intentionally interrupt the network during the PoC and verify that:
- The gateway detects loss within the agreed time.
- The edge unit stores data locally.
- The dashboard shows Delayed or Unknown rather than a frozen Normal.
- Data is resent with original timestamps after recovery.
- Duplicate events are not created and any unrecoverable gap is visible.
- If local decisions are required during loss, the alternative notification path works.
In an illustrative design, one-minute features might be buffered for 72 hours. FAT would confirm capacity and full-buffer behaviour; SAT might include a controlled two-hour outage and resend. Seventy-two hours is not a recommended universal value. It must be derived from recovery objectives, communication quality, and data volume.
SKF’s 8 September 2026 announcement describes its Insight bearing concept, which measures conditions such as load, speed, temperature, and vibration inside the bearing, along with an integrated hardware platform for sensing, processing, and communications. This remains a vendor statement, not a general performance guarantee. It does highlight the need to identify in acceptance tests whether a gap originated in sensing, processing, or communications. Telit Cinterion’s 10 September 2026 announcement also previews edge fault detection and recovery on a live robotic line. Even when recovery is automated, the factory needs an audit trail of what recovered, outage duration, lost data, and actions performed.
Connect condition monitoring to the maintenance work order
When equipment-diagnostics IoT ends at a dashboard, the shop floor must check another screen and retype the event. The PoC should include the path from Caution-or-higher events into a CMMS work order or inspection request. If automatic creation is too risky, an approval button that creates a pre-filled draft is acceptable.
Transfer at least:
- unique plant, line, and asset ID;
- sensor ID and measurement point;
- alarm level, start time, and duration;
- candidate failure mode and confidence where available;
- links to waveform, trend, and operating conditions;
- recommended next check, not an unreviewed shutdown command;
- source event ID; and
- assignee, due date, and priority.
At closure, send the result back to the anomaly event. Use controlled codes such as No Fault Found, Retightened, Lubricated, Bearing Replaced, Sensor Fault, or Operating-Condition Effect—not free text alone. Without this closed loop, the system cannot learn from false alarms, and management cannot explain avoided downtime or maintenance value.
Work-order acceptance scenarios
| Scenario | Input | Expected result |
|---|---|---|
| First Caution | Valid asset event | One WO draft with asset ID and evidence link |
| Continuing same event | Same failure mode within 30 minutes (assumed) | Update existing event; do not flood new WOs |
| Escalation to Critical | Escalation from Caution | Raise existing WO priority and notify supervisor |
| Confirmed false alarm | Inspection finds healthy asset and a cause | Return result code to anomaly event |
| Resend after recovery | Historical-timestamp event arrives | Keep occurrence and receipt time separate; no duplicate WO |
| Asset master mismatch | Unknown asset ID | Quarantine instead of auto-creating; send configuration alert |
Define responsibility down to the first troubleshooting action
When sensor values disappear, maintenance, IT, the integrator, and the device vendor can all wait for someone else. A responsibility matrix must include detection, first check, evidence, recovery objective, and escalation—not merely an owner name.
| Layer | Example main responsibility | First check | Required evidence |
|---|---|---|---|
| Sensor and mounting | Maintenance + device provider | Power, fixation, direction, calibration, damage | Installation photo, model, calibration, point ID |
| Edge | Integrator / system provider | Process, capacity, clock, buffer | Edge log, configuration version, restart history |
| OT network | Plant OT/IT | Switch, VLAN, firewall, wireless quality | Connection log, change record |
| Analytics and alarm | Provider + reliability engineer | Model version, threshold, suppression, input quality | Inference log, setting diff |
| Notification and CMMS | IT + maintenance planning | API, identity, asset master, queue | API response, event ID, WO number |
| Maintenance action | Plant maintenance | Receipt, inspection, safety procedure, closure | Acknowledgement time, findings, work result |
Also distinguish cybersecurity anomaly detection from physical asset anomaly detection. NIST IR 8219 evaluates behavioural anomaly-detection techniques for manufacturing industrial control systems, focusing on abnormal network behaviour. That is a different purpose from vibration- and temperature-based equipment diagnosis. Yet adding monitoring devices to an OT network still makes asset inventory, traffic paths, logs, access rights, and change control part of acceptance. Do not combine “machine anomaly” and “communication anomaly” behind one red lamp; route them to different owners and procedures.
Use FAT to remove defects before site installation
Factory Acceptance Testing takes place in a test environment before site installation. Data replay and signal simulation can verify many requirements without the physical machine.
FAT checklist
- Tags and units: asset ID, point, direction, unit, and sampling condition match the data dictionary.
- Time: edge, gateway, server, and CMMS time zones and synchronization are consistent; for example, store UTC and display ICT under a documented rule.
- Four-level transition: simulate Normal→Watch→Caution→Critical and recovery.
- Hysteresis: the alarm does not chatter near a boundary.
- Missing data: inject null, stuck, out-of-range, and reversed-time values and confirm Unknown or Invalid.
- Outage: test disconnect, buffering, full buffer, recovery, resend, and de-duplication.
- Notification: Thai, English, and Japanese names display correctly; routing and suppression are correct.
- CMMS: test create, update, failure, retry, duplicate prevention, and asset-master mismatch.
- Access: separate view, acknowledge, threshold-change, and administrator permissions.
- Audit: trace who changed a setting, acknowledged an alarm, and cleared it.
- Version: record model, rule, firmware, configuration, and data-dictionary versions.
- Export: retrieve raw data, features, events, and work results in the agreed format.
Passing FAT does not prove field performance. It means the system is testable according to the specification and known interface defects have not been carried to site. Every open item needs severity, workaround, due date, and owner. Installation should not proceed with unresolved issues involving safety, data integrity, or duplicate work orders.
Use SAT to accept field conditions and the operating workflow
Site Acceptance Testing occurs after installation on the real equipment and network. It covers mounting direction, cabling, power, wireless conditions, asset IDs, speed, load, ambient vibration, cleaning, and temperature—conditions a bench cannot reproduce.
SAT checklist
- Match sensor position, direction, fastening, and identification label to the drawing.
- Measure the noise floor while stopped and normal baselines at representative speed and load.
- Record startup, shutdown, changeover, cleaning, and idling as separate operating states.
- Verify the display from a field terminal, maintenance office, and approved remote route.
- Conduct a planned communication interruption and verify display, local storage, recovery, and resend.
- Push a simulated Watch/Caution event through notification, CMMS creation, and closure feedback.
- Test night-shift and holiday escalation routes.
- Ask actual users, including Thai-speaking staff, to explain the alarm reason and next action.

FAT and SAT evidence should not be a folder of screenshots. Link requirement ID, test procedure, expected result, actual result, timestamp, data file, executor, approver, and defect ID. ISO 13374-3 addresses communication of condition monitoring information between systems, while ISO 13374-4 addresses presentation of diagnostic, prognostic, advisory, and recommendation information. Data reaching a screen is not enough; its meaning and decision context must survive the journey.
Illustrative predictive-maintenance PoC estimate: show decision value, not invented savings
The following is not a customer case. It is an illustrative model assumption for monitoring four critical rotating assets for 90 days. Costs vary substantially with sensor count, interfaces, travel, and CMMS design, so the table shows effort and deliverables rather than currency.
| Work package | Assumed effort | Main deliverable |
|---|---|---|
| Asset audit and failure-mode definition | 4 person-days | Scope, points, exclusions |
| Acceptance specification and responsibility | 3 person-days | KPIs, RACI, exit rules |
| FAT | 4 person-days | Signed tests, defect register |
| Installation and SAT | 6 person-days | Installation record, baseline, SAT report |
| Monitoring and weekly review | 12 person-days | Event register, setting-change log |
| CMMS interface and operating test | 5 person-days | API map, work-order evidence |
| Final evaluation and handover | 4 person-days | Decision, scale plan, open items |
| Total | 38 person-days (assumed) | Complete 90-day PoC package |
Avoid claiming “one shutdown was prevented” without evidence. Use scenarios:
Expected annual avoided loss = annual frequency of the target failure × impact per shutdown × avoidable share after detection
Enter a range for every factor and compare low, middle, and high cases. If no failure occurs during the PoC, the factory can still test whether detection lead time and work-order routing meet the time required by the scenario. The avoidable share remains an assumption, however, and must not be presented as a measured fact.
False-alarm workload is easier to observe:
Monthly false-alarm effort = monitored assets × FP per asset-day × operating days × review time per event
In an illustrative case of four assets, 0.1 FP per asset-day, 26 operating days, and 15 minutes per review, the result is 2.6 hours per month: 4 × 0.1 × 26 × 0.25 = 2.6. At 1.0 FP per asset-day under the same assumptions, it becomes 26 hours. Translating performance into labour makes operability clearer than an extra decimal place of “accuracy.”
Define the PoC exit before it starts
Automatically extending the trial on Day 90 because “we need more data” is not an exit. Define three outcomes at the start.
Go
- No critical FAT/SAT nonconformity remains; every minor item has an owner and deadline.
- Each target failure mode has a result at the agreed evidence level.
- False-alarm burden and misses fall inside the pre-agreed boundary.
- Missing data and outages create Unknown, and recovery/resend works.
- An alarm can be traced through the work order and maintenance result.
- Plant staff can own daily operation.
Conditional Go
Use this only when remaining issues do not compromise safety or data integrity and their deadline, cost ownership, retest, and responsible party are agreed. Examples may include inconsistent naming for a non-critical asset or a pending Thai-language report correction. Do not conditionally accept missed critical events, hidden data loss, duplicate work orders, or weak access control as “operational workarounds.”
No-Go or redesign
- The monitored signal does not fit the target failure mode.
- Routine transients cannot be separated from faults and false-alarm burden is unacceptable.
- A critical event was missed and corrective retesting is incomplete.
- An outage is displayed as Normal or gaps cannot be traced.
- CMMS integration cannot prevent duplicate or wrong-asset work orders.
- Ownership of configuration and alarm decisions cannot be agreed.
No-Go is not necessarily failure. If the PoC shows that the factory should choose another asset, collect an additional signal, use route-based measurements, or fix CMMS master data first, it has prevented a larger and more expensive mistake.
Contract the deliverables, not just sensor rental and dashboard access
If a predictive-maintenance PoC is purchased only as sensors plus temporary dashboard access, the factory may finish without data or settings. Confirm at least the following deliverables and usage rights before contract signature:
- asset, measurement-point, and failure-mode register;
- sensor installation drawings, model, settings, and calibration information;
- data dictionary, time rules, units, and missing-value codes;
- raw data or export at an agreed resolution;
- features, alarms, configuration changes, and acknowledgements;
- model/rule/threshold version and change rationale;
- FAT/SAT procedures and signed results;
- API and CMMS interface specification with asset-ID mapping;
- incident, backup, recovery, and support boundary;
- removal, retention, and account-deletion procedure after the PoC;
- production licence, connectivity, maintenance, and per-asset cost structure; and
- training material and plant operating procedure.
ISO 13374-1 gives general guidance for software specifications concerning processing, communication, and presentation of machine condition information. Part 3 addresses information exchange, and Part 4 presentation for technical analysis and decision support. Rather than casually claiming standards compliance, turn those concerns into procurement questions about portability, preserved meaning, display, and information for each responsible role.
FAQ: a 90-day equipment anomaly detection PoC
Can an equipment anomaly detection system prove accuracy in 90 days?
For low-failure-frequency assets, not enough real failures may occur to prove predictive accuracy statistically. The 90-day scope can validate signal quality, operating-state separation, historical replay and safe scenarios, false-alarm workload, missing-data handling, notification, CMMS integration, and role ownership. If real failures are scarce, state the evidence level and avoid calling the result a proof of predictive accuracy.
What false-alarm rate is acceptable for a condition monitoring system?
There is no universal percentage. Severity, asset count, review effort, and operating patterns change the acceptable level. Express it as actionable false alarms per asset-day and monthly labour, then check missed-event count and severity at the same time. Freeze the denominator, event-grouping rule, and exclusions before testing.
What should equipment-diagnostics IoT do when communications fail?
It should not leave the last value displayed as current Normal. Show a separate state such as Delayed, Missing, or Invalid and set equipment condition to Unknown. FAT/SAT should verify edge buffering, resend with original time, de-duplication, gap visibility, and any alternative notification route. Retention time should be engineered from the plant recovery objective and data volume.
Are both FAT and SAT required for a predictive-maintenance PoC?
They answer different questions. FAT checks tags, clocks, alarm transitions, missing data, outages, CMMS behaviour, and permissions before installation. SAT checks mounting, real load and noise, network, real users, and shift coverage. Separate test sheets make it easier to distinguish software defects from site conditions, even in a small PoC.
Are four alarm levels too many?
They are too many if each level has no different action. In the example, Normal, Watch, Caution, and Critical separate observation, secondary diagnosis, planned inspection, and immediate risk decision. A plant with an established three-level rule need not add a fourth colour. The requirement is unambiguous condition, data quality, deadline, action, and clearance.
Is the PoC rejected if no natural failure occurs?
Not automatically. Historical replay, safe input simulation, communication loss, sensor gaps, work-order integration, and operating response can still be tested. The team cannot claim a successful failure prediction, however. Scale-up should be gated, with a new review after enough confirmed physical events have accumulated.
Conclusion: a 90-day PoC is an acceptance project, not a product demo
For equipment anomaly detection in a Thai factory, a sensor value and an AI score are not the finish line. Freeze the acceptance specification in Day 1–5 and connect detection lead time, false alarms and misses, four-level alarms, data quality, communication recovery, CMMS work orders, and responsibility boundaries in one test system. Use FAT to remove interface and failure-path defects, SAT to verify site conditions and human actions, and Day 80 to freeze the configuration. On Day 90, decide Go, Conditional Go, or No-Go from traceable evidence.
Even when natural failures are rare, a factory can state exactly what the 90 days did and did not prove. A sound PoC replaces “the accuracy looked good” with a specific answer: for this asset, failure mode, and operating state, how early did the system detect a valid change, who acted within what time, and which record proves it?
TOMAS TECH can help at the candidate-comparison stage, including a 90-day acceptance matrix, FAT/SAT test sheets, and the scope of CMMS integration for a Thai plant. We first separate what the PoC can demonstrate from what must remain unproven. Contact TOMAS TECH to discuss your current equipment and maintenance workflow.
References and sources
- NTN Corporation, “Launch of Condition Monitoring System for Industrial Equipment” (Japanese, 2 Sep 2026)
https://www.ntn.co.jp/japan/news/new_products/news202600062.html
- Emerson, “Emerson Updates Turbomachinery Protection and Asset Health Monitoring for Safer Operations” (8 Sep 2026)
- SKF, “SKF advances Insight bearing technology through collaboration with Sentea” (8 Sep 2026)
- Telit Cinterion, “deviceWISE to Demonstrate Agentic AI and Automated Fault Detection & Recovery on Live Robotic Lines at IMTS 2026” (10 Sep 2026)
https://www.telit.com/press/devicewise-to-demonstrate-agentic-ai-robotic-at-imts-2026/
- NIST IR 8219, *Securing Manufacturing Industrial Control Systems: Behavioral Anomaly Detection* (2020)
https://nvlpubs.nist.gov/nistpubs/ir/2020/NIST.IR.8219.pdf
- ISO 17359:2018, *Condition monitoring and diagnostics of machines — General guidelines*
https://www.iso.org/standard/71194.html
- ISO 13374-1:2003, *Condition monitoring and diagnostics of machines — Data processing, communication and presentation — Part 1: General guidelines*
https://www.iso.org/standard/21832.html
- ISO 13374-3:2012, *Condition monitoring and diagnostics of machines — Data processing, communication and presentation — Part 3: Communication*
https://www.iso.org/standard/37611.html
- ISO 13374-4:2015, *Condition monitoring and diagnostics of machine systems — Data processing, communication and presentation — Part 4: Presentation*