Condition-Based Maintenance (CBM) 2026: A 90-Day Acceptance Blueprint
Condition-based maintenance (CBM) does not reduce downtime merely because sensors and a dashboard are installed. If alerts do not become controlled maintenance work, the project only creates more information. If thresholds are too sensitive, alarm fatigue follows. This guide shows how a Thailand factory can connect asset criticality, failure modes, baselines, dynamic thresholds, work orders, OT security, FAT/SAT, and a 90-day proof of concept into one auditable decision process.
CBM is a closed operating loop, not a measurement project
ISO 17359:2018 provides general procedures for establishing condition-monitoring programmes. ISO reports that this edition was reviewed and confirmed in 2023 and remains current. It is not a sensor-specific threshold catalogue. The useful procurement unit is therefore not “12 sensors”; it is a programme that detects an applicable failure mode, supports a decision, triggers safe work, and proves that the asset returned to an acceptable condition.

| Loop stage | Required input | Evidence produced | Accountable role |
|---|---|---|---|
| Select | Asset register, event history, criticality | Included and excluded assets | Production, maintenance |
| Design | FMEA, repair records, drawings | Failure-mode-to-signal map | Reliability, engineering |
| Measure | Location, unit, interval, operating state | Quality-checked time series | OT, supplier |
| Diagnose | Baseline, thresholds, suppression rules | Alert with rationale | Diagnostician |
| Act | Priority, due date, safety conditions | Work order and owner | Planner |
| Close | Repair, repeat measurement, cause | Recovery and learning record | Maintenance manager |
The PoC acceptance target is evidence that this loop works. A polished screen without a closed job is not maintenance DX.
Separate asset criticality from failure-mode detectability
Criticality describes consequences; detectability describes whether a physical precursor appears in the selected signal. A highly critical machine dominated by PLC faults or material jams may gain little from vibration alone.
| Class | Consequence | Detectability | PoC treatment |
|---|---|---|---|
| A | High | High | First priority; connect spares and work history |
| B | High | Low or unknown | Investigate failure physics first |
| C | Medium | High | Useful learning and tuning candidate |
| D | Low | Low | Exclude; retain routine inspection |
Score safety, quality, environmental impact, delivery, redundancy, repair time, and local spare availability—not only lost production. Then map each failure mode to a precursor and an action.
| Failure mode | Expected precursor | Context needed | Action | Limitation |
|---|---|---|---|---|
| Imbalance | Increase at rotational frequency | Speed, load | Clean and balance | Load can change amplitude |
| Misalignment | Axial change and harmonics | Coupling, temperature | Alignment check | Resonance may look similar |
| Looseness | Harmonics, waveform distortion | Base and bolt history | Inspect fastening | Poor sensor mounting can mimic it |
| Bearing damage | High-frequency impacts, envelope | Speed, bearing type | Lubricate or plan replacement | Bandwidth and mounting matter |
| Overheating | Temperature trend | Ambient and load | Check cooling/lubrication | Temperature alone does not identify cause |
| Electrical/control fault | Current and event log | PLC/VFD state | Electrical diagnosis | Often invisible to vibration |
ISO 13379-1:2025 addresses common concepts, technical characteristics, and guidance for selecting diagnostic approaches. It does not mandate a particular AI model. Ask suppliers which failures, inputs, operating assumptions, explanations, and limitations apply.
Build baselines around operating states, not a magic number of days
This illustrative plan allocates 28 days for an initial baseline; that is a planning assumption, not an ISO minimum. A month of one operating condition can still be inadequate. Separate startup, steady load, high and low load, product change, cleaning, and idle periods. For variable-speed assets, compare within speed bands. For pumps, add flow or valve position; for compressors, distinguish load and unload.
| Baseline control | Acceptance condition | Failure example | Correction |
|---|---|---|---|
| Missing data | Within the agreed tolerance | Gateway outage silently interpolated | Repair communications and time sync |
| Operating state | Main states classified | Startup mixed with steady state | Split state models |
| Mounting | Direction, point, and fixing recorded | Sensor placed on a cover | Relocate and recollect |
| Unit and scaling | Source and display agree | g confused with mm/s | Verify conversion and calibration |
| Healthy reference | Maintenance confirms healthy period | Degrading asset learned as normal | Rebuild after repair |
| Time | PLC, gateway, and CMMS align | Time-zone mismatch | Standardize NTP and display rules |
ISO 20816-3:2022 covers vibration evaluation for the industrial machinery within its scope: above 15 kW and 120–30,000 r/min. Do not copy its evaluation values to small motors, low-speed machines, transient states, or other equipment outside that scope. Use the relevant standard, manufacturer guidance, and machine-specific baseline together. Our vibration sensor equipment diagnosis guide explains the measurement boundaries in more detail.
Make dynamic thresholds and false-positive control testable

A dynamic threshold should compare like with like: the same speed, load, product, or operating mode. “AI-adjusted automatically” is not an acceptance criterion. Specify the learning window, excluded periods, update frequency, hard limits, human approval, model version, and rollback method.
Use at least two response levels. A warning begins verification; an alarm changes maintenance planning or an operating decision. Combine persistence, consecutive samples, multiple signals, and operating state. An analytics alarm should not be connected directly to an emergency stop without a separate safety design.
| Metric | Definition | Illustrative 90-day target | Guardrail |
|---|---|---|---|
| True positive | Actionable condition correctly alerted | Few real faults may prevent a meaningful rate | Inject faults only where safe |
| False alert | Reviewed, no action required | ≤0.25 per asset per week | Planning example, not benchmark |
| Miss | Later evidence shows an unalerted precursor | Aim for zero critical misses | Report denominator and observation period |
| Acknowledgement time | Alert to owner review | Within 4 business hours | Adapt to shift coverage |
| Work conversion | Valid alert receives a tracked decision | 100% documented | Record why no work was required |
| Closure | Post-repair measurement completed | ≥90% by due date | Categorize overdue reasons |
Classify each alert as real degradation, operating-state change, sensor/communications issue, known work, or insufficient evidence. Threshold changes need an owner, reason, before-and-after value, affected assets, approval, and effective date.
Turn every valid alert into a work order and close it after repair
An alert should state asset and point, unit, operating state, difference from baseline, rate of change, suspected failure mode, recommended check, priority, due time, and graph. When confidence is low, recommend a safe confirmation—portable measurement or lubrication inspection—not immediate replacement.
| Status | Owner | Required record | Exit condition |
|---|---|---|---|
| New | Monitor | Time, asset, evidence | Duplicates and planned work checked |
| Reviewing | Diagnostician | Waveform, state, field observation | Action decision made |
| Planned | Planner | Work, parts, outage, safety | Approved and ready |
| Executed | Technician | Repair details, photo, readings | Repeat measurement possible |
| Verified | Diagnostician | Before/after values, residual issue | Recovered or repeat work |
| Closed | Manager | Cause, lesson, register update | Evidence complete |
Carry a common alert ID into the CMMS or ERP. API integration is optional for a PoC; controlled CSV or a linked field can prove the workflow. Repair is not closure. Repeat the measurement under a comparable operating state and confirm recovery. This produces valuable labeled data.
Time-based preventive maintenance triggers work by interval. CBM uses observed condition to decide when work is needed. Predictive analysis may additionally estimate future condition or remaining life. Operational consistency matters more than the label. See our predictive versus preventive maintenance comparison.
Design OT security around segmentation, least privilege, and recovery
NIST SP 800-82 Rev. 3 addresses OT security while recognizing performance, reliability, and safety constraints. Segment sensor/gateway networks, monitoring services, and IT or cloud connections; permit only required flows. Use named, least-privilege, time-limited supplier accounts. Remote support should require a request, approval, time window, logging, and deactivation.
| Control | FAT check | SAT check | Operating evidence |
|---|---|---|---|
| Asset inventory | Model, version, owner | Matches installation | Change history |
| Network flow | Ports and direction | Allowed and denied tests | Firewall review |
| Accounts | Privilege and expiry | Login and disable | Periodic recertification |
| Time | NTP design | PLC/GW/cloud alignment | Drift monitoring |
| Backup | Scope, schedule, protection | Successful capture | Test record |
| Restore | Procedure and owner | Representative restore | Recovery exercise |
NIST SP 1339, published in June 2026, calls for backups to be integrated with change management, created and tested regularly, and reviewed during recovery exercises. For CBM, protect gateway configuration, sensor mapping, thresholds, model versions, dashboards, interfaces, and work history. A file existing is not proof of recovery: restore a representative configuration and verify tags, units, thresholds, and time.
Use FAT and SAT to test function, data, workflow, and recovery
FAT tests specifications with simulated inputs before site work. SAT tests the installed sensors, real network, actual operating states, and real users. “The dashboard opens” is insufficient.
| Test area | FAT example | SAT example | Evidence |
|---|---|---|---|
| Function | Input, unit, warning/alarm | Real signal and notification | Cases, screens, logs |
| Data quality | Missing/range/time faults | Disconnect and recovery | Missing-rate report |
| Diagnostics | Simulated waveform and state | Healthy values, known event | Expected/actual comparison |
| Workflow | Alert, approval, job ID | One complete field cycle | Work order and closure |
| Security | Access denied, audit trail | Firewall and remote session | Access records |
| Recovery | Create backup | Restore a representative setup | Time and reconciliation |
| Documentation | Procedures and register | User practical exercise | Versions and training record |
Write expected results before testing. For example: in state A, when signal X meets condition Y for Z samples, generate a warning to role P, suppress duplicates for N minutes, and retain an audit record. Keep defect severity, temporary control, owner, deadline, and retest condition. Unresolved major data loss, excess privilege, failed restore, or an unclosable workflow should hold acceptance.
Run the 90-day PoC through three decision gates

Ninety days is an illustrative planning period, not a market standard or success promise. Real faults may be too infrequent for a statistically meaningful detection rate. Evaluate data quality, replay of known events, alert burden, work closure, recovery, and user competence as well.
| Period | Main work | Gate question | Deliverable |
|---|---|---|---|
| Days 1–15 | Register, criticality, history, FMEA, network survey | Can the target failures be measured? | 12 assets and failure-signal map |
| Days 16–30 | FAT, installation, SAT, security, and notification checks | Can it be installed safely and produce trustworthy data? | FAT/SAT records, drawings, flow matrix |
| Days 31–45 | State-aware baseline collection, part 1 | Are the main operating states identifiable? | Initial data-quality report, state tags, exclusions |
| Days 46–60 | Complete and approve the 28-day baseline | Is “normal” comparable and approved? | Approved 28-day baseline and initial warnings |
| Days 61–75 | False-alert review, jobs, changes | Can the team absorb the work? | Change log and closure records |
| Days 76–90 | Recovery exercise, skills test, final review | Scale, modify, or stop? | Acceptance and next-stage plan |
By day 30, FAT, installation, and SAT must be complete. If data loss, time alignment, units, or notification tests fail, do not start baseline collection on day 31. The illustrative 28-day window runs from day 31 through day 58; days 59–60 are reserved to review and approve its state tags, exclusions, and quality. Scale when the loop works, serious risks are closed, and additional assets fit the validated failure modes. Modify when the value hypothesis remains but data, state tags, or workflow needs correction. Stop when the dominant failures are not visible, work cannot be closed, or security and recovery conditions cannot be met. A well-supported stop is a useful PoC result.
Use conservative economics as a decision aid
This is an illustrative 12-asset planning scenario, not a price benchmark, customer result, or promise.
| Cost | Assumption (THB) | Arithmetic |
|---|---|---|
| Sensors, gateways, software | 420,000 | Planning allowance |
| Integration, training, FAT/SAT | 280,000 | Planning allowance |
| Contingency | 105,000 | (420,000 + 280,000) × 15% |
| Internal review | 93,600 | 6 h/week × 13 × THB 1,200/h |
| 90-day total | 898,600 | 420,000 + 280,000 + 105,000 + 93,600 |
If two events each avoid five hours at an assumed THB 70,000/h, gross avoided loss is 2 × 5 × 70,000 = THB 700,000. If 12 unnecessary replacements of THB 12,000 are avoided, that adds THB 144,000. Illustrative gross benefit is THB 844,000; 844,000 − 898,600 = negative THB 54,600 at day 90. Do not force a positive ROI. The PoC also purchases evidence about detectability, alert workload, closure, and recovery.
| Decision dimension | Scale | Modify and continue | Stop or use another approach |
|---|---|---|---|
| Failure-mode fit | Dominant failures have observable precursors | Additional signals are required for part of the scope | Dominant failures are outside the measured phenomena |
| Data quality | Stable and reproducible | Correct communications or state tags | Cannot remain trustworthy in operation |
| Team workload | Alerts are handled within the agreed SLA | Redesign suppression or ownership | Work remains chronically overdue |
| Security and recovery | Requirements are met | Limited corrective actions remain | Material risk remains unresolved |
| Economics | Still reasonable under changed assumptions | Extend observation before committing | A clearly better alternative exists |
The official BOI/OSOS 1H 2026 release reports approximately THB 13.1 billion of machinery, automation, and robotics applications across 82 projects, and 132 Smart and Sustainable Industry applications worth approximately THB 17.2 billion. This is market context only. Eligibility and approval are project-specific; confirm them individually and do not budget an incentive before approval.
Ten questions for a CBM supplier
- Which failure modes are included and excluded?
- What signal, mounting, sample, and operating-state assumptions apply?
- Exactly where does ISO 20816-3 apply?
- What are the baseline, exclusion, and relearning rules?
- Who approves dynamic-threshold changes, and can versions be rolled back?
- How are false alerts, misses, and insufficient evidence reported?
- How does an alert reach a closed, post-repair work order?
- What are the OT zones, data flows, remote-access rules, and log retention?
- What is backed up and restored during FAT/SAT?
- What evidence decides scale, modify, or stop after 90 days?
Frequently asked questions
Is condition-based maintenance the same as predictive maintenance?
No. CBM uses observed condition to trigger maintenance. Predictive analysis may estimate future condition or remaining life. CBM can be valuable without a remaining-life model if its decisions and work closure are controlled.
Should maintenance DX start with sensors?
Start with event history, criticality, failure modes, and the work process. Choose a sensor only after defining the precursor that must be measured.
Can motor vibration thresholds come only from ISO values?
No. ISO 20816-3 is useful within its stated machinery scope. Combine applicable guidance with mounting, operating state, manufacturer information, and a machine-specific baseline.
How many false alerts are acceptable?
There is no universal number. Set a PoC planning target based on criticality, shift coverage, and review capacity. Classify causes rather than merely disabling alerts.
Is a PoC a failure if no real fault occurs?
Not necessarily. You can verify data quality, known-event replay, workflow timing, restoration, and user competence. Do not claim an observed detection rate or savings without real events.
Does BOI support automatically apply to CBM?
No. BOI figures provide investment context; project eligibility and approval require individual confirmation.
Conclusion
Successful condition-based maintenance connects criticality and failure physics to state-aware baselines, controlled thresholds, actionable alerts, post-repair verification, segmented OT architecture, and tested recovery. FAT/SAT and a 90-day PoC should produce evidence for all three valid decisions: scale, modify, or stop.
TOMAS TECH can help a Thailand or ASEAN factory structure asset selection, failure-signal mapping, FAT/SAT, and 90-day acceptance criteria before a product is fixed. Share the current event history and network constraints through our contact page if an implementation review would help.