Blog

2026.07.30

Predictive Maintenance System 2026: Cost Layers and How to Start

Predictive Maintenance System 2026: Cost Layers and How to Start

Few things weigh on a plant manager more than a phone call at two in the morning. The compressor has stopped. The hydraulic pump on the injection moulding machine is making a strange noise. Tomorrow’s shipment will not be ready. At Japanese-affiliated factories in Thailand and Vietnam, calls like these are routine. At the same time, the veteran maintenance engineers who knew every machine by ear are leaving through repatriation or retirement, and the transfer of that skill to local staff is only half finished. A predictive maintenance system is the mechanism that shifts your organisation from “running around after something breaks” to “planning before something breaks.” This article works through what predictive maintenance is, how the cost breaks down, which equipment to start with, and the failure patterns that show up again and again on the shop floor — all from the practical viewpoint of an ASEAN production site.

What Is a Predictive Maintenance System? Breakdown, Preventive and Predictive Compared

The three maintenance approaches

Maintenance is easiest to organise if you divide it into three approaches.

Breakdown maintenance (BM) means fixing the machine after it fails. The budget is easy to plan, but you do not get to choose when the stoppage happens. If it lands in the middle of a production plan, the lost production opportunity costs far more than the repair itself. BM is a rational choice for low-criticality equipment, for machines with a standby unit, and for machines that can be recovered in a few minutes. It should not be applied to bottleneck equipment that takes the whole line down.

Preventive maintenance (PM) means replacing parts or servicing the machine before it fails, based on elapsed time or accumulated running hours. When the trigger is the calendar or the hour meter — “every three months,” “every 2,000 operating hours” — it is called TBM (time based maintenance). Most factory maintenance schedules are built on TBM. It is highly reliable, but because parts that still have plenty of life are replaced uniformly, both part costs and labour hours tend to be excessive. Conversely, if degradation runs faster than assumed, the machine fails before the next scheduled service arrives.

Predictive maintenance (PdM) means measuring the condition of the equipment continuously and acting when signs of degradation appear. Because the trigger is condition rather than time, it is used in much the same context as CBM (condition based maintenance). Strictly speaking, CBM refers to the general principle of judging by condition, while PdM usually refers to a further step: analysing the trend to forecast roughly when and how the machine will fail. Some plants call it “early-warning maintenance,” but in practice the terms can be treated as equivalent.

The difference between preventive and predictive maintenance, in one line

The difference between preventive and predictive maintenance is whether the decision is based on time or on condition. Preventive maintenance says, “six months have passed since the last replacement, so replace it.” Predictive maintenance says, “the vibration level has been creeping up steadily for three months, so replace it at the next planned shutdown.”

The point most often misunderstood is that predictive maintenance does not replace preventive maintenance. In a real factory you will use all three approaches, selected according to how critical each machine is. Adopting a predictive maintenance system does not mean fitting sensors to every machine. In fact, a plan to instrument everything is a textbook example of the failure patterns described later in this article.

ISO 17359 and the idea of a “programme”

To make equipment failure prediction work as a system rather than as an experiment, it helps to know the international standards framework.

  • ISO 17359 is the high-level standard that defines the framework for a complete condition monitoring and diagnostics programme. The important point is that it is not limited to vibration. It is designed as a way of thinking that applies to any condition measurement technique, including oil analysis, thermography (thermal imaging) and acoustic emission. It asks you to define, as a programme, the whole sequence: which equipment is in scope, which parameters are measured, on what criteria a judgement is made, and who does what as a result.
  • The ISO 13373 series deals with vibration condition monitoring in more concrete terms. Part 1 covers measurement procedures; Part 2 covers data processing, analysis and presentation.
  • The ISO 20816 series specifies vibration acceptance levels — vibration severity. It is the starting point when you ask, “at this measurement point on this machine, how many mm/s is still acceptable?”
  • ISO 18436-2 specifies the certification of personnel who perform vibration diagnosis. In other words, the standards bodies themselves do not assume that simply installing hardware produces a judgement.

When companies evaluate predictive maintenance vendors, the discussion tends to drift towards sensor accuracy and AI algorithms. What ISO 17359 points out is the obvious but easily forgotten fact that predictive maintenance is a programme — a way of working — not a tool. Who looks at the numbers, who makes the call, who orders the parts, and who negotiates the timing of the shutdown with production planning? If those roles are not assigned, no amount of hardware performance will produce results.

Predictive Maintenance System 2026: Cost Layers and How to Start - figure 1

Why Interest in Predictive Maintenance Systems Is Rising in 2026

Downtime is getting worse per event, not in event count

Siemens’ report “The True Cost of Downtime 2024” is one of the most frequently cited studies in this field. According to the report, the world’s top 500 companies by revenue lose approximately USD 1.4 trillion a year to unplanned downtime, equivalent to about 11% of their total revenue. Using the conversion assumptions in this article (1 USD = 32 THB and 1 THB = around 4.5 JPY, approximate as of July 2026, therefore 1 USD = 144 JPY), that is on the order of JPY 200 trillion.

There is also a figure with a more tangible shop-floor feel. At major automotive plants, the cost of a production line stoppage is reported to reach up to USD 2.3 million per hour. On the same assumptions that converts to roughly 73.6 million THB per hour, or about 331 million JPY. This is an extreme case — a final vehicle assembly line — so it should not be applied directly to your own site. Compared with the in-house estimate later in this article, it is more than three orders of magnitude larger. Still, it is a useful number for demonstrating to management why you need to know what one hour of downtime costs.

More interesting is the change in the character of stoppages. The same report puts unplanned downtime at large plants at 25 events per month (42 in 2019) and 27 lost hours per month (39 in 2019). Calculating the average duration per event, 2019 gives 39 ÷ 42 ≈ 0.93 hours and 2024 gives 27 ÷ 25 = 1.08 hours — meaning each event is now about 16% longer (calculated with rounding at the third decimal place). And that is despite the event count falling by about 40% and total downtime hours falling by about 31%.

This structure matches what maintenance teams report. Basic equipment reliability has improved and minor trouble has decreased. But when a machine does stop, root-cause investigation on an increasingly complex machine takes longer, dedicated spare parts are not immediately available, and recovery drags on. The main battleground of downtime reduction has moved from “reduce the number of events” to “stop each event from dragging on — and better still, convert it into a planned shutdown.” That is the situation in 2026.

The gap between wanting to and actually doing it

Adoption, however, is far from smooth. According to maintenance statistics published by MaintainX (2026 edition), the share of organisations practising predictive maintenance actually slipped slightly, from 30% in 2024 to 27% in 2025. At the same time, 65% of maintenance teams said they plan to use AI by the end of 2026.

In other words, there is a clear gap: intent is high, implementation is not following. Where does the gap come from? In our experience on the ground in Thailand, the causes are less technical than organisational, and three stand out. First, nobody has calculated the company’s own cost of one hour of downtime, which is exactly the number the investment decision needs. Second, the PoC (proof of concept) is owned by the production engineering department, while the maintenance department that will actually operate the system is not involved at the planning stage. Third, it is unclear who has the authority and responsibility to set alarm thresholds. None of these are technology problems; all three are internal alignment problems.

The Siemens figure that nearly half of surveyed companies now have a dedicated predictive maintenance team, double the 2019 level is the mirror image of this. Companies that get results are investing in organisation, not in tools.

Thailand today: the phase is extending existing assets, not building new plants

Conditions in Thailand are also pushing predictive maintenance up the agenda. As The Nation Thailand reported in May 2026, factory closures in the first quarter of 2026 reached 156 (up 11.4% year on year) while new openings were 139 (down 63.9%) — the first time in ten quarters that closures exceeded openings. Working backwards from the reported rates of change, the previous year’s figures were about 140 closures and about 385 openings (139 ÷ 0.361). New openings shrank to roughly a third in a single year.

Yet over the same period, expansion investment at 106 existing sites reached 152.5 billion THB. Against 8.5 billion THB in the same quarter a year earlier, that is roughly 18 times higher. Of that, machinery investment was 52 billion THB, about 34% of total expansion investment. On this article’s conversion assumptions, 152.5 billion THB is about 686.3 billion JPY and 52 billion THB is about 234 billion JPY. Closures are concentrated in metal products and plant-based products; expansion is concentrated in plastics, food, machinery, electrical equipment and electronics.

The contrast is clear. Thai manufacturing has moved from a phase of building new factories to a phase of earning more from the factories that already exist. Lending to SMEs has now been negative for 13 consecutive quarters, which reinforces the same trend. If it is hard to secure budget for wholesale equipment replacement, competitiveness comes from keeping existing machines running and using them for longer. That is the context in which predictive maintenance in Thai factories has become a practical topic rather than a theoretical one.

The macro outlook is cautious too. The World Bank’s Thailand Economic Monitor (February 2026) projects growth of 1.6% in 2026 and about 2.3% in 2027. When you cannot assume a large expansion in demand, the case for investment payback is scrutinised more severely — which is precisely why building the ROI from your own cost-per-hour of downtime, as described later, matters so much.

Vietnam: a different problem created by growth

Vietnam is a different story. According to reporting based on the S&P Global survey, the manufacturing PMI for June 2026 was 51.8 (52.8 in May), remaining above the 50 no-change mark, and output expanded for the 14th consecutive month. According to the National Statistics Office, the index of industrial production (IIP) in the second quarter of 2026 rose 11.2% year on year, with manufacturing up 11.3% — the strongest first-half growth since 2019.

In a plant running at high utilisation, it is hard to secure a slot for a planned shutdown, so the machines never get to rest. Scheduled TBM servicing gets pushed to “next month,” and the risk of unplanned stoppages accumulates. Growth is exactly what makes it impossible to stop the line — and that is the background to the rising number of predictive maintenance enquiries in Vietnam.

The Architecture of a Predictive Maintenance System: Sensor, Collection and Diagnostic Layers

Once you start evaluating predictive maintenance systems, it is easy to get lost in feature-by-feature comparison of product catalogues. To keep sight of the whole picture, we recommend thinking of the system as three layers: the sensor layer, the collection layer and the diagnostic layer. Responsibility boundaries, cost ownership and troubleshooting all break down along exactly these same lines.

Layer 1: the sensor layer (what you measure)

This is what gets mounted on the machine and converts a physical quantity into an electrical signal. Machine diagnosis using vibration sensors is the core of predictive maintenance, but it is not the whole of it.

  • Vibration: the technique that carries the most information for rotating machine condition monitoring. Accelerometers capture the vibration waveform, and frequency analysis lets you separate unbalance, misalignment, bearing damage, broken gear teeth and so on. Motor condition monitoring by vibration is the most common entry point into predictive maintenance.
  • Temperature: temperature sensor monitoring is used for bearing temperature, motor winding temperature, oil temperature, cabinet internal temperature and similar points. It is inexpensive, easy to install, and intuitive for the shop floor to explain when something goes wrong. However, by the time temperature rises, damage has often already progressed, so warning tends to come later than with vibration.
  • Current: estimating load abnormalities or rotor bar damage from the motor current waveform. It is useful where a sensor cannot be attached directly to the rotating part, and because the measurement is taken inside the panel, installation is easier.
  • Oil analysis: detecting wear progression from metal particles and viscosity changes in the lubricant. Effective on gearboxes and large hydraulic systems.
  • Ultrasound and acoustic emission: used to detect air leaks, compressor valve faults and early-stage bearing damage. Compressed-air leak surveys are one application where the payback is easy to see.
  • Thermography: capturing heat generation at switchgear connection points, or localised overheating on a motor, as a two-dimensional image.

You do not need all of them. What ISO 17359 asks for is a specific order of work: identify the failure modes first, then choose the parameter in which those failure modes appear. The disappointing outcome — “we fitted vibration sensors, but the failure that actually occurred was seal degradation on an air cylinder” — is what happens when that order is reversed.

Layer 2: the collection layer (how you move the data)

This layer gathers sensor values and delivers them to the diagnostic layer. It is the most underestimated layer in a predictive maintenance system, and the one where installation costs pile up.

  • Edge gateways: receive sensor data, pre-process it if required, and forward it upstream. Raw vibration waveforms are large, so the usual design is to perform FFT (frequency analysis) and RMS calculation at the gateway and transmit only the summarised features. This keeps both data volume and cloud costs down.
  • Wireless or wired: for retrofitting onto existing machines, wireless is overwhelmingly advantageous. But a factory is a hostile RF environment full of metal walls and machinery, and proceeding on the assumption that “it should reach” without a prior site survey guarantees rework. Battery life on wireless sensors also varies greatly with sampling interval and transmission frequency, so both must be decided at design time.
  • Taking data from existing equipment: if the machine already has a PLC or an inverter, you can often extract run signals and operating conditions from it. A vibration value alone cannot tell you whether the reading is low because the machine is stopped or because it is running normally and stably. Capturing a running/not-running signal alongside the vibration data has a large effect on predictive maintenance accuracy.
  • Network segregation: whether predictive maintenance traffic may share the production network is a point that must be agreed with the IT department. We cover this in detail in our practical guide to OT security.

Layer 3: the diagnostic layer (how you judge, and how you connect it to work)

This layer stores the collected data, detects abnormalities and converts them into human action. Equipment anomaly detection breaks down into three stages.

  1. Threshold judgement: the simplest method — raise an alert when a measured value exceeds a predefined limit. The vibration severity levels in ISO 20816 are the starting point for the criteria. Beginning here is the realistic choice in the early phase.
  2. Trend monitoring: looking at the rate of change over time rather than the absolute value. Catching the state described as “still inside the acceptable range, but 1.8 times higher than three months ago” is the use closest to the essence of predictive maintenance.
  3. Model-based anomaly detection: building a machine learning model from normal-condition data and flagging deviations from it as anomalies. It can be started even with few recorded failure cases, but it is hard to explain why a given reading was judged abnormal, which makes it hard to win acceptance on the shop floor.

There is one more element that must be part of the diagnostic layer: digitising maintenance records. Unless a CMMS (computerised maintenance management system) records when, by whom, which part was replaced and why, you cannot link an anomaly detected by a sensor to what the failure actually was. An architecture in which a maintenance planning system and a predictive maintenance system are procured separately and never integrated is one you will always regret later. The fuel that grows an anomaly detection model is not sensor data but correctly labelled maintenance history.

The full flow from sensors to plant-wide visibility is also covered in our factory IoT implementation guide 2026. Predictive maintenance is one domain within factory IoT, and if you already have an operational monitoring platform, it can often be reused as the foundation.

Predictive Maintenance System 2026: Cost Layers and How to Start - figure 2

Predictive Maintenance System Cost: A Five-Layer Breakdown

The order of magnitude to start from

When companies look at predictive maintenance cost, most want to know the price “for the whole package.” But the cost of a predictive maintenance system is determined almost entirely by the number and type of machines in scope, so a package price is meaningless.

As a starting point, here is the general range published by vendors in Japan, presented as an indication only. Retrofitting a vibration sensor onto a rotating machine (motor, pump or compressor) and connecting it to cloud analytics costs, in the minimum configuration, roughly 300,000 to 800,000 JPY up front, with monthly fees starting from the tens of thousands of yen. This is an indication based on figures published by various vendors, not a fixed price.

Broken into the five layers, that minimum configuration is roughly: sensor layer 100,000–250,000 JPY; collection layer 80,000–200,000 JPY; diagnostic layer initial 0–100,000 JPY; implementation engineering 120,000–200,000 JPY; operations and people, initial 0–50,000 JPY. Added together, the total falls within the 300,000–800,000 JPY range.

Breakdown for 10 machines and 20 measurement points

In practice, a project rarely stops at one machine. The table below is presented as an indication for a configuration commonly seen at Thai sites, assuming 10 machines and 20 measurement points (two points per machine). Please treat these as indications, not fixed prices.

LayerMain cost itemsIndicative initial costIndicative running costWhose budget (internal owner)
1. Sensor layerVibration and temperature sensors, wireless nodes, mounting hardware, dust/water protection600,000–2,000,000 JPYBattery replacement and calibration (from tens of thousands of yen per year)Maintenance dept. (equipment budget)
2. Collection layerEdge gateways, wireless APs, cabling and power work, panel modification400,000–1,500,000 JPYCommunication fees (5,000–20,000 JPY per month)Maintenance dept. + IT dept.
3. Diagnostic layerCloud or on-premise platform, licences, CMMS integration200,000–800,000 JPY50,000–200,000 JPY per monthIT dept. or head office IT
4. Implementation engineeringScope selection, measurement point design, installation work, initial threshold setting, commissioning, report design600,000–2,000,000 JPYRecurs each time additional machines are rolled outProject budget (capex application)
5. Operations and peopleMaintenance process redesign, local staff training, annual support, threshold tuning200,000–800,000 JPY300,000–1,200,000 JPY per yearMaintenance dept. (opex) + training budget
Total2,000,000–7,100,000 JPY (approx. 444,000–1,578,000 THB)about 1,000,000–4,000,000 JPY per year (approx. 222,000–889,000 THB)

Note: these figures start from prices published by vendors in Japan. At sites in Thailand and Vietnam they move up or down with installation labour, import costs for the hardware, and whether local support is included. Treat them as a way to sanity-check a local quotation, not as a price list.

The most important message in this table is not the amounts but the fact that each layer is paid for by a different part of the organisation. The sensor and collection layers come out of the equipment budget, the diagnostic layer out of the IT budget, and operations and people out of operating expense. Most predictive maintenance projects that fail to get approval fail not for technical reasons but because this split of budget ownership was never sorted out in advance.

Three things to check when you read a quotation

When you collect quotations from several vendors, the reason for the price differences can be hard to see. Checking the following three points puts everyone on the same footing.

1. Is implementation engineering included?

A quotation covering only hardware and licences looks cheap. But if measurement point design, installation, initial threshold setting and commissioning are not included, that work becomes internal labour hours. Initial threshold setting in particular requires normal-condition data to be collected over a period first, which means one to three months of operation after installation. Make sure the engineering support during that period is in the quotation.

2. What exactly is the running fee charged against?

Platform pricing is usually based on measurement points, data volume, user count or feature tiers, and different bases produce very different totals a few years later. Always ask, at quotation stage, for the monthly fee if you later expand to 20 machines and 40 measurement points. Whether the unit price falls or simply doubles in proportion can create a five-year difference larger than the initial cost.

3. Removal, relocation and replacement costs

Layout changes are common in ASEAN factories. Check the cost of moving a sensor to a different machine, whether licences can be transferred, and the export format of accumulated data (can you get CSV files, or is there an API?). A contract that does not allow you to take your data out is vendor lock-in in all but name.

Building internal agreement

A predictive maintenance capex application has a structural difficulty: you have to present “the failure that did not happen” as the result. The following sequence has worked in practice.

Start by collecting two years of unplanned stoppage records. In most factories they are scattered across daily reports and shift handover notebooks. Aggregate them by machine and by cause, and produce a list of how many times and for how many hours each machine stopped the line. This exercise is, in fact, the single most valuable step in the project. Once aggregated, the machines at the top of the list are often not the ones you expected.

Next, agree the cost per hour of downtime with finance and production planning. If the technical department decides this alone, the application comes back with “where does that unit cost come from?” Agree in writing, before the capex application, whether the calculation is based on contribution margin or revenue, and whether labour costs incurred during the stoppage are included.

Finally, apply for a small number covering a narrow scope. A plant-wide rollout with a large figure raises the approval hurdle sharply. As in the ROI calculation in the next section, narrowing the scope and showing a payback period is both easier to get approved and more likely to produce a real result.

Getting Results: How to Choose Which Equipment to Start With

Four criteria for narrowing down target equipment

Whether a predictive maintenance system produces results is decided almost entirely by which machines you select. Start with machines that satisfy all four of the following criteria.

Criterion 1: if it stops, the line stops (criticality)

Machines with no standby unit and no bypass route, where one unit stopping halts production as a whole. Conversely, machines with a spare that can be switched in can be given lower priority.

Criterion 2: there is a track record of unplanned stoppages (frequency)

Choose machines that have stopped unexpectedly more than once in the past two years. If you select a machine because “it has never failed but it worries me,” you will not be able to measure any effect — and a project whose effect cannot be proven will not get the next budget.

Criterion 3: degradation progresses gradually (predictability)

Failure modes that develop over time — bearing wear, growing unbalance, belt slackening — are suited to prediction. By contrast, sudden death of an electronic board, burnout from an instantaneous overcurrent, and damage from external causes are difficult to catch in advance with sensors. Rotating machines, compressors, injection moulding machine hydraulic pumps, and conveyor motors and gearboxes are usually the first candidates because they satisfy this criterion easily.

Criterion 4: a sensor can physically be attached (installability)

High-temperature areas, surfaces that are constantly washed down, explosion-proof zones, and moving parts where cabling cannot be routed all raise the installation hurdle. Confirm in advance whether wireless reaches, whether power is available, and whether the mounting surface is flat.

Points to watch by equipment type

Equipment typeMain failure modesEffective measurementsEase of implementation
Compressor (air supply)Bearing wear, valve faults, air leaksVibration, temperature, ultrasound, currentHigh (easy to access, and a stoppage affects every line)
Hydraulic pump on moulding machinePump wear, oil degradation, filter cloggingVibration, oil temperature, oil analysis, pressureMedium (heat and oil environment need care)
Conveyor motors and gearboxesBearing damage, gear wear, misalignmentVibration, temperatureHigh (clear measurement points, easy to interpret)
Pumps and blowers (utilities)Cavitation, bearing wear, impeller wearVibration, current, flow rateHigh (a utility stoppage has wide impact)
Machine tool spindlesBearing damage, thermal displacementVibration, temperatureMedium to low (must be considered together with quality impact)
Control panels with electronic boardsSudden failure, capacitor degradationPanel internal temperature, thermographyLow (prediction is difficult)

For a first target, we consider compressors and utility pumps to be the safest choices. Their failure has wide impact, their failure modes are easy to capture with vibration and temperature, and because you are not touching the production line itself, resistance from the production department is low — a very practical advantage.

How to think about ROI: build it as a formula

When explaining payback, build the argument in three simple steps.

Formula 1: loss per hour of downtime

“`

Loss per hour of downtime

= (units produced per hour × contribution margin per unit)

+ (labour cost incurred even while stopped)

“`

Formula 2: annual saving

“`

Annual saving

= annual unplanned downtime hours caused by target equipment

× reduction rate

× loss per hour of downtime

“`

Formula 3: payback period

“`

Payback period (years) = initial investment ÷ (annual saving − annual running cost)

“`

A worked example (conversion assumptions: 1 USD = 32 THB, 1 THB = around 4.5 JPY, approximate as of July 2026)

Consider a mid-sized factory in Thailand with a scope of two compressors, four moulding machine hydraulic pumps and four conveyor motors — 10 machines and 20 measurement points. The figures below are placeholder assumptions; replace them with your own.

Deriving the loss figure

  • Units produced per hour on the target line: 1,200 pieces
  • Contribution margin per unit: 35 THB
  • Lost production profit: 1,200 × 35 = 42,000 THB per hour
  • Labour cost incurred while stopped: 12 people × 200 THB per hour = 2,400 THB per hour
  • Loss per hour of downtime: 44,400 THB (about 199,800 JPY, roughly 200,000 JPY)

For reference, compared with the automotive plant figure of USD 2.3 million per hour (about 73.6 million THB) cited earlier, this estimate is roughly one 1,660th of that scale. The gap illustrates exactly why you cannot borrow another company’s case-study numbers.

Current losses

  • Unplanned stoppages attributable to the target equipment: 14 times a year, averaging 3.5 hours each → 49 hours a year
  • Annual loss: 49 × 44,400 = 2,175,600 THB = about 9.79 million JPY per year

Payback by reduction rate (assuming 3.5 million JPY initial investment and 1.8 million JPY annual running cost)

The 3.5 million JPY initial figure sits in the lower-to-middle part of the table in the previous section. The 1.8 million JPY annual running cost covers the main items only: 120,000 JPY per month for the diagnostic layer (1.44 million JPY a year) plus 360,000 JPY a year of support. Communication fees and battery replacement or calibration are not included in this 1.8 million JPY; add them from the table in the previous section when you build your own case.

Reduction in unplanned downtimeHours saved per yearAnnual savingAnnual benefit after running costPayback on 3.5 million JPY
20%9.8 hoursabout 1.96 million JPYabout 160,000 JPYabout 22 years (effectively no payback)
40%19.6 hoursabout 3.92 million JPYabout 2.12 million JPYabout 1.7 years (about 20 months)
60%29.4 hoursabout 5.87 million JPYabout 4.07 million JPYabout 0.9 years (about 10 months)

Note: amounts are rounded to the nearest 10,000 JPY. Savings are calculated as “hours saved × 44,400 THB × 4.5 JPY.”

How to read this table

The most important row is not the middle one — it is the top one. If the reduction rate stops at 20%, the annual saving is about 1.96 million JPY while running costs are 1.8 million JPY, so the net benefit is only about 160,000 JPY and recovering the initial investment becomes effectively impossible. The investment decision on a predictive maintenance system has a structure in which success and failure swap places within the narrow band between a 20% and a 40% reduction.

And what determines the reduction rate is not sensor performance. It is whether, when an alert appears, the action “replace this part at the next planned shutdown” is actually carried out, reliably and without delay. With the same system installed, the gap in reduction rate is wide between factories that rebuilt the work process far enough to reschedule maintenance around an alert, and factories that merely receive alerts by email (we are not aware of a published statistic that quantifies this gap; this is what we have observed on site). Understand it this way: the moment you cut investment in layer 5, operations and people, you fall into the top row of this table.

Attach the sensitivity analysis to the capex application

We recommend attaching the sensitivity table above, with the reduction rate varied, directly to the approval document. Rather than asserting “we will achieve a 40% reduction,” it is more persuasive to management to say: “at 20% we cannot recover the investment; to reach 40% we will redesign the maintenance process and train local staff as part of the same project.” It also makes explicit inside the company that the success measure for the project is not “the system is running” but “downtime hours are reduced.”

Five Common Failure Patterns in Predictive Maintenance Projects, and How to Avoid Them

Failure 1: the project never gets past the PoC

The most common failure of all. A sensor is fitted to one machine, data is collected for six months, and the project ends with a report stating, “we confirmed that anomalies can be detected.” No further budget is approved, and the sensor is left in place, unused.

Why it happens: the PoC success criterion was set as “can we detect it technically?” Technically, you usually can. But what management wants to know is how much money was gained.

How to avoid it: before the PoC begins, define in numbers the criteria for proceeding to full deployment. For example: “if during the PoC period we detect two or more precursors that lead to an actual failure, and convert at least one of them into a planned shutdown, we proceed to full deployment.” In addition, choose PoC target equipment that satisfies Criterion 2 above (a track record of unplanned stoppages). If you run a PoC on a machine that has not failed in two years, nothing happening for six months is the expected outcome.

Failure 2: thresholds never get settled, and alerts lose credibility

A problem that typically appears a few months after go-live. Set the threshold tight and false alarms multiply until the shop floor starts ignoring them. Set it loose and you miss the real anomaly.

Why it happens: initial thresholds were set purely from the equipment maker’s generic values or standard recommendations, without adjustment against real data from your own machines. Even for identical motor models, normal vibration levels vary considerably with foundation stiffness, load conditions and vibration from nearby equipment.

How to avoid it: in the early phase, it is effective to accept a trade-off and make “recording the trend” the objective rather than “raising alerts.” Position the first two to three months as a period for accumulating normal-condition data, and understand the seasonal variation (in Thailand, the temperature and humidity difference between dry and rainy seasons) and the variation caused by product changeovers. Then set thresholds from the distribution of your own data, referring to the ISO 20816 vibration severity levels. Beyond that, make threshold review a standing quarterly meeting, recording both “alerts that turned out to be false” and “failures that were missed,” and adjust accordingly. Whether that recurring meeting exists is what creates the difference a year later.

Failure 3: the maintenance process is never changed

The system is running, but the way maintenance work is done is exactly the same as before. Alerts are being raised, but nobody looks at them — or people look at them but the next action has not been defined.

Why it happens: predictive maintenance was treated as an IT implementation project rather than as a business process reform project.

How to avoid it: at the same time as implementation, decide the following four points in writing.

  1. Who sees the alert first (primary and backup owner, and how out-of-hours alerts are handled)
  2. Alert severity levels and the response deadline for each level (for example: level 1, check at the next scheduled inspection; level 2, inspect on site within one week; level 3, immediate stop/no-stop decision)
  3. Authority to order parts (the value ceiling for ordering parts in advance at the precursor stage, and who approves it)
  4. Who negotiates conversion to a planned shutdown with production planning

Point 3, ordering parts at the precursor stage, is particularly easy to overlook. If you detect the anomaly but the part takes six weeks to arrive, half the value of the prediction is gone. The effect of predictive maintenance is determined by detection accuracy × speed of response.

Failure 4: underestimating retrofit work on existing machines

Retrofitting sensors onto existing equipment is sometimes described as “just sticking a sensor on,” but the actual installation work is not that simple.

Why it happens: the site survey at quotation stage was inadequate.

What actually happens on site:

  • The mounting surface is not flat because of paint or rust, so the accelerometer does not couple properly (high-frequency content in the vibration data is not captured correctly)
  • There is no power source near the sensor, so cabling from a panel is required
  • Wireless is blocked by a metal partition, so a repeater has to be added
  • The equipment maker’s warranty conditions include “no modifications,” and confirming whether mounting is permitted takes several weeks
  • In an explosion-proof zone, switching to explosion-proof rated sensors sends the cost up sharply
  • The machine has to be stopped for installation, so no work can be done until the next planned shutdown three months away

How to avoid it: always carry out a site survey before placing the order, and compile the mounting method, power source and communication route for every measurement point into a single table. A quotation issued without that table will always generate additional costs. And because installation has to coincide with planned shutdowns, the project schedule must be reconciled against the production plan. On retrofit IoT projects, it is usually securing an installation window — not technical difficulty — that determines the schedule.

Failure 5: nobody owns the data

This one is more serious than it first appears: no decision has been made about who manages the collected data, who can access it, and how it is handed over when someone resigns or transfers.

Why it happens: at implementation the project proceeds on “let’s just store it in the cloud for now,” and no responsible owner is named.

Where it becomes a concrete problem:

  • It is unclear whether the local subsidiary or the Japanese head office is responsible for the data, so improvement instructions from head office do not land locally
  • On terminating the vendor contract, several years of accumulated data cannot be extracted
  • The local IT staff member resigns and the administrator account password is lost
  • Data containing a customer’s production information may conflict with the confidentiality agreement signed with that customer

How to avoid it: at implementation, document four things — the data owner (responsible department), the access rights list, the handover procedure on resignation or transfer, and the data return format at contract termination. In particular, manage administrator accounts under a role name or shared account rather than an individual’s name, and set a rule for password storage. This is also mandatory housekeeping from an OT security point of view.

Predictive Maintenance System 2026: Cost Layers and How to Start - figure 3

Implementation Steps: What to Do in the First 90 Days

Here is the implementation of a predictive maintenance system organised into the first 90 days. Reliably bringing one machine online in 90 days and correcting the design with what you learn beats spending a year perfecting the overall design before you move — and it gets you to results sooner.

PhasePeriodMain activitiesCompletion criteria (deliverables)
Phase 0: understand the current stateDays 1–15Collect two years of unplanned stoppage records and aggregate by machine and by cause. Take stock of the existing maintenance schedule and CMMSA table of stoppage counts and stoppage hours by machine
Phase 1: select scope and agree the loss figureDays 16–30Narrow down target equipment using the four criteria. Agree the cost per hour of downtime with finance and production planningTarget equipment list (3–10 machines) and a written agreement on the loss figure
Phase 2: site survey and designDays 31–50Design measurement points, confirm mounting method, power and communication route, run a wireless site survey, agree the network policy with the IT departmentMeasurement point table, installation specification, network diagram
Phase 3: procurement and installationDays 51–70Procure equipment, install during a planned shutdown, verify communicationsData being received from every measurement point
Phase 4: baseline capture and operational designDays 71–90Begin accumulating normal-condition data. Document alert severity levels and response rules, train local staffOperating rules document and a list of trained personnel
Phase 5: threshold setting and full operationDay 91 onwardsSet thresholds from accumulated data, start the quarterly review meeting, measure resultsRecorded reduction in downtime hours (monthly report)

The phase most often skipped in this 90-day plan is Phase 0. It gets dropped because “the records are scattered and aggregating them takes too long,” but skipping it removes the basis for selecting target equipment and makes later measurement of results impossible. Think of it this way: the two weeks spent aggregating past stoppage records are the highest-return two weeks in the entire predictive maintenance project.

Also, because the Phase 3 installation depends on a planned shutdown, it may not fit inside the 90-day window. In that case, complete Phases 0 to 2 first and use the wait for an installation slot to progress the operational design (part of Phase 4), so that no time is wasted.

Practical Issues That Deserve Extra Attention at Thai and ASEAN Sites

Factor BOI and depa support schemes into your plan

Thailand offers public support for capital investment and human resource development. BOI-certified companies can claim a 200% tax deduction on approved training expenses. In other words, twice the amount actually spent on training can be recorded as a deductible expense, reducing taxable income accordingly. Because, as discussed above, investment in “operations and people” is what decides success in predictive maintenance, preferential treatment of training costs has a substantial real-world effect.

In addition, depa grants of up to 1 million THB (about 4.5 million JPY on this article’s conversion assumptions) are available for advanced training and tools, and eligible hardware includes IoT sensors, edge computing devices and industrial robots. That overlaps with layer 1 (sensors) and layer 2 (collection) of a predictive maintenance system.

However, this information is current as of July 2026, and eligibility must always be confirmed against the latest BOI and depa announcements and through individual review. Scheme content and eligible scope are revised over time, so please do not base an investment decision on the description in this article. In practice, the safe approach is to confirm eligibility in parallel with collecting quotations, and to prepare a budget plan that also works if the grant is not awarded.

Power and network conditions differ from Japan

If you design on the same assumptions as a Japanese factory, you will hit the following issues in Thailand and Vietnam.

  • Voltage dips and outages: momentary voltage dips caused by lightning in the rainy season are not unusual. Assume a UPS for edge gateways and servers, and confirm that the configuration restarts automatically and resumes measurement after a power failure. If recovery requires a manual restart, you will end up with weeks of missing data that nobody noticed.
  • Surge protection: consider surge protection devices on outdoor cabling and long signal runs. It is usually the gateway and communication equipment, rather than the sensor itself, that gets damaged.
  • Earthing: in older plants, earthing is sometimes inadequate, which becomes a source of noise and makes readings unstable. It affects current measurement and communication stability more than vibration measurement.
  • Line quality: with a cloud-based system, the design must ensure that an internet outage does not become a measurement outage. Confirm that the gateway buffers a certain period of data and transmits it in bulk once the connection is restored.
  • Heat, humidity and dust: internal panel temperatures run higher than expected, which affects equipment life. For installations close to outdoor areas, it is safer to specify one grade higher in dust and water protection rating.

Multilingual operation of the local maintenance team

At factories in Thailand and Vietnam, Japanese managers, local maintenance leaders and operators do not share a single working language. A predictive maintenance system whose screens are only in English — or only in Japanese — will stumble at the very first step of operation.

In practice, deciding the following three points at implementation makes operation far more stable.

  1. Display language: the screens local staff use daily (alert lists, inspection record entry) should ideally display in the local language, while the aggregated screens head office looks at can be in Japanese or English. That split is the realistic arrangement.
  2. A glossary: build a cross-language table of the terms that carry a judgement — “abnormal,” “caution,” “under observation,” “trend monitoring” — in Japanese, English and the local language. Without it, people’s perception of how serious the same alert is will diverge.
  3. Scope of training: as ISO 18436-2’s certification of vibration analysis personnel implies, interpreting vibration data requires a degree of training. Rather than demanding advanced diagnostic ability from everyone, a two-tier structure works better and is far more resilient to staff turnover: everyone understands who to escalate an anomaly to, while actual diagnosis is handled by designated specialists.

Designing reporting to the Japanese head office

Most Japanese-affiliated plants in ASEAN are subsidiaries whose capital investment is approved by a parent company in Japan, so how you design head office reporting largely determines whether the project survives.

The purpose of reporting is not to show that the system is running; it is to show that the investment decision was correct. A monthly report is therefore effective when it includes downtime hours for the target equipment (compared with the pre-implementation baseline), the number of alerts detected and the action taken, the number converted into planned shutdowns, and the estimated loss avoided as a result. All of these can be calculated directly with the ROI formulas in the previous section.

Conversely, it is better not to send raw vibration graphs to head office. They cannot be interpreted there, the questions come back to the site, and the only outcome is more work. The division of labour that works is: judge locally, and report the conclusion and its monetary value to head office.

OT security

Installing a predictive maintenance system means adding a new communication path to your production equipment. If the gateway connects to the internet and talks to a cloud service, that is a new attack surface.

At minimum, confirm the following points at implementation.

  • On which device, and by what mechanism, the production network and the predictive maintenance network are segregated
  • Whether the gateway-to-cloud communication is designed to be outbound only
  • Whether the remote maintenance connection path is enabled only when required, rather than left permanently open
  • Who performs firmware updates and when, and whether that procedure is documented
  • Whether the contract states explicitly what range of data the vendor’s engineers can access

We cover these points in detail in our practical guide to OT security. If you submit the predictive maintenance application and the security application separately, the latter can be sent back for revision and stall the whole project, so we recommend designing them as one package from the start.

Frequently Asked Questions (FAQ)

What is predictive maintenance?

It is a maintenance approach that measures equipment condition continuously, catches signs of degradation, and acts before a failure occurs. In English it is called predictive maintenance (PdM). In contrast to preventive maintenance (TBM), which is triggered by time, it is triggered by condition, so it is used in much the same context as condition based maintenance (CBM). Among international standards, ISO 17359 defines the framework for a complete condition monitoring programme.

What is the difference between preventive and predictive maintenance?

The basis for the decision. Preventive maintenance replaces a part because a set period has passed since the last replacement — a time-based trigger. Predictive maintenance replaces it because the trend in measured values indicates that degradation is progressing — a condition-based trigger. Preventive maintenance replaces parts that still have life left, so part costs and labour hours tend to be excessive; predictive maintenance lets you act only when needed, but requires up-front investment in the measurement, judgement and response mechanism. In practice you do not choose one over the other: you apply breakdown, preventive and predictive maintenance selectively according to how critical each machine is.

How much does a predictive maintenance system cost?

There is no package price, because the cost varies greatly with the number and type of machines in scope. As an indication published by Japanese domestic vendors, retrofitting a vibration sensor onto a single rotating machine and connecting it to cloud analytics costs roughly 300,000 to 800,000 JPY up front, with monthly fees starting from the tens of thousands of yen (an indication based on published vendor figures, not a fixed price). For a configuration of 10 machines and 20 measurement points, one indication is 2,000,000–7,100,000 JPY initial and about 1,000,000–4,000,000 JPY per year running. What matters most is that the cost splits into five layers — sensors, collection, diagnostics, implementation engineering and operations — and that each layer is paid for by a different part of the organisation.

Which equipment should I start with?

Start with machines that satisfy all four of these conditions: (1) if it stops, the line stops; (2) it has a track record of unplanned stoppages in the past two years; (3) its failure mode develops gradually; and (4) a sensor can physically be attached. The machines that most often fit are compressors, utility pumps and blowers, moulding machine hydraulic pumps, and conveyor motors and gearboxes. Compressors in particular are an easy first choice: their failure has wide impact, and their condition is straightforward to capture with vibration and temperature.

Can sensors be retrofitted to old machines?

In most cases, yes. External vibration and temperature sensors can be fitted without touching the machine’s control system, so equipment more than 20 years old can still be brought into condition monitoring. In fact, predictive maintenance shows its value precisely when you want to extend the life of existing assets. That said, retrofitting involves many installation checks: flatness of the mounting surface, securing power for the sensor, wireless coverage, the equipment maker’s warranty conditions, explosion-proof zone requirements, and securing a planned shutdown slot for the work. If these are not confirmed measurement point by measurement point in a pre-order site survey, additional costs and schedule delays will follow.

Is it worthwhile for a small factory?

Yes, provided you keep the scope narrow. The fewer machines a factory has, the fewer alternatives exist when one stops, and the more directly the stoppage hits. Even at a scale of one air compressor and two motors on the main line, the return can hold up if that one machine stopping means half a day of no production. The criterion is not the size of the factory but “how much do we lose if that machine stops for one hour?” Start by calculating that figure for your own site. If the figure is small, securing a standby unit or holding spare parts in stock may be a more rational answer than predictive maintenance.

How much will downtime actually fall after implementation?

There is no single answer. In the estimate in this article, a 20% reduction in unplanned downtime makes recovering the initial investment effectively impossible, 40% pays back in about 1.7 years, and 60% pays back in a little under a year. What decides the difference is not sensor performance but whether there is a mechanism that reliably executes the action “replace the part at the next planned shutdown” after an alert is raised. Documenting the response rules and decision authority during the evaluation phase is what moves the reduction rate.

Will AI detect anomalies automatically?

Within limits, yes — but be careful about where you place your expectations. Building a machine learning model from normal-condition data and flagging deviations from it is at a practical stage. On the other hand, it is hard to explain why a reading was judged abnormal, which makes shop-floor acceptance difficult. Improving model accuracy also requires more than sensor data: it requires labelled maintenance history that records what the failure actually was. Digitising maintenance records first is a precondition for using AI. In practice we recommend the order of getting operations running with threshold judgement and trend monitoring, and adding model-based detection afterwards.

How long does implementation take?

If you narrow the scope to 3–10 machines, use roughly 90 days as a guide from understanding the current state to starting full operation. The breakdown is two weeks to aggregate past records, two weeks to select scope and agree the loss figure, three weeks for the site survey and design, three weeks for procurement and installation, and three weeks for baseline capture and operational design. However, because installation must coincide with a planned shutdown, the schedule extends if the next slot is three months away. Full threshold setting requires a further two to three months of normal-condition data, so it is realistic to expect that you can begin measuring results at around the six-month mark.

Summary

A predictive maintenance system is the mechanism for converting unplanned stoppages into planned shutdowns. Here are the key points of this article.

  • Breakdown, preventive and predictive maintenance are applied selectively according to equipment criticality. A plan to fit sensors to every machine is likely to fail. As ISO 17359 makes clear, predictive maintenance is a programme — a way of working — not a tool.
  • The structure of downtime has changed. The number of unplanned stoppages has fallen, but each event now lasts about 16% longer. The centre of gravity has moved from “reduce the number of events” to “stop each event from dragging on.”
  • Thailand has moved from new construction to making more of existing assets. In Q1 2026, factory closures exceeded new openings, while expansion investment at existing sites reached about 18 times the level of the same quarter a year earlier. Keeping the machines you already have from stopping now translates directly into competitiveness.
  • Cost splits into five layers, and each is paid for by a different part of the organisation. Most failed approval applications fail because of that split in budget ownership, not because of technology.
  • The boundary between success and failure sits between a 20% and a 40% reduction. And what decides which side you land on is not sensor performance but whether the action after the alert actually gets executed. The moment you cut “operations and people” from the five cost layers, the investment stops paying back.
  • In the first 90 days, concentrate on narrowing the scope and bringing one thing online. In particular, the first two weeks spent aggregating two years of stoppage records are the highest-return step in the whole project.

Predictive maintenance is not a mechanism that delivers results automatically once installed. But if you narrow the scope, calculate the loss figure for your own site, and decide the action to take after an alert, the path to payback becomes something you can genuinely explain. Start by calculating your own cost per hour of downtime.

TOMAS TECH is based in Bangkok and supports Japanese-affiliated manufacturers with production management systems, energy management systems, and factory IoT/OT implementation. We are happy to talk at an early evaluation stage — questions such as “which machine should we start with,” “we would like a second opinion on whether this quotation is reasonable,” or “we first want to work out whether we need this at all” are all welcome. We can also start from how to organise your past stoppage records. Please feel free to get in touch through our contact page.

References