“We installed an equipment alert notification system and downtime did not move.” This is the single most common conversation we have with plant managers in Thailand. The cause is almost never that alerts fail to arrive. It is that nobody designed what happens after they arrive — who acts, within how many minutes, and who takes over if that person does not. This article breaks notification into three layers (Detection, Delivery, Record) and shows how to judge the investment by which part of MTTR it actually removes.
What Is an Equipment Alert Notification System, and How Does It Differ From Andon?
An equipment alert notification system is the full chain that automatically detects an abnormality on a machine or process, reliably delivers it to a person who can act, and retains a record of what was done. The same thing circulates under many names — factory call system, shop-floor alerting, operator paging, machine downtime alert — but the underlying object is usually identical.
Andon is something narrower. It originated in the Toyota Production System as a way to stop, call, and wait: a visual signal, usually a stack light or a large display board, that shows line status in colour so that people nearby notice it. As a concept it is close to perfect, and as the many explainers of the basic andon principle attest, it remains the reference point most manufacturing engineers reach for.
Andon Is One Delivery Method, Not the System Itself
This is the starting point for everything that follows. An andon light or board is one way of getting an abnormality in front of a human being. It is not the notification system. A stack light reaches only the people who happen to be looking at it. In a 24-hour plant, if the night-shift maintenance technician is in the building on the other side of the site, the lamp rotates for nobody.
So andon’s weakness is not conceptual. It is physical: limited reach. That gap is exactly what drives searches for andon notifications on smartphones or equipment alerts on smartwatches. The correct framing is not replacing andon with phones. It is adding a second delivery method alongside the one you already have.
The Implementation Range Hiding Behind One Term
What plants call an equipment alert notification system spans a very wide range of maturity.
| Level | What is actually installed | Delivery method | Record kept |
|---|---|---|---|
| L0 | Machine buzzer and stack light only | Local sound and light | None |
| L1 | Stack light plus handwritten shift report | Local, plus paper after the fact | Paper (never aggregated) |
| L2 | Signals collected and shown on a large monitor | Shop-floor monitor | Partial logs |
| L3 | Collected signals pushed to phones/watches | Personal devices | Notification log |
| L4 | Escalation plus ACK plus performance analysis | Personal devices plus staged handoff | Full response history |
Most plants sit at L1 or L2. And a striking number stop as soon as they reach L3, satisfied that alerts now appear on phones. As the cost section below shows, the investment only pays back when you reach L4. Understanding that gap first changes how you read every quotation you receive.
Alerts Arrive but Nothing Moves: Three Failure Patterns
When we break down the plants where the system did not deliver results, the causes are remarkably consistent. They reduce to three.
Failure 1: No Definition of Who, Within How Many Minutes, and Then Whom
By far the most common. An abnormality occurs and every phone in the maintenance group rings at once. When everyone is alerted, everyone assumes somebody else is going. The bystander effect, straight out of social psychology, plays out on the factory floor.
What is needed is not broadcast but escalation design: who takes it first, how many minutes of silence before it moves on, and where the chain terminates. A system where those three things are not defined per equipment group increases alert volume without changing response speed.
As a practical test, if you cannot write down the following three items, the design is incomplete.
- Primary owner — a named person or a named role. “The group” does not qualify
- Response deadline in minutes, varied by the cost of that machine stopping
- Secondary and tertiary escalation targets, and the change of channel that goes with them (watch to phone call, for example)
Failure 2: Too Many Alerts, and Alert Fatigue Sets In
Once a plant builds out to L3, it almost always ends up over-notifying. Every signal the sensors can produce gets loaded onto the notification path because it is there. Once several dozen to over a hundred alerts fly per shift, people first stop reading them, then start leaving the device in the locker.
Alert fatigue is a volume problem, not a quality problem. However correct each message is, past a certain count it will be ignored. And once ignoring has become a habit, rebuilding trust costs more than building the system correctly from scratch, because what you lost was credibility rather than configuration.
Failure 3: Alerts Leave No Record, So Nothing Improves
The third is the most overlooked. The alert went out — but who acknowledged it, when, how long recovery took, and what the cause was are all missing. The result is that the monthly improvement meeting gains no new data at all.
Minor stoppages (*choko-tei* in Japanese) are especially prone to this: because the operator clears them on the spot, they rarely reach any record, and the real picture stays invisible (definition and characteristics of choko-tei, overview of short stops — both Japanese-language references). A notification system is the only realistic mechanism for turning those unrecorded stops into data automatically. Not using it means discarding half of what you paid for.
Splitting Notification Design Into Three Layers: Detection, Delivery, Record

Invert the three failures and the design target splits cleanly into three parts. Throughout this article we call them the Detection layer, the Delivery layer, and the Record layer, and every section from here states which layer it belongs to.
| Layer | What it decides | Main components | What happens when it is weak |
|---|---|---|---|
| Detection | What counts as an abnormality | Dry contacts, PLC tags, current/vibration/temperature sensors, cycle-time monitoring | Missed events, or a flood of false alarms |
| Delivery | Who is reached, and how | Wireless LAN/LTE, notification platform, phones, watches, andon, IP radio | Alerts do not arrive, or arrive and nobody moves |
| Record | What is retained, and how it is used | Notification logs, ACK, response history, cause codes, OEE integration | No feedback into improvement, no way to prove ROI |
Measure by Which Part of MTTR You Are Removing
The reason to adopt three layers is that it lets you measure return on investment with a single consistent ruler. Mean time to repair, from stop to restart, decomposes in practice as follows.
Occurrence → detection → notification → awareness → response starts → repair complete → restart confirmed
Of that chain, a notification system can only shorten the detect-to-response segment. The actual repair time (response start to recovery) is a function of spare parts and technician skill, and no alerting product will compress it. Deploy without understanding that distinction and the verdict becomes “we installed it and downtime barely changed.”
The inverse also holds: the longer your detect-to-response wait, the larger the effect. There is exactly one number worth measuring before you buy. The observed time from abnormality to a person physically arriving at the machine. Time twenty events with a stopwatch and the distribution becomes visible. If the median is under three minutes, this investment is low priority. If it is over five minutes, you have a strong basis for a capital request. The three-to-five-minute band in between is the zone where the effect is real but payback swings widely depending on how the cost is structured. The calculation later in this article deliberately assumes that band (four minutes), because a case built on where most plants actually sit survives internal review better than one built on the best-case plant.
Claims that automated notification and early warning can improve MTTR by 30 to 50 percent appear mainly in reports published by solution vendors. The direction is plausible, but it is safest to treat those figures as vendor claims and keep some distance from them. Putting another company’s improvement rate into your capital request without your own baseline leaves you unable to explain the number later.
The Detection Layer Is Harder to Shrink Than to Grow
The common misconception in detection design is that more sensors are better. In practice, narrowing the set of abnormalities you put on the notification path from day one is what makes the system stick.
The first release should carry only abnormalities that require a person to go and fix them. Minor errors that self-recover, and abnormalities the operator will certainly notice unaided, belong in the Record layer but not the Delivery layer. That single split prevents most alert fatigue.
Moving from knowing after the fact to knowing before it happens is the next stage. Vibration and current-waveform early warning belongs to the discussion in implementing a predictive maintenance system, and is far less likely to fail if you build the notification foundation first. Reverse the order and you end up detecting precursors with nowhere to send them.
Comparing Delivery Methods: Andon, Phones, Smartwatches, and Paging Devices

This section is entirely about the Delivery layer. Below are the options people compare when searching for andon-to-smartphone or smartwatch alerting, compared on the axes that actually matter inside a plant.
| Delivery method | Reach | Under high noise | While walking | Hands occupied | ACK possible | Device cost guide | Main weakness |
|---|---|---|---|---|---|---|---|
| Andon (stack light, board) | Line of sight only | Excellent (light) | Poor | Excellent (just look) | No | Low to medium | Useless if nobody is looking |
| Plant-wide PA | Whole building | Fair (loses in noisy buildings) | Good | Good | No | Low | Ambiguous recipient, no record |
| Rugged smartphone (push) | Wireless coverage area | Fair (sound loses) | Excellent | Fair (must take it out) | Excellent | Medium | Hard to operate with gloves or in Ex zones |
| Smartwatch (vibration) | Wireless coverage area | Excellent (vibration) | Excellent | Excellent (glance at wrist) | Good | Medium | Little information, charging routine |
| IP radio / dedicated pager | Dedicated network area | Good | Excellent | Good | Good | Medium to high | A separate network to maintain |
| Anywhere | Poor | Fair | Poor | No | Near zero | No immediacy; unfit for urgent alerts |
The Practical Conclusion
What the comparison implies is simple. Do not try to do it all with one channel. A realistic configuration looks like this.
- Primary alert = smartwatch vibration. It reaches the individual reliably even under high noise, and can be read by glancing at the wrist with gloves on. In a factory environment, vibration is a dramatically stronger delivery channel than sound
- Detail check = rugged smartphone. Which machine, which alarm, and whether the same thing has happened before
- Status sharing = andon or large display. Line state, not individual assignment, shared with everyone walking past
- Escalation on no response = phone call or IP radio. Changing the channel raises the probability of reaching someone
The classic smartwatch deployment failure is trying to do everything on the watch. The screen is too small to read detail, so the operator ends up walking to the machine every time anyway. Designs that treat the watch as a receive-and-acknowledge device only, and leave judgement to a phone or a terminal at the machine, are the ones that stick.
Delivery Quality Is Capped by Your Wireless Infrastructure
This gets overlooked constantly: Delivery layer performance is determined by wireless design, not by the quality of the notification software. When coverage drops, a push notification either arrives late or does not arrive. A plant full of metal is a hostile RF environment, and placing access points with an office mindset guarantees dead zones.
What to verify before buying is not the notification product’s datasheet but how your own in-plant wireless network is designed. That connects directly to the design thinking in building factory wireless LAN and industrial networks. Cases where the symptom is “alerts do not arrive” and the cause turns out to be an RF hole are not rare at all.
Escalation Design and Countermeasures for Alert Fatigue
This is the hardest part of Delivery layer design. As a software configuration task it looks mundane, but most of the return on investment is decided here.
How to Write an Escalation Rule
Escalation timings vary by how critical the machine is. Applying a single rule to every asset means either critical equipment gets a slow response, or people burn out on trivial ones.
| Severity | Scope | Primary alert | On no response | Final escalation |
|---|---|---|---|---|
| P1 (whole line down) | Bottleneck process, shared utilities | Owner’s watch, immediate | Supervisor at 2 min | Production manager plus phone call at 5 min |
| P2 (single machine down) | Individual equipment | Owner’s watch, immediate | Supervisor at 5 min | Maintenance lead at 15 min |
| P3 (minor, monitor only) | Precursors, quality trends | Batched within the shift | No escalation | — |
The critical discipline here is the courage not to notify P3 at all. P3 accumulates in the Record layer and gets used in the end-of-shift review and the weekly improvement meeting. Not pushing it in real time is the single largest countermeasure against alert fatigue.
Make ACK Mandatory
For escalation to function, the recipient needs a way to declare “I am going.” That is the acknowledgement. One button on the watch or one tap on the phone is enough, but without it you cannot evaluate the condition “no response, pass it on.”
ACK has a second benefit. It makes visible who is being called, and how often. If acknowledgements concentrate on one individual, that is a sign of key-person dependency, and it becomes input for staffing and training priorities. In Thailand, where turnover runs high, that visibility translates directly into early detection of handover risk.
Manage Alert Volume as a KPI
The way to prevent alert fatigue is to make alert volume itself a managed metric. The thresholds we set first are as follows — these are our internal design baselines drawn from operational experience, not external statistics.
- Keep actionable alerts to ten or fewer per person per eight-hour shift
- Aggregate repeat alarms from the same machine within a fixed window (for example ten minutes) into a single notification
- For signals that trip and clear repeatedly (flapping), notify only after n occurrences within m minutes
- Review alert count, ACK rate, and count of escalations due to no response monthly, and adjust thresholds
A system without that review cycle will be back to over-notification within six months, because equipment gets added, sensors get added, and only the thresholds stay untouched. Budget for the fact that notification design is a continuous tuning activity, not a one-time configuration task, and reflect it in the operating model you cost out.
Making Minor Stoppages Visible: Connecting Notification Logs to OEE
Now we are in the Record layer — and in the payback discussion, it is in fact the layer with the highest value.
Minor and Major Stoppages Erode Different Parts of OEE
First, terms. Minor stoppages (*choko-tei*) are short stops of tens of seconds to about ten minutes, cleared on the spot by the operator. Major breakdowns (*doka-tei*) are long stoppages of an hour or more, recognised as failures and written up in a report.
Overall equipment effectiveness decomposes into availability multiplied by performance multiplied by quality. Major breakdowns pull down availability; minor stoppages pull down performance. Management usually watches availability, so major breakdowns reach the agenda and minor stoppages do not. Yet in aggregate, minor stoppages are frequently the larger loss.
Why Minor Stoppages Never Reach the Record
The reason is straightforward. For an operator, there is no rational case for writing down a jam that took twenty seconds to clear. Writing it takes longer than fixing it. So it is not recorded. Because it is not recorded, it cannot be aggregated. Because it cannot be aggregated, nobody knows which machine and which failure mode dominates.
The Record layer is what breaks that structure. Nobody has to write anything — the fact that the machine stopped exists as a signal. How many times the alert fired, time to ACK, and time to restart all accumulate automatically. That is what visibility into minor stoppages actually consists of, and it is a part you can never reach by redesigning the shift report form.
The Minimum Fields a Record Layer Must Capture
| Field | How it is captured | What it is used for |
|---|---|---|
| Timestamp, machine, alarm type | Automatic (signal) | Pareto analysis of occurrence frequency |
| Recipient, ACK timestamp, who acknowledged | Automatic (system) | Detect-to-response time, key-person dependency |
| Restart timestamp | Automatic (signal) | Measured downtime, MTTR calculation |
| Cause code | Semi-automatic (picklist on device) | Prioritising countermeasures |
| Response notes | Optional free text | Knowledge capture, training material |
Making the cause code a picklist is a condition for adoption; free text simply does not get filled in. On Thai shop floors, where Thai, Japanese, and English coexist, a design that has people select a code (number or icon) and switches only the displayed label by language is what keeps operations from breaking down.
The logs you collect gain far more value when merged into the wider plant data design rather than kept in a silo. Putting stoppage data on the same footing as utilisation, yield, and energy consumption is the same line of thinking set out in traceability system build costs and design, where the question is likewise how to make shop-floor records usable beyond the system that generated them. Not letting the notification system close in on itself is what decides its asset value five years out.
Cost of an Equipment Alert Notification System: Five Cost Layers and Payback (THB)

Search for the cost of an equipment alert notification system and the ranges are so wide as to be useless. That is because the scope included in a quotation differs completely between vendors. Here we split cost into five layers and make the variables explicit.
The Five Cost Layers (Assuming a 20-Line Plant in Thailand)
The figures below assume a mid-sized Japanese-managed plant in Thailand at roughly 20 lines. They move substantially with machine age, existing wiring, and building structure, so treat them strictly as numbers for getting the order of magnitude right in early study.
| Layer | What it includes | Approximate range (THB) | Main variables |
|---|---|---|---|
| 1. Detection | Contact/PLC tag pickup, added sensors, I/O units | 8,000–35,000 per line | Existing PLC, spare contacts available, machine age |
| 2. Network | Industrial APs, cabling, gateways, power work | 150,000–600,000 per plant | Floor area, metal shielding, reusability of existing wireless |
| 3. Platform | Notification engine, server or cloud, licences | 200,000–900,000 (initial) | On-premises vs cloud, user count, multilingual support |
| 4. Devices | Smartwatches, rugged phones, andon, display boards | 4,000–25,000 per unit | IP rating, Ex requirement, quantity |
| 5. Deployment and support | Requirements, installation, training, annual support | 15–25% of initial cost per year | Number of languages, SLA, remote support feasibility |
Rolling out all 20 lines at once and assuming 60 devices (one per maintenance technician and supervisor across two shifts, plus spares), initial cost typically lands in a range of 1.2 to 4.5 million THB. The rough composition is as follows.
- Low end: 160k + 150k + 200k + 240k (4,000 THB x 60 units) + roughly 450k deployment effort = approximately 1.2 million THB
- High end: 700k + 600k + 900k + 1,500k (25,000 THB x 60 units) + roughly 800k deployment effort = approximately 4.5 million THB
One caution on how to read this: these five layers are cut by where the money is spent, which is a different axis from the three layers used earlier in this article (Detection, Delivery, Record). The costs that correspond to the Record layer — storing notification logs, the aggregation screens, the connection into OEE — sit inside layer 3, the Platform layer. Cutting them is the easiest way to bring the initial figure down, and it is also the fastest way to discard the layer this article argues returns the most value.
Note that the 15 to 25 percent of initial cost in layer 5 is the annual support guide; the one-time deployment costs above (requirements definition, installation, training) are carried separately on top of the initial figures. Deployment effort does not scale linearly with plant size — a fixed quantum is always required — so as a ratio to hardware it runs higher for smaller sites (in the low-end case it works out to roughly 60 percent of hardware cost, in the high-end case roughly 20 percent). It is worth assuming this is not a line item you will negotiate down easily. The reason the spread exceeds three times is that three things dominate: whether existing signals can be reused in the Detection layer, whether wireless must be built from nothing in the Network layer, and which device grade and how many units in the Device layer. That also means investigating those three points alone raises the accuracy of your estimate by a full step.
The Running Costs People Forget
Discussion fixates on initial cost, but over five years layers 3 and 5 come to dominate.
- Cloud subscription (five-year total swings heavily on per-user versus per-device pricing)
- Smartwatch replacement as batteries degrade (in real operation, plan on two to three years)
- Replacing lost and damaged devices (assume a few percent attrition per year on shop-floor hardware)
- Threshold tuning effort (the monthly review from the previous section)
Unless you compare on five-year total cost of ownership, the cheapest initial proposal ends up the expensive one. Per-user cloud pricing in particular inflates faster than expected in shift-based plants with large headcounts.
Our Own Calculation: Annual Effect of Cutting Detect-to-Response From 4 Minutes to 1
To make an investment decision, put the benefit side in the same currency. Everything below is a calculation built on assumed values, not external statistics. Substitute your own numbers.
Assumptions (all assumed)
- Plant scale: 20 lines, two shifts, 250 operating days per year
- Minor stoppage frequency: 8 events per line per day
- Detect-to-response wait: average 4 minutes, reduced to 1 minute (a 3-minute saving per event)
- Opportunity loss per hour of line stoppage: assumed at 1,200 THB
Calculation
| Step | Formula | Result |
|---|---|---|
| Annual minutes saved | 20 lines x 8 events/day x 3 min/event x 250 days | 120,000 minutes/year |
| Converted to hours | 120,000 minutes ÷ 60 | 2,000 hours/year |
| Converted to money | 2,000 hours x 1,200 THB/hour (assumed) | 2,400,000 THB/year |
That produces a figure of roughly 2.4 million THB per year. But dropping that number straight into a capital request is dangerous, for two reasons.
- 1,200 THB per hour is an assumption. Real opportunity loss depends heavily on whether that line is a bottleneck against actual orders. A stoppage on a line with spare capacity is not a loss at all if overtime absorbs it
- There is no guarantee the recovered time converts into revenue. If downstream processes cannot absorb the extra output, less downtime just means more inventory
On the cost of unplanned downtime in manufacturing, several studies exist and their estimates vary widely — from tens of thousands of USD per hour as an industry average in some, to hundreds of thousands of USD per hour in others, with automotive consistently reported as the outlier at the top end (analysis of the cost of unplanned downtime, manufacturing downtime cost outlook). Because the spread between studies is so large, we strongly recommend against importing any single figure as your own assumption.
State Payback as a Range, Not a Number
Set the benefit figure against the cost range. The practical addition here is a column that takes a conservative view of realisation rate — the proportion of the designed time saving you actually achieve.
| Case | Initial cost | Expected annual benefit | Simple payback |
|---|---|---|---|
| Optimistic (existing signals and wireless reusable, 100% realisation) | 1.2M THB | 2.4M THB | approx. 0.5 years |
| Standard (some installation work, 60% realisation) | 2.5M THB | 1.44M THB | approx. 1.7 years |
| Conservative (full installation, multilingual, 40% realisation) | 4.5M THB | 0.96M THB | approx. 4.7 years |
Payback scatters across roughly 0.5 to 5 years. That width is precisely why this article argues for the three-layer model. What separates the optimistic case from the conservative one is not product pricing but how reusable the existing Detection layer is, and what realisation rate you achieve. In other words, separating out how much each layer needs moves payback far more than negotiating a discount.
And what determines realisation rate is the implementation level introduced at the top of this article. Stop at L3 (push notifications arrive) and the benefit is limited to whatever anyone happened to notice sooner. Only when you reach L4 — escalation, ACK, and performance analysis — does the reduction in waiting time become a repeatable number. In the table above, 100 percent realisation assumes L4 is embedded in daily practice, 60 percent assumes L4 was built but adoption is still in progress, and 40 percent assumes the deployment stopped at L3. When comparing quotations, check whether the proposal stops at L3 or includes L4 before you compare the price.
Additional Considerations for Plants in Thailand and ASEAN
Import a Japanese domestic reference design as-is and Thailand will always add requirements. Here are the ones worth building in at design stage.
Solve Multilingual Requirements by Coding, Not Translating
Thai plants commonly run with Thai, Japanese, and English, and often Myanmar or Khmer speakers on the floor as well. Preparing notification text in every language collapses the moment alarms start being added, because each addition triggers a translation task.
The workable design is to define alarms as a code plus an icon, and switch display language by user setting. The notification body consists of machine name, alarm code, and colour; only the detailed description lives in a language table. With that structure, adding an alarm means adding one row to the language table.
Turnover and Key-Person Dependency
Thailand’s working-age population has entered decline, and competition for talent with newly arriving investors, EV-related manufacturers among them, is reported to be pushing up salaries for manufacturing technicians (commentary on Thai labour trends, analysis of Thailand’s population and labour structure — both Japanese-language references).
This bears directly on notification design. An operating model that assumes a veteran technician is always on site watching the machines is not sustainable where labour mobility is high. The moment that veteran leaves, the tacit knowledge of which sound means which fault disappears with them. Accumulating alarm types and response notes in the Record layer is insurance against training cost and handover risk.
Mining the ACK history for alarms that only one person has ever handled tells you in advance which machine will stop at the next resignation. That is less a side benefit of a notification system than one of the primary reasons to have a Record layer at all.
Noise and Environmental Conditions
Thai plants frequently run in non-air-conditioned or open-structure buildings, with rainy-season downpours added to the ambient noise. In high-noise buildings, PA announcements and buzzers are severely disadvantaged as a delivery channel. The advantage of vibration devices (smartwatches) shows up more strongly here than in an air-conditioned Japanese plant.
Beyond that, design needs to cover device ingress protection ratings suited to heat, humidity and dust; lightning surge protection (rainy-season lightning countermeasures are a commonly expected item locally, and worth confirming specifically for APs and gateways with outdoor cabling); and maintaining the notification path during power outages (UPS on the server is not enough — the access points need it too).
The BOI Angle
The Thailand Board of Investment maintains ongoing incentives for automation and digital-related investment (BOI investment promotion policy for automation and robotics industries). Scope and conditions are revised periodically, so please work on the basis that you should never decide specific rates or eligibility from an article like this one, and must confirm the current conditions at the time of application.
The practical point is not eligibility but timing. Cases where a company realises after committing the investment that it would have qualified do occur. A notification system on its own may be too small to matter, but folded into a wider automation investment package it can come into scope. Checking once with your internal finance function or a BOI consultant at the early stage of the capital request costs very little.
Deployment Steps and How to Run a PoC
Here is a realistic sequence aligned to the three-layer model. Not doing every line at once is the single largest risk reduction available.
Step 1: Measure Your Current MTTR (1–2 Weeks)
As above, what you measure is the time from abnormality to a person arriving at the machine. No special mechanism is needed. A stopwatch and a paper log across 20 to 30 events are enough to see the distribution. Record who responded and which alarm it was at the same time.
That measurement becomes the baseline for verifying results later. Deploy without it and you will never have grounds for claiming an improvement. Projects that begrudge these two weeks always stall at the second round of capital approval.
Step 2: Run a Narrow PoC (4–6 Weeks)
Limit it to one or two lines, five or fewer alarm types on the notification path, and ten or fewer participants. What the PoC needs to verify is not product features but these four things.
| Verification item | How to measure | Judgement guide |
|---|---|---|
| Delivery rate | Device receipts against messages sent | Below 99% and the wireless design is the suspect |
| ACK rate | Acknowledgements against notifications | Below 80% means an operating rule or volume problem |
| Median detect-to-response | System logs | Is it clearly shorter than the baseline? |
| False alarm rate | “Went, found nothing” against total notifications | Above 10% and trust starts to break down |
Roll out to production with a high false alarm rate and you go straight into Failure 2 (alert fatigue). If the false alarm rate will not come down during the PoC, the cause is almost always Detection layer thresholds. Changing the Delivery layer product will not fix it.
Step 3: Fix the Go/No-Go Criteria in Advance
Document what constitutes grounds to proceed before the PoC starts. Set criteria afterwards and, having already spent money, the interpretation drifts towards success. A workable set: delivery rate at or above 99 percent, ACK rate at or above 80 percent, median detect-to-response reduced by 40 percent or more against baseline, and false alarm rate at or below 10 percent.
Step 4: Phased Rollout and Standing Up the Record Layer (3–6 Months)
Start the rollout at the bottleneck process. The financial effect is largest there and the internal case is strongest. In parallel, begin the Record layer’s monthly review from the very first line. The Record layer is not something you build after rollout finishes. Unless it starts running from line one, you will reach every line without ever forming the review habit.
Frequently Asked Questions
What is the difference between an equipment alert notification system and an andon system?
An andon system is one delivery method that conveys an abnormality through light or a display to the people physically present. An equipment alert notification system refers to the whole mechanism spanning three layers — detection, delivery, and record. Andon sits inside the Delivery layer and carries a physical constraint: it cannot reach anyone outside line of sight. The two are not alternatives. The usual configuration keeps andon in place and adds notification to personal devices on top.
How much does an equipment alert notification system cost?
For a 20-line plant in Thailand, the cost section above sets out that initial cost typically falls in a range of roughly 1.2 to 4.5 million THB. The two factors driving that spread are whether signals can be reused from existing PLCs (Detection layer) and whether in-plant wireless must be built from scratch (Network layer). On top of that, annual support runs at roughly 15 to 25 percent of initial cost, and cloud-based options add a subscription. Compare on five-year total cost of ownership.
Smartwatch or rugged smartphone — which is better suited?
They serve different roles, so running both is more realistic than choosing one. In summary:
| Aspect | Smartwatch | Rugged smartphone |
|---|---|---|
| Strength | Receiving under noise, hands occupied | Detail check, history lookup, cause entry |
| Information volume | Low (machine name plus type) | High |
| Wearability | Worn continuously | Must be taken out of a pocket |
| Suitable role | Receiving the primary alert and acknowledging | Secondary confirmation and record entry |
In the high-noise plants typical of Thailand, the value of a vibration-based primary alert is especially high.
Can a notification system be deployed on old existing equipment?
Yes. Even for machines with no PLC, the practical order of options is as follows.
- Borrow the existing contact signal. Branch off the wiring to the stack light or buzzer. The cheapest route by a wide margin
- Add external sensors. A current transformer to read motor on/off, or a photoelectric sensor to watch work flow
- Read the machine’s status lamp with a light sensor. Non-contact state capture even on equipment that cannot be modified
In our experience, option 1 alone covers a large share of target equipment in most plants, and starting by inspecting the existing stack light and buzzer wiring is the standard opening move. Stopping the evaluation on the assumption that the equipment is too old is the most common pattern we see.
How do you stop alerts from becoming so frequent that people ignore them?
There are four countermeasures, listed in descending order of impact.
- Narrow what goes on the notification path. Only abnormalities that require a person to go and fix them. Minor self-recovering errors go into the Record layer only
- Split routes by severity. P3 (minor, monitor only) is never pushed in real time
- Aggregate and suppress flapping. Repeat alarms from the same machine collapse into one within a fixed window
- Review alert volume as a monthly KPI. Look at counts, ACK rate, and no-response counts, and adjust thresholds
Rebuilding a floor that has already learned to ignore alerts costs far more than designing it correctly from the start. Be most restrictive in the early phase of deployment.
Summary
Investment decisions on an equipment alert notification system always blur if you start from product comparison. The first move is to measure how much of your own MTTR sits in detect-to-response. Short, and the priority is low. Long, and the effect is substantial.
Split the design into three layers. In the Detection layer, narrow what counts as an abnormality and put only human-required events on the notification path. In the Delivery layer, do not depend on one channel: vibration (watch) for the primary alert, phone for detail, andon for shared status, and a different channel for escalation on no response. Delivery quality here is set by in-plant wireless design, not by notification software. In the Record layer, automatically accumulate the minor stoppages that previously left no trace, and connect them to performance rate improvement and OEE analysis.
View cost across five layers — detection, network, platform, devices, deployment and support. For a 20-line plant in Thailand that lands at 1.2 to 4.5 million THB initially, with the variance concentrated in whether existing signals are reusable and whether wireless work is required. In the calculation above (all assumed values), cutting detect-to-response from four minutes to one produces an annual effect of roughly 2.4 million THB, and simple payback scatters across roughly 0.5 to 5 years depending on realisation rate. The key to narrowing that range is not discounting but separating out how much each layer actually needs.
Then build the Thailand-specific considerations in at design stage — solve multilingual requirements by coding, treat the Record layer as insurance against high turnover, recognise that vibration devices have an outsized advantage in noisy environments, and confirm BOI incentives before committing the investment. Every one of them is more expensive to add later.
If you have measured your current MTTR but cannot judge which layer to start from, or you want to know whether keeping your existing andon and adding notification to personal devices is realistic in your plant, we are happy to talk at that stage. TOMAS TECH is based in Bangkok and supports Japanese manufacturers across Thailand and ASEAN with equipment monitoring, IoT traceability, and factory network design. Even just the exercise of separating out which of the detection, delivery, and record layers is your bottleneck tends to surface the real issues faster with an outside perspective. Reach us through the contact form.
References
- ReliaMag, “Cost of Unplanned Downtime in Manufacturing”
- IDS Data, “Manufacturing Downtime Costs and Forecasting 2026”
- BOI, “Investment Promotion Policy for Automation and Robotics Industries”
- Skillnote, “What is choko-tei” (Japanese)
- Keyence, “Short stops (choko-tei)” (Japanese)
- evort, andon explainer (Japanese)
- Kuno, “Latest trends in Thai labour management” (Japanese)
- IDE-JETRO IDE Square, on Thailand’s population and labour force (Japanese)