Blog

2026.08.06

Minor Stoppage Countermeasures 2026 – Where OEE Actually Hides Them

Minor Stoppage Countermeasures 2026 - Where OEE Actually Hides Them

Stand in front of a machine for an hour and the line will stop again and again. An operator clears a jammed part by hand, presses reset, and twenty seconds later the machine is running as if nothing happened. Nobody writes it on the shift report. Yet at the end of the month the utilisation figure comes back at 95 percent, and it bears no relationship to what anyone on the floor actually experienced. Minor stoppage countermeasures miss the target so often because what the floor calls “a stop” and what the OEE (Overall Equipment Effectiveness) formula counts as downtime are not the same event at all. This article traces where that gap is created, and which metric you have to look at to catch it.

Why minor stoppage countermeasures miss the target

When an improvement programme for minor stoppages (also called micro-stops or chokotei) begins, it usually starts with a slogan about raising equipment utilisation. A target percentage is set, production engineering is made the owner, and the deadline is the end of the half. Six months later the metric has barely moved and the floor reports that nothing feels different. That ending is not a failure of the people running the programme. It is a structural mismatch between the thing being measured and the thing you wanted to reduce.

What the floor calls “stopped” is not what OEE calls downtime

Ask an operator how many times the machine stopped today and you will get an honest answer. Ten times. Maybe fifteen. But that definition of “stopped” belongs to the person standing in front of the machine. A part caught and they reached in. A sensor misfired and the tower light turned amber. The conveyor ran empty waiting for material. Every one of those is a stop to the person who dealt with it.

What a production monitoring system or MES records as downtime, on the other hand, is only the subset of events that satisfy a condition somebody defined in advance. In most installations that condition is “the running signal from the machine has been off continuously for longer than a threshold.” If the threshold is set at five minutes, a jam that cleared in twenty seconds does not exist in the record at all. The operator says fifteen stops. The system says two. That difference is the definition of a minor stoppage.

The important part is that the unrecorded time does not disappear. On paper the machine was running, but no product came out during those minutes. So the loss is pushed into the “running but not producing enough” bucket, which is Performance, not Availability. The Availability column stays clean and only the Performance column sits stubbornly low for no visible reason. When someone says “our Availability is high but OEE will not improve,” this displacement is frequently what is happening.

Who decides the threshold that separates minor stoppages from major breakdowns

The boundary between a minor stoppage and a major breakdown is not a physical phenomenon. It is an operating decision that somebody has to make. TeepTrak’s 2026 OEE benchmark report cites under 5 minutes as the common threshold for a micro-stop. The same report finds that, with variation by industry, stops shorter than five minutes can account for 18 to 38 percent of total losses. At a plant near the upper end of that range, close to 40 percent of all loss consists of stops that most shop floors never record.

The problem is that in many plants nobody ever made that five minute decision explicitly. The value is whatever happened to be in the standard PLC function block, or the machine builder’s default, or a number a predecessor typed into a settings screen years ago. Choosing a threshold is really a decision about which size of loss enters management’s field of view, and that is not something production engineering should settle alone. In practice it usually gets settled quietly, inside a configuration dialog.

Lower the threshold and you capture more stops, but you buy a different problem. Drop it to ten seconds and cycle waits, pauses during changeover, and momentary sensor chatter all register as downtime events. You end up with thousands of events a day and the reason code workflow collapses under its own weight. Threshold design is the act of balancing the granularity of the phenomena you want to see against the number of entries the floor can realistically keep making. Install sensors before settling this and you build a dashboard that produces data nobody looks at.

Why OEE from handwritten logs comes out 8 to 15 points too high

Plenty of plants calculate a baseline OEE from their existing paper shift reports, get 75 percent, and conclude they are in decent shape. The same TeepTrak report notes that OEE derived from handwritten records tends to come out 8 to 15 points higher than the real figure. That gap is not falsification or laziness. It arises naturally from the nature of recording by hand.

First, people do not record problems they solved themselves. They kicked the jam loose, hit reset, nudged the rotary table by hand. Anything that recovered in tens of seconds is not perceived as an incident, so it never reaches the form. Second, records are written retrospectively. Filled in from memory before the end of shift, only the memorable large stops survive and the small ones get rounded away. Third, the unit of recording is coarse. If the form is laid out in five or ten minute increments, anything shorter is structurally impossible to write down.

A second distortion acts on the reason column. Even with ten options available, the floor gravitates to the one or two that are easiest to select. Entries pile up under “material defect” and “other,” and when you aggregate them the top of the Pareto chart is “other.” This is not the floor cutting corners. It is the safest choice when someone is busy and uncertain how to classify what just happened. Build a countermeasure plan on handwritten data and you are building on top of both distortions at once, small stops falling out and reasons being flattened.

Locate yourself on the four quadrants of downtime

To work out where your own minor stoppages sit, it helps to cut downtime along two axes. One is whether the system records it or not. The other is whether a person has to intervene to recover, or the machine recovers by itself.

Minor Stoppage Countermeasures 2026 - Where OEE Actually Hides Them - figure 1

Crossing those two axes gives four quadrants, and it becomes obvious that each one calls for a different response.

Recorded or not, by recovery modeA person intervenes to recoverThe machine recovers automatically
RecordedBreakdowns, changeovers, material outages. Already visible as Availability lossAutomatic retries that exceeded the threshold. Present in machine logs but the reason field is usually blank
Not recordedClearing a jam in tens of seconds, manual resets, small adjustments. Absent from both the shift report and the machine logSensor chatter, recovery after a retry, waits inside the cycle. On paper the machine is running

The top left quadrant is already visible. If a breakdown stops the line for thirty minutes the line leader knows about it and the maintenance record exists. This is where the improvement cycle is already turning, so the marginal return on additional investment here is relatively small. The problem lives in the bottom two quadrants, and especially in the bottom right.

The main battlefield for minor stoppages is the “not recorded and automatic recovery” quadrant. Nobody is called, so no one remembers it. The machine recovers on its own, so it usually leaves no alarm history either. And yet because the frequency is high, the cumulative time is anything but negligible. The loss from this quadrant is booked inside Performance rather than Availability. Which means that no amount of shouting about raising Availability reaches this quadrant at all. The instruction is addressed to the wrong metric.

The bottom left, “not recorded and a person intervenes,” deserves attention too. This is the territory absorbed by operator skill, and because the better operators fix things faster, the machines run by the better operators look problem-free in the data. A plant where output drops the moment staff change is very likely leaning on this quadrant. In Thailand and across ASEAN, where a degree of workforce turnover is normal, making this quadrant visible is worth more than it would be at a plant in Japan.

You can get a rough read on which quadrant your losses sit in from data you already have. Put the count of downtime events in the machine log next to the number of stops operators describe when you interview them. The larger the difference between those two numbers, the more loss has sunk into the lower two quadrants.

The three OEE factors and where minor stoppages live, aligned with ISO 22400-2

Most internal arguments about minor stoppages fail to converge because people are using different definitions of the same metric. Take a single word like “utilisation.” One person in the meeting means the share of time the machine had power on. Another means the share of planned time spent producing good parts. Start discussing numbers before aligning definitions and the discussion will stall halfway through, every time.

Minor Stoppage Countermeasures 2026 - Where OEE Actually Hides Them - figure 2

For manufacturing KPIs there is an international standard, ISO 22400-2, which defines 34 KPIs including Availability, Performance, Quality and OEE. Adopting that standard as the internal reference point lets the Japanese head office, the Thai plant and the machine builder use the same vocabulary. It matters most when you want to compare sites. If each location uses its own homegrown definition of utilisation, the comparison is not merely inaccurate, it is meaningless.

Which of Availability, Performance and Quality absorbs the loss

OEE is the product of three rates. The kinds of loss each one absorbs can be laid out as follows.

FactorWhat it measuresLosses it mainly absorbsRelationship to minor stoppages
AvailabilityShare of planned production time the machine was actually runningBreakdowns, changeover, waiting for material, stops longer than the thresholdOnly stops that exceed the threshold land here
PerformanceActual output against what ideal cycle time says should have been produced in the running timeSpeed loss, idling, short stops below the thresholdThe bulk of minor stoppages dissolve here
QualityShare of good parts in total parts producedDefects, rework, start-up lossIndirect, when minor stoppages induce defects

The second row of that table is the core of this article. A stop below the threshold means, by definition, that the machine was running. But no product came out during that time, so actual output falls short of what ideal cycle time predicts. Performance therefore drops. Not a trace of it remains in the Availability column.

So when a programme is labelled “minor stoppage countermeasures” and given an Availability target, the activity automatically drifts toward breakdown prevention and changeover improvement. Neither is a bad thing to do, but neither touches the minor stoppages you set out to reduce. Six months later, “utilisation is up but nothing feels different” is the predictable end of that path.

Performance will not move until you fix the ideal cycle time

To calculate Performance you need an ideal cycle time in the denominator, the theoretical fastest time per piece. It is remarkable how many plants have never fixed this value. The machine builder’s catalogue figure, the trial figure from commissioning, the fastest run of the last three months, the standard time used in cost accounting. There are several candidates, and the Performance figure moves substantially depending on which one you pick.

Take the catalogue figure and Performance reads low, because the real material and fixturing conditions are not in that number. Take the recent fastest run and your baseline is a value that happened on a day when conditions were good, producing a target that is chronically unreachable. Take the cost accounting standard and the allowance baked into it makes the baseline generous enough that Performance can read close to 100 percent even while minor stoppages are occurring.

In practice the most workable approach is to measure actual cycle times per part number during a stable stretch where good parts are coming out continuously, and set the initial value near the median of that distribution. What matters is not the absolute correctness of the number but that the reasoning behind it is documented, and that it only changes by agreement. A metric whose baseline shifts cannot be compared over time, which makes it useless for measuring whether an improvement worked.

There is one more constraint. Ideal cycle time differs by part number. Set a single cycle time on a line that runs a mixed model schedule and Performance will move whenever the product mix changes, making it impossible to distinguish from a change in minor stoppages. If you want minor stoppage visibility on a line with frequent changeover, linking the part master to production data collection is a prerequisite rather than a nice-to-have. That is an information systems task more than an equipment installation task, and it follows the same reasoning as estimating factory IoT cost across five layers.

An OEE calculation you can follow by hand

To make the definitions concrete, here is a worked example. The numbers below are illustrative and do not describe any particular installation. They exist so you can follow the arithmetic yourself.

Take one shift on a machine with 480 minutes of planned production time. Of the stops during that shift, the recorded ones, which is to say breakdowns, changeover and waiting for material, totalled 60 minutes. Total parts produced were 630, of which 54 were defective, leaving 576 good parts. Ideal cycle time is 30 seconds per piece.

Availability is running time of 480 minus 60, or 420 minutes, divided by planned production time, so 420 divided by 480 equals 87.5 percent. Performance is the time ideal cycle time says 630 parts should take, 630 times 30 seconds equals 18,900 seconds or 315 minutes, divided by the 420 minutes the machine was actually running, so 315 divided by 420 equals 75.0 percent. Quality is 576 divided by 630, or 91.4 percent. OEE is the product of the three, 0.875 times 0.750 times 0.914, which is 60.0 percent.

That 60.0 percent sits exactly at the median OEE of 60 percent reported in TeepTrak’s 2026 benchmark across 450 plants. In the same benchmark the top quartile reaches 75 percent and what is called world class is 85 percent. So this example machine is squarely in the middle of the pack.

Now the interesting part. Suppose that during this same shift there were 45 minutes in total of stops shorter than five minutes. Because the threshold is set at five minutes, those 45 minutes were never recorded as downtime and the machine is treated as having been running. What happens to the numbers if you can record those 45 minutes as stops?

Running time becomes 420 minus 45, or 375 minutes, so Availability falls to 375 divided by 480, which is 78.1 percent. At the same time the denominator of Performance becomes 375 minutes, so Performance rises to 315 divided by 375, or 84.0 percent. Quality is unchanged at 91.4 percent. The product is 0.781 times 0.840 times 0.914, which is 60.0 percent. The OEE value itself has not moved at all.

That is precisely what making minor stoppages visible does. The headline OEE stays put, but the loss relocates into Availability. Only when Availability drops from 87.5 percent to 78.1 percent does the fact that 9.4 points worth of downtime is sitting right there appear as a number. Put the other way round, if you have no visibility system and you are only watching OEE, minor stoppages are undetectable in principle. A high Availability figure is not evidence that you have few minor stoppages.

The five layers of a minor stoppage visibility system

So how do you actually build it? Laying out the production data collection architecture as five layers, from the bottom up, makes it easier to spot the gaps in a design. For each layer, here is what has to be decided and what happens if it is not.

LayerWhat the layer doesWhat must be decided hereWhat happens if it is not
Layer 1 signal acquisitionExtract run, stop and part count from the machine as electrical signalsWhich points, how to tap them (dry contact, signal tower, PLC comms), what one count pulse meansPart counts disagree with reality and nobody trusts Performance
Layer 2 stop detectionCut discrete downtime events out of the signal time seriesThreshold for calling it a stop, chatter filtering, exclusion of planned stopsEither thousands of events a day, or almost none
Layer 3 reason entryAttach a cause to each downtime eventReason code structure, who enters it, how long they have to enter it“Other” takes the majority and analysis becomes impossible
Layer 4 aggregation and analysisProduce OEE and Pareto chartsAggregation unit (machine, part number, shift), ideal cycle time, rules for revising baselinesNumbers change between meetings and cannot be compared
Layer 5 improvement routineTurn numbers into countermeasuresWho looks at what and when, who decides the countermeasure, how effect is measuredThe dashboard runs but nobody opens it

These five layers look as though difficulty increases from the bottom up. In reality Layers 3 and 5 are the hard ones. Layers 1 and 2 are engineering problems, and once the requirements are settled the work is straightforward. From Layer 3 onward it becomes a question of operating routine and organisation.

The first decision in Layer 1 is where to take the signal from. If the PLC has an Ethernet port and the machine builder will give you an address map, communications give you by far the richest data. Older machines will not offer that. In that case you fall back on non-invasive methods, reading the voltage on the signal tower lamps, which are also the common display device in an andon system, or tapping the existing counter output with a contact in parallel. The signal tower has a real weakness, in that it only lights for conditions the machine itself judged to be abnormal, but the installation work is small and you can start without taking existing machines out of service. That trade-off is worked through in more detail in the article on how to retrofit existing equipment for IoT.

Part counting has a pitfall of its own. Unless you confirm machine by machine whether one pulse means one piece or one shot of a multi-cavity tool, Performance will be wrong across the board. Machines that count trial shots or dry cycles need those excluded. Start up without resolving this and at the first monthly review somebody will say “these numbers do not match the physical count,” and after that nobody will open the system again. Trust in data is decided on the first attempt.

Layer 2 stop detection needs chatter filtering in addition to the threshold. Signal tower and sensor signals can flip momentarily due to mechanical vibration or wiring condition. Without a filter that ignores transitions shorter than a few hundred milliseconds you will generate large volumes of downtime that never happened. Make the filter too aggressive, however, and genuine short stops vanish with the noise, so schedule a period of a few days on the floor to look at raw logs and tune both parameters.

Layer 3, reason entry, is where the design philosophy of the whole system shows most clearly. Finer reason codes give finer analysis but raise the entry burden on the floor. What sustains itself in practice is a first tier limited to five to seven codes, with a second tier only on the machines that need it. Whether the operator or the line leader enters the reason, and whether it is entered on the spot or batched during a break, also has to be decided in advance. Skip that and you end up with whoever can enter something entering it whenever they can, which means the missing data clusters on particular shifts and particular machines. Data that is missing unevenly is worse than data that is missing entirely.

For Layers 4 and 5, deciding “who looks at this, when, and what do they decide” is more effective than producing more metrics. Without a routine such as reading the top three Pareto items from yesterday at the morning meeting, or reviewing countermeasure progress weekly, the investment in Layers 1 through 3 never pays back.

When killing the top Pareto item does not help

Once the visibility system is live, a Pareto chart appears in the first month. Design a countermeasure against the top item, execute it, check the effect next month. That sequence is correct, and yet the expected effect often fails to materialise. Most of the time the cause is how the chart was sorted.

Minor Stoppage Countermeasures 2026 - Where OEE Actually Hides Them - figure 3

Rank by frequency or by cumulative time

There are at least two axes for ranking downtime reasons. Occurrence count, that is frequency, and cumulative stopped time. The two orderings frequently disagree.

Compare reason A, a jam lasting 20 seconds that happens 600 times a month, with reason B, an adjustment lasting 40 minutes that happens 4 times a month. By count, A is 600 and B is 4, so A dominates. By cumulative time, A is 200 minutes and B is 160 minutes, and the gap has narrowed considerably. A modest reduction in the count would flip the ranking outright. A Pareto sorted by frequency and a Pareto sorted by time are two different pictures.

Which one to look at depends on the nature of the countermeasure. If you are considering a permanent fix such as changing the machine mechanism or rebuilding a fixture, cumulative time makes it easier to judge return on investment. If you are thinking about operator workload or quality risk, frequency is the one that matters. A 20 second jam occurring 600 times a month means 27 times a day, which means the operator is reaching into the machine more than a dozen times per shift. That frequency carries risk in the act of intervening itself, injury, parts knocked out of position, loss of repeatability, plus the time not spent watching the other machines. Watch only cumulative time and none of that workload enters your field of view.

The safe practice is to produce both rankings and start with the reasons that appear near the top of both. A reason that ranks high on only one of them should be left alone until you can explain why it ranks high on only one. A ranking you cannot explain usually points at a problem in how the data is being collected.

The other trap is skew in the reason codes themselves. As noted earlier, the floor gravitates to whichever code is easiest to enter. A Pareto chart with “other” or “material defect” at the top is describing the distribution of data entry behaviour, not the distribution of the phenomena. Trying to attack the top item in that state gets you nowhere because the target cannot be identified. If “other” is at the top of your Pareto, fixing the reason code design and the entry routine comes first.

Changeover reduction or minor stoppage countermeasures first

How to split limited resources between the two is a common question, and the evidence you need is in your own numbers. Break down the Availability loss, then compare the time occupied by changeover against the time that would migrate into Availability once minor stoppages are made visible.

As a general tendency, lines that switch product several times a day carry a heavy changeover share, while lines running few products for long stretches carry a heavy minor stoppage share. But there is a sequencing problem in that judgement. Without minor stoppage visibility you have nothing to compare against in the first place. Conclude from a high Availability figure that “changeover is our problem” and, as described above, you will overlook the time sunk in Performance.

The two also differ in how their results appear. Changeover improvement delivers a large saving per event and the result is easy to see. Minor stoppage countermeasures deliver a small saving per event and only add up through volume, so confirming the effect takes longer. Where short-term results are required, starting with changeover is a realistic choice, but even then it is better to begin collecting minor stoppage data in parallel, because data takes time to accumulate. And if you want to judge the two against the whole flow, including the waiting that accumulates between processes and the inventory it creates, that wider lead time view is worth building in parallel rather than treating as a later phase.

Cost and payback for a Thai plant with 20 machines

On to money. The example below assumes a plant in Thailand with 20 target machines, two shifts, and 22 operating days per month, with initial cost estimated along the five layer structure. Real quotations vary with machine type, cable runs and the state of the existing network, so read the following as a way to get a feel for the order of magnitude.

ItemAmount (THB)
Layer 1 signal acquisition units 8,000 x 20 machines160,000
Layer 2 IoT gateways 25,000 x 4 units100,000
Layer 3 shop floor entry terminals 12,000 x 5 units60,000
Layer 4 server and visualisation software, initial350,000
Layer 5 design, wiring and training430,000
Total initial cost1,100,000

What stands out in this breakdown is that design, wiring and training at THB 430,000 is larger than the hardware itself. Designing where to take signals from, panel work, and training the floor all scale with machine count, and the labour is hard to predict because it depends on the condition of the existing equipment. Trying to compress that line item by doing it in house can delay go-live by several months.

On top of the initial cost, budget annual running cost of THB 240,000, made up of maintenance at 180,000, communications at 36,000 and spare parts at 24,000. On a monthly basis that is 240,000 divided by 12, or THB 20,000.

Now the benefit side, where it needs stating plainly that assumptions are being introduced. The calculation below is a framework, an invitation to substitute your own numbers in the same order, and not a guarantee of any particular amount. Replace the assumptions in bold with your own values and run the same arithmetic.

Assumption 1. Operating hours across the target machines, the denominator corresponding to planned production time in OEE, are 20 machines times 16 hours per day times 22 days per month, giving 7,040 machine hours per month.

Assumption 2. Gross profit contribution per machine hour is provisionally set at 450 THB. This is the variable that differs most between plants, moving with unit price, material cost ratio, and whether the machine is a bottleneck. You can approximate your own value by dividing the monthly gross profit of the product families passing through the target machines by those machines’ monthly operating machine hours. If the process is not a bottleneck this value comes out smaller and the investment case becomes correspondingly conservative.

Assumption 3. Visibility plus an improvement routine raises OEE by 3 points. The equivalent additional running time is 7,040 times 0.03, or 211.2 machine hours per month. In money that is 211.2 times 450, or 95,040 THB per month. Subtract the monthly cost of 20,000 THB and the net gain is 95,040 minus 20,000, or 75,040 THB per month. Divide the initial cost of 1,100,000 THB by that and payback is 1,100,000 divided by 75,040, roughly 15 months.

The catch is that 3 points is not a given. As a contrast case, run the same arithmetic for an improvement that stalls at 1 point. 7,040 times 0.01 is 70.4 machine hours per month, 70.4 times 450 is 31,680 THB per month, and the net gain is 31,680 minus 20,000, or 11,680 THB per month. Payback becomes 1,100,000 divided by 11,680, roughly 94 months, or about eight years. Given equipment replacement cycles and software obsolescence, that is a level at which you should conclude the investment does not pay back at all.

Same initial cost, same number of sensors, same software, and payback splits between 15 months and eight years. The variable driving that split is not the hardware specification. It is the single question of how many points you can move with the data once you have it. Which means that debating whether to proceed on the basis of how many sensors to install puts the decision criterion in the wrong place. The questions to ask are whether the floor can keep entering reason codes, whether there is someone who reviews the top of the Pareto every week and decides a countermeasure, and whether the effect can be measured against a consistent baseline. In other words, whether Layers 3 and 5 will turn.

If you are not confident about the Layer 5 routine, the alternative is not “do not invest” but “start smaller.” Narrow the scope from 20 machines to a handful, initial cost falls, and you can widen the deployment once you have confirmed the routine works. For an investment this sensitive to payback assumptions, verifying that the routine is repeatable is more rational than getting the scale right first.

What is different in Thailand and ASEAN

Everything so far is structural and applies anywhere. Plants in Thailand and across ASEAN carry a few additional considerations on top.

First, multilingual reason codes. Showing only Japanese on the shop floor entry terminal will not sustain entry. You need Thai, and depending on the workforce, Burmese or Khmer as well. What matters here is not translation quality but the separation of code from label. Unless aggregation runs on the code and only the display switches by language, every language you add later fragments your historical aggregation. Adding an icon or a photograph to each label also allows entry that does not depend on reading and writing.

Second, workforce turnover and the entry routine. The “not recorded and a person intervenes” quadrant described earlier is territory absorbed by experienced operator skill. When people change, it surfaces. Put the other way round, making that quadrant visible turns training into something concrete. Once you can show in data which part numbers jam most and which machines get hands put into them most often, handover content shifts from anecdote to procedure.

Third, external support for the investment. Thailand’s Board of Investment supports factory smartening investment under its Smart and Sustainable Industry framework, and has published that the first half of 2026 saw 132 applications worth THB 17.2 billion. Whether a monitoring and data collection project qualifies is decided case by case, but where the amounts are significant it is worth checking in advance. Scheme requirements change, so go directly to the BOI’s published material for current conditions.

Fourth, the business environment. Thailand’s manufacturing production index fell 3.1 percent year on year in June 2026, with automobile production down 7.55 percent in the same month, as reported by Business Recorder. When volumes are not growing, investment in recovering what existing machines are losing tends to get more consideration than investment in adding machines. At the same time, measuring OEE while utilisation is depressed makes the figure highly sensitive to how planned production time is defined. Decide in advance whether non-production caused by demand is excluded as planned downtime or included, or month to month comparison becomes impossible.

Fifth, the comparison against labour cost. According to Thai Law Online, the 2026 Thai minimum wage ranges from 337 to 400 THB per day depending on the province. Against that level, the idea of absorbing the data entry burden with people looks superficially viable. But the purpose of recording is not to reduce labour hours, it is to obtain data that supports decisions on a continuing basis. The earlier finding that handwritten records make OEE look 8 to 15 points better carries weight here. Adding people does not readily close the structural gap in which small stops never reach the record at all.

Sixth, time difference and network design when a Japanese head office wants to see overseas plant status. Thailand is two hours behind Japan, so the start of the Thai morning shift lands around 10am Japan time for anyone watching from the head office. Demanding real time visibility makes the network and server architecture heavier, so it is often more realistic to design for the head office consuming daily aggregates while real time stays a local decision support tool. Connecting the plant network directly to head office systems brings security requirements and questions of operational ownership with it, so involve the information systems function early in requirements definition. If you intend to extend condition monitoring toward more advanced maintenance, the considerations overlap with those covered in the article on deciding whether to adopt a predictive maintenance system.

What to do in the first 90 days

Structure the 90 days on the assumption that you do not deploy across all lines at once. The objective is not to produce a result, it is to establish whether Layers 3 and 5 can turn in your organisation.

Spend the first 30 days aligning definitions. Narrow the scope to three to five machines and document the definitions of Availability, Performance and Quality against ISO 22400-2. Set a provisional downtime threshold, five minutes is a common starting point, and build the list of what is excluded as planned downtime. In parallel, measure actual cycle times for the target part numbers and record the ideal cycle time along with the reasoning behind it. Not a single sensor needs to be installed during these 30 days.

Use the next 30 days to stand up Layers 1 and 2. Fix the signal acquisition method, install on the target machines, and tune chatter filtering and the threshold while watching several days of raw logs. The one thing to verify without fail is that the system’s part count agrees with the counter on the floor. Move forward while they disagree and every number downstream loses its credibility. At the same time, put the number of stops operators describe in interviews next to the number of events the system detected. That difference is your clue as to which of the four quadrants your losses have sunk into.

Devote the final 30 days to Layers 3 and 5. Run reason codes provisionally with five to seven options and watch the entry rate and the share of “other” every day. If the entry rate falls on particular shifts or at particular times, adjust the number of codes or the timing of entry. Once a week, produce the Pareto both by frequency and by cumulative time, pick one reason that appears near the top of both, and decide a countermeasure. If even one countermeasure has gone through the loop in those 90 days and its effect could be measured against a consistent baseline, you can judge that you are ready to expand. If it did not turn, adding machines before resolving why will simply reproduce the same outcome at greater scale. Fixing the routine before adding equipment makes the rework far cheaper.

Frequently asked questions

How many minutes counts as a minor stoppage?

There is no definitive international definition. TeepTrak’s 2026 report cites under 5 minutes as the common threshold for a micro-stop, and in practice many plants take that as a starting point. As discussed above, however, the threshold is an operating decision about which size of loss you want to make visible, and the appropriate value shifts with the machine’s cycle time. Apply a five minute threshold to a machine with a cycle time of a few seconds and dozens of cycles’ worth of loss hides underneath a single threshold.

How much do minor stoppage countermeasures cost?

The worked example above shows an initial cost of THB 1,100,000 and an annual cost of THB 240,000 for 20 machines on two shifts. That is one configuration, and the figure moves substantially with how easily signals can be taken from existing machines, cable distances, and whether a network already exists. Starting with a handful of machines lowers the absolute initial cost, but because the server and software portion does not scale with machine count, the cost per machine goes up.

What OEE should we target?

TeepTrak’s 2026 benchmark across 450 plants puts the median at 60 percent, the top quartile at 75 percent, and what is called world class at 85 percent. What level is reasonable depends on industry, equipment configuration and product mix, so tracking your own trend on the same machines under the same definitions is more useful than adopting another company’s number as a target. Note especially that immediately after switching from handwritten shift reports to system measurement, losses that used to fall out of the record surface, and OEE can appear to have got worse than before. That is not deterioration, it is improved measurement accuracy, and unless this is shared with management in advance the programme tends to stall right after it starts.

Can minor stoppages be captured on older equipment?

Yes. Even without a communications port, you can read the voltage on the signal tower lamps, tap the existing counter output with a contact in parallel, or detect motion with an external proximity sensor. The granularity of what you can capture does drop. The signal tower only tells you about stops the machine itself judged to be abnormal, so waiting for material and stops caused by human judgement have to be covered another way. Understand what you cannot capture at the outset and you can fill that gap in the reason entry design.

Is automatic part counting alone worth anything?

It is. Once part counts are collected automatically, the distribution of actual cycle times becomes visible, which gives you the evidence base for setting an ideal cycle time. If the distribution has a long tail, meaning occasional extremely slow cycles, you can infer that minor stoppages are hiding there. Without stop detection and reason entry you still cannot say why the cycle was slow, so it will not get you to a specific countermeasure. If you want to phase the work, starting with part counting, adding stop detection, and layering reason entry on last is a realistic sequence.

Summary

Minor stoppage countermeasures miss because of how the metrics are constructed. Short stops below the threshold are not recorded as Availability loss, they dissolve into Performance. So activity organised around an Availability target never reaches them. The place to catch them is Performance, and to do that you have to fix an ideal cycle time, fix the downtime threshold, and design reason codes at a granularity the floor can actually keep entering. All three should be settled before you buy equipment or sensors.

The essence of the investment decision, then, is sensitivity rather than sensor count. In the worked example, the same THB 1,100,000 investment pays back in roughly 15 months if you can move OEE by 3 points, and takes roughly 94 months, which is to say it does not pay back, if you only move it by 1 point. What determines which branch you land on is the reason code design and the routine of reviewing numbers weekly and deciding countermeasures. Layers 3 and 5 of the five layer implementation. Settle internally who runs those two layers and how, before you request hardware quotations, and the investment discussion becomes far more concrete.

TOMAS TECH works with Japanese-affiliated plants in Thailand and across ASEAN on equipment visibility, covering everything from signal acquisition at the machine through reason analysis and the improvement routine. We are glad to help at the early stage of sorting things out, questions such as where to put the threshold and reason codes, how much signal can realistically be taken from your existing machines, or how to estimate your own gross profit contribution. Please use us as material for your evaluation and get in touch through our contact page.

References