Blog

2026.08.09

Factory KPI Management 2026 — Cut by Frequency, Measure the Bottleneck

Factory KPI Management 2026 — Cut by Frequency, Measure the Bottleneck

Factory KPI management rarely stalls because the wrong metrics were chosen. It stalls because nobody decided by what formula, from what source, at what frequency, and for whom each number is produced. A dashboard that lists metric names alone moves no one, even when every metric on it is the right one. This article walks through a practical design sequence — define each KPI as a five-part set, cut the metric set by update frequency rather than by hierarchy, and measure only the bottleneck — together with a cost model built around a factory in Thailand.

Two metrics measured the same factory and pointed in opposite directions, both correctly

What happened in Thai manufacturing in 2026

Thai manufacturing in 2026 offers two headline numbers that appear to point in opposite directions.

The first is the Manufacturing Production Index, or MPI, published by the Ministry of Industry. In June 2026 the MPI was −3.10% year on year, and −1.79% for the second quarter. Average capacity utilization over the same period was 57.47%. Read on their own, these figures say that output is shrinking and that equipment is running at only a little over half of capacity.

The second is the Thai manufacturing PMI published by S&P Global. The July 2026 PMI came in at 54.2, up from 53.6 in June. Since 50 is the line between expansion and contraction, this reading says the sector is expanding.

The same country, the same industry, the same period — one number says contraction, the other says expansion. Is one of them wrong?

Why both are correct — different definitions, different data sources

Neither is wrong. They measure different things, and they collect data in different ways.

The MPI is an index built from actual production volume. It aggregates results reported up from factories, so it shows the change in how many units were actually made compared with the same period a year earlier. It is a record of what has already happened.

The PMI is a survey of purchasing managers. It asks about new orders, output, employment, input prices and inventories — essentially, whether each is better or worse than last month — and converts those answers into an index. It aggregates direction rather than quantity, and because it reflects the order flow that respondents are already seeing on their desks, it tends to lead actual results.

In short, one measures past volume and the other measures current direction. The definitions differ, the data sources differ, and the update frequencies differ. That is why two correct measurements of the same reality can point opposite ways, and it is not a contradiction.

The problem starts when both land on the same meeting slide and the discussion becomes “which one is true.” The question worth debating is not which is true, but which metric has which countermeasure attached to it.

This happens inside your factory every day

Exactly the same thing happens inside a single factory every day.

The utilization rate reported by production control differs from the utilization rate reported by maintenance. The defect rate used by quality assurance differs from the defect rate accounting uses for costing. The on-time delivery rate that sales looks at differs from the one the plant looks at.

Every one of those owners is calculating correctly according to their own definition. Yet the meeting still opens with “how did you produce that number,” and the time available for deciding countermeasures drains away. What happens at the level of national macro indicators happens inside a factory once for every metric it tracks.

This article deals with the design sequence for getting out of that state. It is not about choosing metrics. It is about turning metrics into something that moves people.

Five symptoms that factory KPI management is not working

Check your own situation first. If two or more of the five symptoms below apply, redefining the metrics you already have will pay off faster than adding new ones.

Symptom 1 — meetings start with “how did you produce that number”

Most of the time in the monthly review or the daily morning meeting goes into explaining the numbers themselves. When the numerator, the denominator, the exclusion rules and the covered period are not shared, participants silently reinterpret the figure using the definition in their own head. The moment that reinterpretation happens, the number has stopped being a shared basis for judgment.

Symptom 2 — the dashboard is running, but nobody opens it

Everyone watched it at go-live, and six months later nobody opens it. This is a common ending. The cause is not a hard-to-use screen. The cause is that nothing was decided about what to do after looking at it. If looking produces no action, not looking is the rational behavior.

Symptom 3 — the numbers improve but profit does not

Utilization is up, first-pass yield is up, and the P&L has not moved. This happens without fail when the improvement occurred on equipment that is not the bottleneck. Raising the utilization of non-bottleneck equipment only increases work in process. A metric can improve without any profit impact if the improvement did not occur where money is being lost.

Symptom 4 — problems surface only after the month-end close

Yield and unit consumption become visible only at the monthly cost close. In that case the problem is learned about at least two weeks after it occurred. The information needed to trace the cause has already disappeared from the shop floor, and recurrence prevention degrades into “let’s be more careful.”

Symptom 5 — the metric changes meaning when the owner changes

A successor inherits a spreadsheet built by a predecessor and reorganizes the aggregation logic in the way that makes sense to them. There is no bad intent; it is a well-meant improvement. But because no definition document exists, comparison with the past becomes impossible from that moment. A multi-year trend that turns out to be three different definitions stitched together is not a rare finding.

A KPI is not a metric name — it moves only as a five-part set

Factory KPI Management 2026 — Cut by Frequency, Measure the Bottleneck - figure 1

Here is the core of it. A KPI is not a metric name such as “utilization rate” or “defect rate.” It becomes a metric that moves people only when the following five elements are all present. If even one of the five is missing, a correctly named metric is still just decoration on a dashboard.

Element 1 — the formula, meaning numerator, denominator and exclusions

Write the numerator and the denominator at a granularity that produces the same calculation no matter who reads it. The exclusion rules matter most. Does planned downtime go into the denominator? How is changeover treated? Do trial pieces count in the numerator? A utilization rate defined without settling those three questions will take a different value in every department.

Element 2 — the data source, meaning automatic, manual or calculated

Write down where the number comes from, at the level of system name and table name. “From the production control system” is not enough. Always state which of the three categories applies — automatically acquired, entered by a person, or calculated from other metrics. This classification feeds directly into the reliability judgment discussed later.

Element 3 — the update frequency

Write down when the value refreshes. Real time, hourly, next morning, third business day of the following month. A metric whose refresh cycle does not match the cycle of its countermeasure will become an empty ritual. The next section covers this in detail.

Element 4 — the owner, meaning the person who acts on a deviation, not the person who compiles the number

This is where a very large share of factories get it wrong. The owner of a KPI is not the person who compiles it. It is the person with the authority to actually move people and equipment when the threshold is breached. Make the compiler the owner, and what follows a bad number is not improvement but a review of the aggregation method.

Element 5 — the countermeasure on deviation, meaning the threshold and who does what

Decide in writing, in advance, who does what when the threshold is breached. “Check it” and “keep an eye on it” are not countermeasures. Only when you can write “the maintenance lead goes to the line within 30 minutes, sorts the cause into one of three categories, and records it” do you have a countermeasure.

A worked example of a KPI definition sheet

Here is one complete metric. A metric you cannot write at this granularity is not yet a KPI.

ItemEntry
Metric nameBottleneck equipment availability
ScopeMolding machine No. 3 on line 2
FormulaNumerator = actual running time, denominator = planned running time
ExclusionsPlanned maintenance and lunch break excluded from the denominator. Changeover included in the denominator and excluded from the numerator
Data sourceEquipment run signal captured by a collection unit, automatic
Update frequencyRefreshed every minute, finalized daily
OwnerLine 2 production section manager
ThresholdDaily value falls below 92%
Countermeasure on deviationPresent the top 3 stoppages at the 8 a.m. morning meeting and have maintenance and production assign a same-day owner
Review cycleValidate the threshold every quarter

Build six of these sheets and 80% of your KPI design work is done. Put another way, buying a platform while you have zero of these sheets produces exactly the dashboard described in Symptom 2.

Three different things the phrase “utilization rate” refers to

Nowhere is the need for the five-part set clearer than with “utilization rate.” Inside a factory this phrase refers to at least three different things.

Meaning A — equipment availability

Running time divided by planned running time. This corresponds to Availability in ISO 22400 and is one of the three components of OEE. It measures how much of its available running time the equipment actually ran. Because the denominator is planned running time, hours in which the line was stopped for lack of orders are in principle excluded from the denominator.

The components of OEE and how to capture minor stoppages are covered in detail in Minor Equipment Stoppages and OEE 2026. This article deals with the layer above that — the design of the metric system as a whole.

Meaning B — output against installed capacity

The ratio of what was actually produced to the maximum output the equipment is capable of. The 57.47% average capacity utilization from the Ministry of Industry cited at the top of this article is this meaning. If orders are thin, this figure falls even when the equipment is in perfect condition. It is a demand indicator, not an equipment management indicator.

Meaning C — direct labor ratio

Direct working time divided by attendance hours. It measures how much of a person’s time went into value-adding work. Because the denominator is human time, it is an entirely different metric from anything about equipment.

Mixing up the three inverts the countermeasure

The correct response to “utilization is only 80%” changes completely depending on which meaning is intended.

MeaningWhat it measuresCountermeasure when 80% is lowWhat happens if you get it wrong
A Equipment availabilityStoppages and breakdownsBreak down stoppage causes, revise the maintenance intervalYou chase more orders and inventory balloons
B Output against capacityGap between order volume and installed capacityWin orders, reallocate models, consolidate equipmentYou repair equipment, the number does not move, maintenance burns out
C Direct labor ratioIndirect work and waitingImprove changeover, cut transport and search timeYou invest in equipment and no human time is freed

When all three are mixed together in one meeting, the countermeasure that gets agreed lands on nobody’s problem. Simply writing “is the scope equipment, people or demand” on the first line of the KPI definition sheet all but eliminates this failure.

ISO 22400 is not a certification scheme — how to use it and where the edition stands

ISO 22400, and in particular ISO 22400-2, was created precisely to kill this confusion. It is a definition set for KPIs in manufacturing operations management, standardizing the numerator, denominator and exclusion rules for metrics such as Availability, Effectiveness and Quality Ratio.

First, clear up a common misunderstanding. ISO 22400 is not a certification scheme. It is not the kind of standard you get certified against or audited for. It is a dictionary you consult when writing your own internal definition sheets.

The current edition is ISO 22400-2:2014, with an amendment covering energy-efficiency KPIs added in 2017 as ISO 22400-2:2014/Amd 1:2017. A revision is under way as ISO/DIS 22400-2, which reached stage 40.99 on 17 September 2024, but as of August 2026 it has not been published. Since the publication timing is not fixed as of this writing, the practical approach is to base your internal definition sheets on the current edition and reflect the differences once the revision appears.

Using it in practice is simple. List the metric names your company uses, check whether ISO 22400 contains the same or a near-equivalent metric, and if it does, borrow its treatment of numerator and denominator. If it does not, state explicitly in the definition sheet that this is an in-house definition. That alone sharply reduces the drift in meaning that occurs at handover.

Cut KPIs by update frequency, not by hierarchy

Factory KPI Management 2026 — Cut by Frequency, Measure the Bottleneck - figure 2

This is the point this article most wants to press.

Four bands — real time, daily, weekly, monthly

Most explanations of KPI design cut the set into three tiers: executive KPIs, departmental KPIs and shop floor KPIs. It follows the org chart, so it is easy to grasp. But what breaks in practice is not the tier. It is the mismatch between update frequency and countermeasure.

Place a daily KPI where no countermeasure can be executed daily, and the shop floor starts producing the number “for the report.” Conversely, you cannot run daily improvement off a yield figure that only appears at the monthly cost close.

So, before choosing metrics, split them into four bands by update frequency and confirm first that a real countermeasure exists in each band.

Comparison of the four bands

BandUpdate frequencyRepresentative metricsWho executesCountermeasureWhat happens with no countermeasure
Real-time bandSeconds to minutesStoppage occurrence, cycle time deviation, threshold alertsOperators and maintenance staffGo to the machine on the spot, classify the stoppage cause and record itAlerts keep sounding until someone mutes them
Daily bandOnce a dayBottleneck equipment availability, daily output, top 3 stoppages of the dayProduction section manager and line leaderAssign an owner and a due date to the top causes at the next morning meetingThe daily report becomes a record of “we worked hard yesterday too”
Weekly bandOnce a weekFirst-pass yield, MTBF and MTTR, progress against plan, changeover timeProduction manager and maintenance managerSwap one improvement theme in and out at the weekly meetingThe same improvement themes sit there for six months
Monthly bandOnce a monthYield, unit consumption, labor productivity, on-time delivery ratePlant manager and managementInvestment decisions, staffing, model allocationNumbers end as a status report and are never used in a decision

The rightmost column of that table is the essential one. Do not place a metric in a band whose countermeasure column you cannot fill in.

Running MTBF and MTTR in the weekly band requires that the asset register granularity and the failure-record entry rules are already in place. Those preconditions are covered in How to Choose an Equipment Maintenance Management System 2026. For unit-consumption KPIs in the monthly band, the implementation side of electricity visualization is explained in Energy Monitoring Systems for Factories 2026.

Why KPI trees do not work — decomposition does not align frequency

A KPI tree decomposes a top-level metric such as operating profit down to shop floor metrics. The logic is sound and the deliverable looks good on a slide. It still frequently fails on the shop floor, and the reason is that decomposition does not align frequency.

Suppose you decompose operating profit into manufacturing cost, then into yield, then into first-pass yield by process. On the tree they are continuous. But operating profit is monthly, yield comes from the monthly cost close, and first-pass yield by process can be produced daily. At that point, if you improve daily first-pass yield, the effect cannot be confirmed as operating profit until more than a month later.

The tree is not the problem. What is missing is the step after the tree is drawn — writing the update frequency onto each node and identifying where the frequency jumps. The places where the frequency jumps are exactly the places where improvement stalls.

Metrics whose frequency does not fit become supporting metrics

A metric you want but whose frequency does not match the countermeasure cycle does not have to be thrown away. Put it in a separate slot as a supporting metric.

Supporting metrics stay off the main meeting pack and are referenced in quarterly reviews or in one-off analysis. Without that distinction, the number of primary KPIs grows without limit and you end up in the “nobody is watching any of them” state described later.

Work backwards from the metrics you can capture automatically

Why manually entered KPIs converge on “good numbers”

Manually entered metrics drift toward “good numbers” over time without fail. This is not a story about fraud.

When writing downtime on a paper daily report, a person records a stoppage that actually lasted seven minutes as “five minutes.” It is a round number, it matches memory, and nobody is aware of having told a lie. The same person applies the same rounding every day, so the aggregate is consistently smaller than reality. Worse, because it is consistent, the data looks clean and is never flagged as an anomaly.

And when the person entering the number knows they are evaluated on it, the direction of rounding automatically aligns with what favors them. What happens in a factory that puts a manually entered metric among its primary KPIs is not improvement. It is number-fixing.

So KPI design must not start from “which metrics do we want.” It works backwards from “which metrics can we capture automatically.

Three routes to automatic acquisition

There are broadly three routes to metrics you can capture automatically.

Equipment signals means taking run signals, stop signals and production counters directly from the machine via a collection unit. There is no human in the loop, so reliability is the highest of the three, and almost every real-time-band and daily-band metric can be built from here. Even on older equipment, picking up a current sensor or a stack-light illumination signal is enough to distinguish running from stopped. How to select the monitoring and collection platform underneath is summarized in What SCADA Is and How to Select an OT Platform 2026.

Shop floor entry terminals means operators scanning barcodes at a line-side terminal or handheld. A person operates the device, but what they enter is a fact — when, which lot, at which process — rather than a number the person composes. Reliability is therefore higher than free-form manual entry. The mechanics of progress tracking and work instructions themselves are covered in Process Management System Costs and Selection 2026.

Core business systems means using order, shipment, inventory and cost data already sitting in the production control system or the ERP. Most monthly-band metrics come from here. However, updates are tied to the close calendar, so this route cannot serve the daily band.

How to handle metrics you cannot automate

Metrics that cannot be automated are handled in one of three ways.

First, demote them to supporting metrics. Take them out of the primary KPI set and keep them for quarterly analysis.

Second, replace them with a proxy. “Operator skill level,” for example, cannot be measured directly, but the variation in cycle time within the same process can be captured automatically from equipment signals. Rather than the concept you want to measure, look for a number that moves with it and can be captured automatically.

Third, change the work so the measurement can be automated. Replace writing stoppage reasons on paper with a three-choice selection on a touch panel linked to the stack-light signal, for instance. Since that is an investment decision, it falls under the cost model discussed below.

Metric reliability by data source

SourceReliabilityWhere it fitsWatch out for
Automatic capture from equipment signalsHighPrimary KPIs in the real-time and daily bandsThe meaning of each signal must be validated machine by machine
Shop floor entry terminalsUpper middleTraceability metrics in the daily and weekly bandsYou need operating rules for missed scans and proxy entry
Existing data in core systemsMiddleManagement KPIs in the monthly bandUpdate frequency is bound to the close calendar
Calculated from other metricsDepends on the source dataSupporting metricsIf one input is manually entered, treat the whole thing as manually entered
Manual entry by peopleLowSupporting metrics onlyPut it among primary KPIs and number-fixing follows

Capturing quality metrics automatically first requires designing the key that links inspection results to lots. That point is covered in detail in Quality Data Management Systems 2026.

How many metrics can actually move

One meeting can follow five or six

The number of metrics a single meeting can actually discuss through to a decided countermeasure is, empirically, five to six at most. Each metric takes a minimum of five to ten minutes to cover the change since last time, break down the causes, decide the countermeasure and assign an owner and a due date. Handle six metrics in a 60-minute meeting and the time is gone.

From the seventh metric onward, the numbers appear in the pack but are not discussed. A metric that is present but never discussed gradually becomes an item that is merely read aloud, and eventually nobody reads it.

The fixed-total rule

So fix the total number of primary KPIs up front. Cap it at five to six metrics per line or per department, and whenever someone wants to add a new metric, demote or retire one of the existing ones. One in, one out — that simple.

The value of this rule is not that it reduces metrics. It is that adding a metric forces the discussion “which of the six we currently follow is this more important than.” Metrics added without that discussion become empty rituals with almost no exceptions.

Why a 30-metric dashboard becomes “nobody is watching any of them”

A dashboard with 30 metrics on it looks comprehensive and feels reassuring. In reality it is not 30 metrics being watched. It is none of them being watched.

There are three reasons. First, the viewer has to filter, on every visit, which of the 30 are relevant to today’s decision, and that filtering cost is itself a reason to stop opening the screen. Second, with 30 metrics, several will be red every single day, so red stops functioning as a signal of abnormality. Third, maintaining 30 metrics requires 30 definition sheets, which in practice never get written, so more than half sit there as numbers of unknown definition.

Comprehensiveness produces reassurance, not action. The number of metrics that move is unrelated to the number of metrics that are visible.

A model cost estimate — a Thai factory with 180 employees, 6 lines and 40 machines

Factory KPI Management 2026 — Cut by Frequency, Measure the Bottleneck - figure 3

From here on we talk money. Everything below is an estimate based on our own assumed model and is not the figures of any specific real company. Read it against your own conditions.

Assumptions for the model factory

Assume a Japanese-owned components factory in Chonburi, Thailand. The plant has 180 employees, 6 production lines and 40 principal machines, and runs two shifts. It operates 300 days a year, with planned running time of 3,600 hours per machine per year. Annual production value is 420,000,000 baht.

From that we derive the machine hour rate. 420,000,000 divided by (40 machines × 3,600 hours) = 2,917 baht per hour. Taking a 30% contribution margin ratio, the contribution margin lost when a bottleneck machine stops for one hour is 2,917 × 30% ≈ 875 baht per hour.

The current state is an 8.5% equipment downtime rate and a 1.8% defect rate. Records move from paper daily reports into a spreadsheet by transcription, and yield first becomes visible at the monthly cost close.

That figure of 875 baht per hour is the reference point running through this entire article. Once you know how much money one hour of each machine represents, the KPI design discussion shifts from “do we want to see this” to “how much can we recover.”

Breaking the cost into five layers

The cost of introducing KPI management breaks into five layers. That breakdown matters later.

LayerContentsScenario AScenario BScenario C
1 Collection layerSignal capture unit 18,000 per machine, wiring and power 7,500 per machine1,020,000204,000306,000
2 Platform layerCollection and visualization software, server or cloud740,000240,000240,000
3 Definition layerKPI definition sheets, agreement on formulas, master data cleanup420,000140,000140,000
4 Screen layerDashboards and standing reports560,000180,000180,000
5 Adoption layerTraining, meeting design, hands-on support360,000160,000260,000
Initial total3,100,000924,0001,126,000
Annual running cost372,000138,600168,900

Amounts are in baht. In Scenario A the mix is 32.9% collection, 23.9% platform, 13.5% definition, 18.1% screen and 11.6% adoption.

What deserves attention is the definition, screen and adoption layers. These three are human work rather than hardware, and they barely scale with the number of machines in scope. Producing six KPI definition sheets takes much the same effort whether the scope is 40 machines or 12.

Scenario A — plant-wide at once, 40 machines and 30 metrics

Fit collection units to all 40 machines and stand up a 30-metric dashboard. Initial 3,100,000, annual running 372,000. Of the 40 machines, 8 are bottlenecks.

Coverage is the highest and management tends to like the proposal, but most of the 1,020,000 baht spent on the collection layer goes into equipment where no money is being lost.

Scenario B — one line first, 8 machines and 6 metrics

Start small with a single line. 8 machines in scope, 6 metrics. Initial 924,000, annual running 138,600. Of the 8 machines, 3 are bottlenecks.

The investment lands at under a third of A. It follows the textbook “start small and expand” pattern and is easy to get approved. Yet, as shown below, this is the option with the longest payback.

Scenario C — bottlenecks across lines, 12 machines and 6 metrics

Ignore line boundaries and select only the bottleneck machines across all 6 lines — 12 of them. Six metrics. Initial 1,126,000, annual running 168,900. All 12 are bottlenecks.

The investment rises slightly above B. Going from 8 machines to 12 makes the collection layer 1.5 times larger, and involving stakeholders from multiple lines thickens the adoption layer. The result is nevertheless dramatically different.

Breakdown of annual benefit

The benefit arrives through three routes.

BenefitFormulaABC
Downtime reductionBottleneck machines × 3,600 hours × reduction points × 875378,000141,750604,800
Reduced transcription and compilation effortTarget hours × 70% reduction × hourly rate284,55047,425110,880
Yield improvementProduction value × reduction points × 60% actual material loss302,40084,000252,000
Total annual benefit964,950273,175967,680
Annual running cost372,000138,600168,900
Annual net benefit592,950134,575798,780

Downtime reduction points are 1.5 points against 8 machines for A, 1.5 points against 3 machines for B, and 1.6 points against 12 machines for C. A and B represent bringing the current 8.5% downtime rate down by 1.5 points to 7.0%. C is set at 1.6 points because every machine in scope is a bottleneck and countermeasures concentrate there, allowing a slightly larger reduction.

Yield reduction points are 0.12 points plant-wide for A, 0.20 points for B against the 70,000,000 baht of production value on one line, and 0.10 points plant-wide for C.

The transcription effort breakdown, shown for Scenario A, is 12 shift leaders at 25 minutes a day over 300 days, which is 1,500 hours, plus 2 production control staff at 90 minutes a day over 300 days, which is 900 hours. The combined 2,400 hours is reduced by 70%. Hourly rates are 145 baht for leaders and 210 baht for production control. B is one line out of six, so it is one sixth of A, or 47,425 baht. C assumes that the daily reports for the target equipment account for 40% of the total, giving 2,400 hours x 40% x 70% = 672 hours at a blended rate of 165 baht, or 110,880 baht.

The annual benefit of A and C is almost identical — 964,950 against 967,680. Yet the initial investment is 3,100,000 versus 1,126,000, a gap of nearly three times.

Payback period and five-year totals

Scenario AScenario BScenario C
Initial investment3,100,000924,0001,126,000
Annual net benefit592,950134,575798,780
Simple payback period5.2 years6.9 years1.4 years
Five-year cumulative net−135,250−251,125+2,867,900
Five-year ROI−4.4%−27.2%+254.7%

Over five years, both A and B are net negative. Only C returns +2,867,900 baht, an ROI of +254.7%.

Why starting small does not, on its own, improve the outcome

This is the part of the article with the most practical bite.

Scenario B invests 70% less than A, yet the payback period worsens from 5.2 years to 6.9 years. The textbook advice to start small betrays you here.

The reason lies in the cost structure. The three design layers — definition, screen and adoption — account for 43.2% of the initial investment in A and 51.9% in B. Scope was cut to a fifth, and the weight of the three design layers went up rather than down. That is what it looks like when design cost is a fixed cost.

In fact, the machine count falls from 40 in A to 8 in B, a fifth, while the initial investment falls only from 3,100,000 to 924,000, roughly to 30%. Writing KPI definition sheets, designing dashboards and running the meeting cadence until it sticks simply do not scale with machine count.

Benefit, on the other hand, scales with bottleneck count. A has 8 bottlenecks among its 40 machines, while B has only 3 among its 8. As a result the annual net benefit falls from 592,950 to 134,575, to just over 20%. The investment falls only to about 30%, while the benefit falls to just over 20%. That is why the payback period worsened from 5.2 years to 6.9 years.

C is fast not because of scale and not because of the number of metrics. It is fast because all 12 machines are bottlenecks. If you are going to pay the same design cost either way, payback is faster when that design cost sits on top of equipment where money is being lost.

The conclusion of this article fits in one sentence. What determines payback is neither the breadth of scope nor the number of metrics, but whether what you measure is the bottleneck.

Sensitivity — what happens when the improvement execution rate falls

Everything above assumes that a countermeasure is always executed against every abnormality that becomes visible. Reality does not work that way. So define the improvement execution rate as the share of visible abnormalities against which a countermeasure was actually executed, apply it uniformly across the whole benefit, and look at the sensitivity. The subject is Scenario C.

Improvement execution rateAnnual benefitAnnual net benefitSimple payback period
100%967,680798,7801.4 years
60%580,608411,7082.7 years
40%387,072218,1725.2 years

If the execution rate falls to 40%, payback stretches from 1.4 years to 5.2 years, roughly 3.7 times longer. Note that this table applies the improvement execution rate uniformly across the whole annual benefit. The transcription and compilation saving is realised the moment capture becomes automatic and does not depend on whether a countermeasure was executed, so the real deterioration is gentler than the table shows. This is deliberately the conservative side of the estimate. You can buy the platform, but if nobody acts after the alert, the investment is not recovered.

Open this table whenever the discussion turns to cutting the 360,000 baht or 260,000 baht of adoption-layer spend. Cutting the adoption layer is the same thing as lowering the execution rate.

Why the conclusion points the other way from our article on temperature and humidity monitoring

Here we should be explicit about a point where this article’s conclusion runs opposite to one we published previously.

In our article on temperature and humidity monitoring systems, we concluded that the more monitoring points you add, the further payback recedes. This article concludes the opposite — narrowing the scope alone does not bring payback closer. At first glance that is a contradiction.

It is not. The cost structures differ.

Temperature and humidity monitoringKPI management
Dominant costSensors and installation work, a variable cost proportional to point countDefinition, screens and adoption, a design cost fixed regardless of point count
Distribution of benefitConcentrated at the few places where quality defects occurConcentrated at bottleneck equipment
Adding pointsMarginal benefit falls off rapidlyFixed cost is spread thinner, but non-bottlenecks produce no benefit
Narrowing scopePayback comes closerFixed cost cannot be spread and payback recedes

Temperature and humidity benefits exist only near the place that is broken — one point in the constant-temperature room, two in the drying process. Monitoring points outside those spots produce almost no benefit against their cost. That is why fewer points means faster payback.

In KPI management, design cost is a fixed cost. Narrow the scope and the effort of definition sheets, screens and meeting cadence does not shrink in proportion to machine count. That is why narrowing alone does not bring payback closer.

And the criterion the two share is exactly the same. Identify where the money is being lost first, then measure. For temperature and humidity that means the process producing defects; for KPIs it means the bottleneck equipment. In both cases, the map of money is drawn before the measuring starts.

How to roll it out — run one band in 90 days

Trying to stand up every band at once will always stall in definition debates. Run the daily band only for 90 days first.

Days 0 to 30 — identify where the money is being lost

Before talking about metrics, draw the map of money. Start with the machine hour rate. Divide annual production value by the number of machines in scope and the annual planned running hours. In the model factory that was 2,917 baht per hour.

Then multiply by the contribution margin ratio to get the loss when a machine stops for one hour. That was 875 baht per hour.

Finally, for each of the 6 lines, identify the one or two machines that govern output. The test is simple — when that machine stops, do the upstream and downstream processes wait? If work in process piles up upstream and downstream goes idle, that is the bottleneck. The deliverables from this period are just two things: the list of bottleneck machines, and the hourly money value of each.

Days 31 to 60 — define six metrics as five-part sets

For the machines you identified, produce six KPI definition sheets. Using the table format shown earlier in this article, fill in the formula, exclusions, data source, update frequency, owner, threshold and countermeasure on deviation.

What consumes the most time in these 30 days is agreeing the exclusion rules. Does planned maintenance go into the denominator? How is changeover handled? Is waiting for material an equipment stoppage or a production control issue? Friction here is normal, and if you skip this discussion and install the platform anyway, the same discussion happens after go-live, this time with numbers already on the screen.

At the same time, confirm how many of the six metrics can be captured automatically. If fewer than 4 are automatic, either reselect the metrics or revisit the scope of the collection layer.

Days 61 to 90 — put it on the meeting cadence and run the countermeasures

Stand up the dashboard and put it on the agenda of the next morning meeting. The purpose of this period is not to look at numbers. It is to confirm that countermeasures actually get executed.

Every morning, present the top 3 stoppages, assign owners and due dates, and check the result the following day. Once that loop starts turning, the improvement execution rate rises. If it does not turn, the cause is not the screen — it is either the choice of owner or the specificity of the countermeasure.

Four things to check on day 90

First, how many of the six metrics are captured automatically? The floor is 4 and the target is 5 or more. Second, what percentage of the countermeasures decided at the morning meeting were executed? As the sensitivity table shows, if this falls below 40% the investment case itself changes. Third, by how many points has the bottleneck downtime rate fallen? Fourth, has anyone wanted to add a metric during these three months? If so, decide which one to drop before adding it.

Considerations specific to sites in Thailand and ASEAN

The KPIs headquarters wants and the KPIs the site can act on diverge

KPIs handed down from headquarters in Japan are usually already defined as a group-wide standard. And they are not necessarily capturable automatically at the local site.

In that situation, do not adopt the headquarters KPI as the site’s primary KPI. Keep the headquarters KPI as a monthly-band reporting metric and place a different, automatically captured metric in the site’s daily band. Share a single mapping table with headquarters and the question “why do we have two versions of this number” disappears.

The mapping table needs five columns: the headquarters KPI name, the corresponding local metric name, the difference in definition, the difference in update frequency, and whether conversion between them is possible. Whether that table exists makes a large difference to the effort of monthly reporting.

The same metric name does not cover the same scope in three languages

A site in Thailand handles the same metrics in three languages — Japanese, English and Thai. What happens constantly is that the scope shifts the moment the term is translated.

The Japanese word for “utilization rate” becomes Availability in meaning A, Capacity Utilization in meaning B, and Direct Labor Ratio in meaning C — three entirely different English terms. Because one word in Japanese splits into three in English, the meaning is fixed at the moment the translator picks a word.

The countermeasure is simple. Write the metric name in all three languages on the KPI definition sheet and write only one formula. State explicitly that the formula, not the translated term, is the authoritative version, and the scope survives crossing languages.

Etiquette when using external indices in an investment proposal

The MPI and PMI discussed at the top of this article are useful when explaining the local market environment in a proposal to headquarters. There is an etiquette to using them.

First, always state the source and the definition alongside the figure. The MPI is based on actual production volume and the PMI is based on a survey of purchasing managers, and it is entirely normal for the two to point in opposite directions. Quote only one and you will be confronted with the other later.

Second, do not use an external index as the basis for your own target values. Average capacity utilization of 57.47% describes the demand environment of Thai manufacturing as a whole and says nothing about the management standard of your own equipment. Use external indices to explain the environment, and build target values from the money at your own bottleneck.

The primary source can be checked at the Industrial Indices published by the Office of Industrial Economics under the Ministry of Industry. For any number that goes into a proposal, verify it at the publishing body rather than in a summary article.

Designing KPIs on the assumption of high turnover

Manufacturing sites in Thailand need to be designed on the assumption that people turn over faster than in Japan. That has two consequences for KPI design.

First, manually entered metrics that depend on individual skill become even less trustworthy. When the person who learned the entry rules leaves, the character of the aggregate changes from that day. The case for prioritizing automatic capture is a notch stronger here than in Japan.

Second, the KPI definition sheet becomes the handover document itself. With a definition sheet, a successor does not have to guess at the predecessor’s aggregation method. In a factory without one, the meaning of a metric mutates at every handover and trend analysis stops being possible.

The minimum wage and the denominator of labor productivity KPIs

Many factories place labor productivity in the monthly band. Because its denominator is human time, the metric is not directly affected by wage revisions. It is affected, however, whenever the metric is evaluated in money terms.

Minimum wage revisions in Thailand are proceeding in stages, and in Bangkok an increase to 400 baht per day has been implemented. Levels differ by province, so the amount applicable to your own location needs to be checked individually.

There is one practical caution. Define the denominator of labor productivity as actual working hours, not headcount. With headcount as the denominator, changes in overtime never register in the metric, and in a period of rising wages you end up in the inexplicable position of “productivity is improving but labor cost is rising.” Writing one line in the denominator field — actual working hours, overtime included — prevents it.

Five common failures and how to avoid them

Failure 1 — buying the platform first. Installing collection units and software, then starting the discussion about what to measure. Definition debates always create friction, so months pass with the platform running and nothing decided. The fix is to write six definition sheets before requesting quotations. With definition sheets in hand, the number of collection points you need is determined automatically.

Failure 2 — instrumenting every machine. Coverage produces reassurance, but investment in equipment that generates no money contributes nothing to payback. This is exactly the difference between Scenario A and Scenario C. The fix is to put bottleneck identification ahead of the investment decision.

Failure 3 — adding metrics forever. Add one every time someone says “I would like to see this number too” and you will have 30 metrics within six months. The fix is the fixed-total rule. Every addition is paired with a deletion.

Failure 4 — making the compiler the owner. When a bad number appears, what starts is a review of the aggregation method. The fix is to write the definition into the owner field — the person with authority to move people and equipment when the threshold is breached.

Failure 5 — cutting the adoption layer. Training and meeting design are the first items cut from a quotation. But as the sensitivity table shows, if the execution rate falls to 40%, Scenario C payback goes from 1.4 years to 5.2 years. The fix is to present the adoption layer in the proposal not as a cost but as an investment in the execution rate.

Summary

Here is what has to be decided to make factory KPI management work.

A KPI is not a metric name. It is a five-part set — formula, data source, update frequency, owner, and countermeasure on deviation. A metric missing any of the five will not move, however correct its name.

“Utilization rate” refers to three different things — equipment availability, output against installed capacity, and the direct labor ratio. ISO 22400 is a definition set rather than a certification scheme, the current edition is the 2014 version plus the 2017 amendment, and the revision remains unpublished as of August 2026.

Cut metrics by update frequency rather than by hierarchy, and confirm that a real countermeasure exists in each of the four bands — real time, daily, weekly and monthly — before choosing the metrics. Never place a metric in a band that has no countermeasure.

Because manually entered metrics always converge on good numbers, design works backwards from which metrics can be captured automatically. Fix the number of primary KPIs at five to six, and drop one for every one you add.

And what determines payback is neither the breadth of scope nor the number of metrics. It is whether what you measure is the bottleneck. In our assumed model, plant-wide at once paid back in 5.2 years, one line first in 6.9 years, and bottlenecks across lines in 1.4 years. Starting small does not by itself bring payback closer, because design cost is a fixed cost.

What to do in the first 90 days is not to choose metrics. It is to work out how much money one hour of each machine represents.

Where to turn while you are still deciding how to design your KPIs

If you want to think this through against your own situation, there is no need yet to get into products or quotations. What is genuinely hard to decide sits earlier than that — which machine is the bottleneck, what one of its hours is worth, and which metrics can realistically be captured automatically.

We work with Japanese-owned manufacturers based in Thailand on collecting and visualizing production and equipment data, and we are happy to talk from the stage before KPIs are chosen — the stage of mapping where money is being lost. A conversation limited to checking which metrics your equipment configuration could capture automatically, or whether signals can be taken from your existing machines, is perfectly fine.

Please get in touch through our contact page.

Frequently asked questions

How many metrics should factory KPI management start with?

Start with a cap of five to six metrics per line or per department. A single meeting can discuss at most five or six metrics through to a decided countermeasure, and anything beyond that appears in the pack without being discussed. We recommend narrowing the start to six daily-band metrics and building the set so that at least 4 of them can be captured automatically. When you want to add a metric, always demote one of the existing ones to a supporting metric first.

What is the difference between utilization rate and OEE?

OEE is the product of three components, one of which is availability. If “utilization rate” is being used in meaning A, equipment availability, then it is one of the components of OEE. Inside a factory, however, the same phrase is also used for output against installed capacity and for the direct labor ratio. The practical difference that matters most is that OEE has a clear definition as the product of three components, while “utilization rate” refers to three different things depending on context. In meetings, never use the phrase on its own — always pair it with the formula.

How far can spreadsheet-based KPI management take you?

For monthly-band metrics, a spreadsheet is entirely adequate. If all you do is process data extracted from the core system once a month, no dedicated platform is required.

Where it stops working is the daily and real-time bands. For capturing stoppages at minute-level granularity, the numbers are rounded the moment a person transcribes them and converge on a downtime figure smaller than reality. The test is simple — is the source data for that metric entered by a person? If it is, that metric should not sit among your primary KPIs, spreadsheet or not.

Does ISO 22400 require certification?

No. ISO 22400 is not a certification scheme. It is a definition set for KPIs. It is not the kind of standard you get audited against or registered under. It is a dictionary you consult for the treatment of numerator, denominator and exclusion rules when writing your own internal KPI definition sheets.

The current edition was published in 2014, with an amendment covering energy-efficiency KPIs added in 2017. The revision is progressing as ISO/DIS 22400-2 and reached stage 40.99 on 17 September 2024, but it has not been published as of August 2026. Basing internal documents on the current edition and reflecting the differences once the revision appears is entirely sufficient.

Should we build a KPI tree?

Building one is worthwhile, but building it is not enough to make it work. A tree shows a logical decomposition and draws a line from operating profit down to shop floor metrics. Decomposition, however, does not align the update frequency of each node.

What we recommend is a step after the tree is drawn — write the update frequency onto every node. Wherever you find a daily node hanging below a monthly node, that is the break point where the effect of an improvement can no longer be confirmed. Use the tree as a diagram of relationships, and do the actual operating design with the four frequency bands.

Why does profit not increase even though the KPIs are improving?

The most common cause is that the equipment that improved is not the bottleneck. Raising the utilization of a non-bottleneck machine only increases work in process; total plant output does not change. The metric improves and the P&L does not move.

The second cause is a denominator that does not match reality. Define labor productivity with headcount as the denominator and the metric will look better even as overtime rises.

The third is that the benefit was never converted into money. If the downtime rate fell by 1.5 points and nobody calculated what that is worth per year, there is no way to connect it to the P&L. Deriving the machine hour rate and the contribution margin ratio first is partly what makes that connection possible.

References

  • Thailand Ministry of Industry, Manufacturing Production Index for June 2026, Xinhua, 27 July 2026 — read the article
  • S&P Global Thailand Manufacturing PMI for July 2026, Xinhua, 3 August 2026 — read the article
  • ISO 22400-2:2014, current edition — official ISO page
  • ISO 22400-2:2014/Amd 1:2017, amendment on energy-efficiency KPIs — official ISO page
  • ISO/DIS 22400-2, under revision, reached stage 40.99 on 17 September 2024 — official ISO page
  • Office of Industrial Economics, Ministry of Industry, Thailand, Industrial Indices — official OIE page
  • JETRO, Minimum Wage in Bangkok Raised to 400 Baht per Day, 4 July 2025 — JETRO business news