Blog

2026.08.01

Demand Forecasting AI ROI for Thai Factories – MAPE and Payback

Demand Forecasting AI ROI for Thai Factories - MAPE and Payback

Demand Forecasting AI ROI for Thai Factories – MAPE and Payback

In 2026, the export outlook for Thailand differs by more than ten percentage points depending on whether you ask a government agency or an industry body. The same year, the same country, two very different numbers. In a year when the external forecasts themselves are that far apart, spending money to buy a forecast that will be right is a poor investment. What you actually need is a mechanism that manages how you are wrong. This article sets out why evaluating demand forecasting AI on MAPE alone will lead you to the wrong decision, what two additional metrics – bias and FVA – add to the picture, and what the cost and payback arithmetic looks like for a plant with roughly 500 million baht in annual revenue. It is written from the perspective of Japanese-affiliated manufacturers operating sites in Thailand and Vietnam.


Thailand’s 2026 demand outlook is split by double digits

Before you look inside your own plant, it is worth confirming the assumptions sitting outside it. For 2026, the public sector and the industry associations in Thailand do not agree with each other in the first place.

Nantapong Jiralertpong, Director-General of the Trade Policy and Strategy Office (TPSO) at Thailand’s Ministry of Commerce, puts the 2026 export outlook at -3.1% to +1.1%. For GDP growth, citing IMF and World Bank projections, the figure is 1.6 to 1.7%. The Thai National Shippers’ Council, on the other hand, revised its 2026 export growth outlook sharply upward, from a previous 2 to 4% to 8 to 10%, on the strength of demand for electronic components and AI-related goods.

For exports from the same country in the same year, the floor is -3.1% and the ceiling is +10%. The spread is more than 13 percentage points. This is not a case of one side being wrong. It is a case of two parties looking at different product mixes under different assumptions. Look at a portfolio weighted toward electronic components and AI-related goods and you turn bullish. Look at the whole economy, including consumer goods and automotive, and you turn bearish. If you summarise 2026 in Thailand as simply growing or not growing without understanding that structure, you will build the wrong production plan.

By market, the directions point opposite ways

The TPSO outlook broken down by market shows the picture in more detail.

Market2026 outlook2025 actual
United States+8.9%+30.6%
China+1.2%+13.6%
Vietnam+1.5%n/a
Indonesia+2.2%n/a
Cambodia-40.8%n/a
Laos+11.3%n/a
India+5.4%n/a
Middle East and UAE+8.3%n/a
Germany-1.4%n/a
United Kingdom-2.4%n/a
Netherlands+4.1%n/a
South Africa-1.6%n/a

Source: Trade Policy and Strategy Office (TPSO), Ministry of Commerce, Thailand

What stands out is that the United States and China are still positive, but the deceleration from 2025 is extreme. Exports to the US go from +30.6% to +8.9%, and to China from +13.6% to +1.2%. TPSO explicitly attributes this to the payback from front-loading in 2025, when shipments were pulled forward in anticipation of tariffs, and the accumulated inventory is now being drawn down. It also cites tariff-driven consumer price increases as a further drag on the demand side.

Then there is Cambodia at -40.8%. This is the single most difficult type of change for a forecasting model to handle. It sits nowhere near the extrapolation of historical data, and statistically it simply looks like an outlier. In reality it is a structural shift, and if you let the model remove it as an outlier, the model will keep making the same error next month and the month after that.

Current manufacturing output is already decelerating

Outlooks are one thing. Actuals are worth checking too. Thailand’s Manufacturing Production Index (MPI) for May 2026 was -0.8% year on year. Automotive production was down -17.94%, a sharp deterioration. The Ministry of Industry’s full-year 2026 MPI outlook is +1.0 to 2.0%. In other words, the current reading is negative while the full year is expected to be slightly positive, which means the number already assumes a recovery somewhere during the year.

For the medium-term picture, Krungsri Research’s Industry Outlook 2026-2028 sees Thai economic growth averaging 2.1% a year over 2026-2028. That is well below the 2010-2019 average of 3.6%. Household debt stands at over 86% of GDP, which structurally caps how much domestic demand can expand.

Vietnam, by contrast, is in an expansion phase

Here an asymmetry emerges that matters a great deal to any Japanese-affiliated manufacturer with sites in both Thailand and Vietnam.

Vietnam’s Index of Industrial Production (IIP) for the first half of 2026 rose +10.8% year on year. That is the strongest growth since 2019, accelerating further from +8.7% in the first half of 2025. Manufacturing and processing grew +11.4%, contributing 8.9 percentage points to the overall increase. The manufacturing PMI for June 2026 was 51.8 (May was 52.8), marking 14 consecutive months of production expansion. Employment, however, has fallen for four consecutive months, which points to a pattern of raising output without adding headcount.

In practice this means that within a single company, the Thai site can be in a slowdown while the Vietnamese site is in an expansion. If you build one group-wide demand forecasting model and push it out to both sites, at least one of them will be badly wrong. We return to this point later.


What demand forecasting AI actually replaces

The term demand forecasting AI is used widely, but what it actually replaces on the shop floor is rarely spelled out. If you secure a budget while that remains vague, the reaction after go-live will be that this is not what we thought we were buying.

Excel plus sales intuition versus AI – granularity, frequency, explainability

In many Japanese-affiliated plants in Thailand, the demand forecast is currently produced something like this.

  • Sales collects advance notices and forecasts from major customers
  • Production control lines up year-on-year comparisons and a three-month moving average in Excel
  • At the monthly supply and demand meeting, the numbers are adjusted using the sales team’s feel for the market
  • The agreed numbers are passed to production planning and purchasing

There is nothing inherently wrong with this. In a plant with few customers, accurate advance notices and short lead times, it works perfectly well. The problem appears when SKUs run into the hundreds, when make-to-stock items without advance notices are mixed in, and when raw material procurement lead times run to two or three months. Done manually, the process breaks down if you make the granularity finer, and it breaks down again if you raise the frequency.

What demand forecasting AI actually replaces is exactly that limit on granularity and frequency. The differences boil down to the following.

DimensionExcel plus sales intuitionDemand forecasting AI
GranularityProduct group and major SKUs at bestCan go down to every SKU by site and by customer
FrequencyMonthly, biweekly at bestAutomatic refresh daily to weekly
External variablesLive in people’s heads but never enter the formulaHolidays, FX, tariffs, temperature and more can be modelled explicitly
ExplainabilityBecause a particular salesperson said soCan be decomposed into contribution factors, depending on the model
ReproducibilityFalls apart when the person in charge changesSurvives as a documented procedure
Strong atAbsorbing exceptions, one-offs and sudden ordersRunning large volumes of average items reliably

The last row is the important one. Demand forecasting AI is not there to override and replace the judgement of your sales team. Information that only sales holds – the probability of closing a new deal, a customer relocating a plant, a competitor exiting a segment – will always live outside the model. What AI takes on is the part that should produce the same answer no matter who does it, done at a finer grain and more often than a human can manage. Miss that distinction and you head straight for the most common failure of all, where the sales team overwrites every number the AI produces.

Statistical models versus machine learning – external variables are the dividing line

Demand forecasting methods fall broadly into three layers.

Layer 1 – time series statistical models (moving average, exponential smoothing, ARIMA, seasonal decomposition)

These build the future from the item’s own history alone. They work with little data, the results are easy to explain, and they are computationally light. For stable-demand items this is often enough, and in fact, when you run the FVA evaluation described later, it is not unusual to conclude that the AI could not beat the statistical model.

Layer 2 – machine learning (gradient boosting, regression families)

In addition to the item’s own history, these accept external variables as explanatory inputs. Holiday calendars, working days, promotions, FX rates, the customer’s own production plan, raw material prices, tariff changes and so on. They can also use similarity across SKUs, filling in for a data-poor SKU using the behaviour of others.

Layer 3 – deep learning (RNN and Transformer-family time series models)

These learn many series simultaneously and capture interactions between series and long-range dependencies. They come into their own at a scale of several thousand SKUs and above, but they demand much more training data and much higher data quality, and explainability drops.

For Japanese-affiliated plants in Thailand, the realistic starting point is almost always Layer 2. The dividing line is clear: do you actually hold, as data, the external factors that move your demand? The way Songkran falls in the calendar, which week Lunar New Year lands in, your customers’ operating calendars, the month a tariff took effect. If you have those as numbers, Layer 2 is meaningful. If you do not, then creating that data is the investment target, not the model.

It is also worth noting that AI accuracy depends heavily on the quality and volume of training data, and analyses of implementation cases repeatedly point to data preparation as the key to success. Before you install an advanced Layer 3 model, confirm that it is not losing to Layer 1. That is the correct order.

How this relates to production planning AI and inventory optimization AI

This is the distinction that causes the most confusion in practice. Laid out properly, the three functions sit in a single line.

“`

Demand forecasting AI -> Inventory optimization AI -> Production planning AI

(what will sell) (how much to hold) (when to make what)

“`

  • Demand forecasting AI outputs future demand volume and its uncertainty, meaning the spread. The critical point here is that the output is not only a point estimate but the width of the distribution.
  • Inventory optimization AI takes that forecast and its spread and sets safety stock, reorder points and order quantities per SKU. With the same forecast, the answer changes depending on your service level target and how you price the cost of a stockout.
  • Production planning AI (the production scheduler) takes required quantities and due dates and decides when to make what and in what sequence, under constraints of equipment capacity, changeover time and labour shifts.

The practical conclusion that follows is that installing demand forecasting AI on its own will not move any of your plant metrics. If forecast accuracy improves but your safety stock formula is still a fixed value set ten years ago, inventory will not fall. If reorder points are still set manually, purchasing will not change. If production runs on fixed weekly lot sizes, production will not change either.

The forecast is an input, not a result. What produces results is the vessel that receives the forecast and acts on it. For the options on the vessel side, we have organised the products actually in use in Thailand in our production scheduler comparison and our factory inventory management system comparison, which are worth reading alongside this article.

TOMAS TECH is not a company that develops or sells demand forecasting engines. Based in Bangkok, we work on the other side of the line, building the production planning, inventory and purchasing vessel that receives the forecast, as part of IT and OT integration for Japanese-affiliated plants. That is precisely why we have seen so many cases of a company that bought the engine and has nowhere to put it.


Reading accuracy metrics – MAPE alone will mislead you

Demand Forecasting AI ROI for Thai Factories - MAPE and Payback - figure 1

The first number that comes up in any vendor comparison is MAPE. But if you decide on MAPE alone, you will very probably decide wrongly. There are three layers of metrics you need.

MAPE benchmarks by segment

MAPE (mean absolute percentage error) tells you by what percentage, on average, the forecast missed the actual. The first thing to internalise is that whether a MAPE is good or bad is determined by the industry segment and the granularity.

SegmentIndicative MAPENotes
FMCG and consumer staples, stable demand, class A and X SKUs10 to 25%Deteriorates to 25 to 35% during promotional periods
Apparel and short-lifecycle items35 to 60%Short lifecycles mean history never accumulates
Industrial and B2B, project-driven items (per SKU)20 to 40%Improves when aggregated to product group level
Range generally regarded as acceptableUnder 10 to 20%Depends on demand variability and lifecycle stage

Most Japanese-affiliated plants in Thailand belong in the third row, industrial and B2B. Which means a SKU-level MAPE of 20 to 40% is not a bad number with room for improvement, but the standard level for that domain. If you assess it without knowing that, and declare that a MAPE of 30% is unacceptable, you are misreading a healthy state as a defect.

Conversely, when a vendor shows you a reference case claiming MAPE of 10%, always ask which segment it came from, at what granularity and over what forecast horizon. MAPE at SKU by week and MAPE at product group by month are entirely different metrics. The latter will produce a better number than the former in almost every case, because aggregation lets the overshoots and undershoots of individual SKUs cancel each other out. A MAPE comparison that does not state the granularity is meaningless.

Separate out bias, the systematic over- or under-read

The fatal weakness of MAPE is that it destroys the direction of the error. Because it takes absolute values, a model that always reads 10% high and a model that scatters randomly at +10% one month and -10% the next both show up as MAPE of 10%.

For a factory, though, those two are completely different animals.

  • Random error is what safety stock is for. Hold inventory in proportion to the spread and you can maintain your service level.
  • Systematic bias is not something safety stock can cover. If you always read high, inventory accumulates monotonically. If you always read low, stockouts occur structurally. It is a problem that requires correcting the median of the forecast itself.

Bias acts on inventory like an integral. Read 5% high every month and inventory grows every month. A plant where MAPE is flat but inventory keeps climbing is almost certainly leaving bias unaddressed.

In practice, review these three side by side, monthly.

  1. MAPE (the size of the error) – used to design safety stock
  2. Bias (the direction and persistence of the error) – used to correct the forecasting process and the adjustment rules
  3. Inventory outcomes (days of inventory, stockout rate, write-off value) – used for the final pass or fail judgement

There are also plenty of situations where MAE (mean absolute error) or MASE is more informative than MAPE. In particular, for items with weeks of zero demand, MAPE has a zero denominator and becomes either uncomputable or divergent, so it simply cannot be used. In plants with many intermittent-demand items, the important thing is not to default to MAPE as the only metric at the point where you select your measures.

FVA (Forecast Value Added) – a forecast that cannot beat a naive forecast is a cost

The third layer is the metric most directly tied to the investment decision.

FVA (Forecast Value Added) measures how much additional value the effort produced, relative to a naive forecast. A naive forecast is the simplest thing you can build at zero cost, such as next month’s demand equals this month’s actual, or next month equals the same month last year.

The evaluation procedure is straightforward.

  1. Compute the MAPE of the naive forecast, over the same period and the same items
  2. Compute the MAPE of the demand forecasting AI
  3. If the difference is positive, that AI is adding value
  4. If the difference is zero or negative, that AI is pure cost

What matters even more is measuring FVA at each step of the process. A real supply and demand process is multi-stage, like this.

“`

Naive forecast -> Statistical model -> Machine learning model -> Sales adjustment -> Consensus number at the supply and demand meeting

“`

Measure MAPE after each step and you often get a surprising result. Specifically, the number after sales adjustment is less accurate than the output of the machine learning model. This is not rare. Sales teams carry strong memories of being blamed for stockouts, so they are structurally inclined to read high. In that case, the correct move is not to make the AI more sophisticated, but to require a documented rationale for every sales adjustment and to stop unsupported upward adjustments. Accuracy improves with zero additional spend.

As a way to design your PoC evaluation, FVA is essential. Sign a contract on the strength of the vendor’s MAPE alone and you will discover, after signing, that you installed an impressive model that turned out to be no more accurate than a year-on-year comparison. Write into the PoC acceptance criteria that FVA against the naive forecast over the same period must be positive.


Where demand forecasting AI works and where it does not

Demand forecasting AI does not deliver the same effect in every factory. Expected value varies considerably with the production mode and the nature of demand.

Expected value differs by make-to-stock, make-to-order and repeat orders

Production modeExpected value of demand forecasting AIReason
Make-to-stock (MTS)HighThe forecast directly sets production volume and inventory level. Error converts directly into cost, so improvements are easy to express in money
Repeat orders with advance noticeMedium to highThe way the advance notice misses becomes the forecasting target. Modelling the gap between notice and actual order pays off
Pure make-to-order (MTO)LowDemand is fixed by the order. The forecast target is not finished goods but procurement volumes for common parts and raw materials
One-off custom ordersAlmost noneThere is no reproducibility in historical data. Pipeline management is more useful

The case most often overlooked is repeat orders. Plants that receive three-month rolling advance notices from customers tend to assume no forecast is needed because the notice exists. In our experience on site, however, plants where the notice and the actual order diverge by tens of percent are not unusual. In that case what you should be forecasting is not demand itself but the error structure of the advance notice. Patterns such as company A’s three-month-out notice runs about 15% high on average, or company B tends to place additional orders at month end, are entirely modellable and tend to produce results quickly.

Conditions where it works poorly

Where the following conditions apply, the return on demand forecasting AI becomes difficult. This does not mean you should abandon the idea, it means you should limit the scope of application.

Condition 1 – project-driven demand with no reproducibility in historical data

Where demand is determined by whether you win a specific large project, learning from past shipment history will not predict the future. What you need here is not demand forecasting AI but probability management of the sales pipeline.

Condition 2 – occurrence frequency of a few units per month per SKU

For items with intermittent demand, there is almost no statistically learnable information. SKU-level MAPE tends to become uncomputable as well. In this domain it is more practical to switch to a qualitative decision about whether to hold stock at all, rather than to forecast.

Condition 3 – procurement and production lead times shorter than the forecast horizon

In a plant where raw material procurement takes a week and production finishes in three days, there is no need to act in advance on the basis of a forecast. You can start once the order arrives and still be on time. Forecast value only arises when lead time is long relative to the forecast horizon. Put the other way round, if you can shorten your lead times, doing that may well offer a better return.

Condition 4 – demand is readable but supply is unstable

Where supplier on-time delivery is poor and the primary cause of stockouts sits on the supply side rather than the demand side, improving the forecast will not reduce inventory. The investment target here is performance management of supplier delivery dates, plus a redesign of safety stock that reflects supply lead time variability.

Why to start at product group level

A common failure at implementation is to start immediately at all SKUs by week. SKU-level MAPE for industrial goods is 20 to 40% as standard, so if you start at fine granularity, the moment stakeholders see the first result they will judge the accuracy too low, and the project stops.

There are three reasons to start at product group level.

  1. Aggregation cancels errors and lowers MAPE. The numbers land at a level people can read, so you start from a state where stakeholders can believe the results.
  2. It matches the granularity of the decision. Line capacity planning, labour shifts and procurement envelopes for common raw materials are usually decided at product group level, not at SKU level in the first place. The principle is to match the granularity of the forecast to the granularity of the decision.
  3. You can descend in stages. Once it works at group level, you can take just the class A items down to SKU level. The opposite direction, starting fine and going coarse, gets read as failure and becomes politically difficult.

The practical question when choosing granularity is this: looking at that forecast, who decides what, and when? A forecast at a granularity where nobody decides anything has no value, however accurate it is.


How to think about cost and payback

This is the core of the investment decision. In the calculations below we use a consistent reference conversion at 1 THB ≈ 4.6 JPY. Amounts are stated primarily in baht, with Japanese yen as a reference figure only.

Rounding policy: amounts are stated in millions of baht to two decimal places (third decimal rounded), days of inventory to one decimal place, and rates as whole percentages. Yen conversions are reference values rounded to the nearest 10,000 yen. As a result of rounding, the sum of components and the stated total may differ in the last digit.

There are four cost buckets, not one

If you estimate the cost of demand forecasting AI as the licence fee, you will almost certainly come up short. Costs land in four buckets.

Bucket 1 – data preparation

This is where estimates are weakest. Cleansing shipment history, deduplicating the item master (the problem of one physical item registered under several codes), separating inter-site transfers from real demand, and putting in place a mechanism to start recording stockouts and lost orders. That last item matters most. Almost every plant has a history of what it sold, and almost none has a history of what it failed to sell. Demand during a stockout leaves no record, so training on the raw history produces a model that structurally under-estimates demand.

Bucket 2 – forecasting model and licences

The SaaS subscription, or the development cost if you build in-house. This is the only bucket that appears explicitly on the vendor’s quotation.

Bucket 3 – integration with existing systems

Interfaces to the production management system, the inventory system and the purchasing system. As long as a person is copying forecast results across by hand in CSV, daily refresh will never survive in operation. Cut this and it always comes back at you later in the form of a PoC that succeeded and a production rollout that will not run.

Bucket 4 – operating structure

Designing the monthly review meeting, documenting the decision criteria, training the people involved, and writing the handover procedure for when they change roles. A system where nobody has defined who looks at the forecast and what they do when it misses will be ignored within three months.

One equation that converts benefits into money

Demand Forecasting AI ROI for Thai Factories - MAPE and Payback - figure 2

The biggest trap when converting benefits into money is counting the same inventory twice. We are going to be strict about this.

Baseline (the only current-state set of values used in this article)

Every calculation that follows is derived from this single baseline, along one consistent line of reasoning.

ItemValueReference in JPY
Annual revenueTHB 500.0 millionapprox. JPY 2.3 billion
Cost of goods sold ratio73%
Annual COGSTHB 365.0 millionapprox. JPY 1,679 million
COGS per dayTHB 1.0 millionapprox. JPY 4.6 million
Days of inventory (COGS basis)60.0 days
Inventory valueTHB 60.0 millionapprox. JPY 276 million
Current MAPE (product group by month)28%
Current MAPE (SKU by week)38%
Inventory holding cost rate (annual)18%
Annual lost sales opportunity from stockouts and late deliveryTHB 6.0 million (1.2% of revenue)approx. JPY 27.6 million
Contribution margin applied to recovered stockout volume25%

COGS per day is set to exactly 1.0 million baht, so one day of inventory reads as 1.0 million baht of inventory value.

Note that the gross margin computed mechanically from the 73% COGS ratio is 27%, but sales made by recovering stockouts carry variable selling costs (freight, promotion, credit), so for converting benefits we use a conservative contribution margin of 25%, two points lower. Every appearance of 25% in the calculations below refers to this contribution margin.

The 18% inventory holding cost rate is a modelling assumption in this article. We assume it breaks down as cost of capital 7%, storage and logistics 6%, write-off, obsolescence and valuation loss 3%, and insurance, stocktaking and administration 2%. When you run this for your own plant, replace those four items with your actual amounts.

Breaking down the 60.0 days of inventory

Split inventory into what is reducible and what is not. Without that split, an unfounded target such as cut inventory by 10% takes on a life of its own.

CategoryRaw materialsWIPFinished goodsTotalReducible
Cycle stock (lot size and transport)11.0 days15.0 days26.0 daysNo, determined by order lot and lead time
Safety stock (buffer against variability)7.0 days13.0 days20.0 daysPartly
Work in process8.0 days8.0 daysNo, this is the domain of process improvement
Dead stock and mis-forecast stock2.0 days4.0 days6.0 daysYes
Total20.0 days8.0 days32.0 days60.0 days

Set exactly one improvement target for forecast error

We take MAPE at product group by month from 28% to 21%, a 25% reduction in error. How that 25% is arrived at is discussed in the next section.

Benefit 1 – inventory reduction (working capital effect)

Safety stock exists to buffer both demand variability and supply variability. Improving forecast accuracy only affects the demand side. Here we assume that of the 20.0 million baht of safety stock, 70% (14.0 million baht) is driven by demand variability and 30% (6.0 million baht) by supply variability.

  • Safety stock driven by demand variability: 20.0 x 70% = THB 14.0 million
  • Reducible through a 25% cut in error: 14.0 x 25% = THB 3.50 million
  • Suppression of dead and mis-forecast stock: 6.0 x 40% = THB 2.40 million
  • Total inventory reduction: 3.50 + 2.40 = THB 5.90 million (approx. JPY 27.14 million)

In days of inventory this is 60.0 days to 54.1 days, a reduction of about 10%.

The annual profit and loss effect is the reduction in holding cost on that inventory.

5.90 million baht x 18% = 1.062, so THB 1.06 million per year (approx. JPY 4.88 million)

Strictly speaking, reducing dead stock should be evaluated separately as avoided write-off loss, but here we conservatively apply the 18% holding cost rate uniformly.

This is the exact point where double counting happens. The 5.90 million baht of inventory you removed is not profit. It is a one-time release of working capital as a cash flow, not a benefit that recurs annually. The only thing you may book to the profit and loss statement every year is the 1.06 million baht of holding cost reduction. The moment you write that reducing inventory by 5.90 million baht produces an annual benefit of 5.90 million baht, your investment case is off by a factor of more than five.

Benefit 2 – fewer stockouts (recovering sales opportunity)

At this point the obvious question arises. If you are reducing inventory, why do stockouts also fall?

The answer is that the inventory you reduce and the inventory causing stockouts are different inventory.

  • What you reduce is inventory that is not moving. Safety stock on SKUs that are in surplus because demand was read high (3.50 million baht), and stock that went stale because the forecast was wrong (2.40 million baht).
  • Where stockouts occur is on SKUs where demand was read low. Those need more inventory, not less.

Improving forecast accuracy is not about reducing the total. It is about becoming able to shift inventory from the SKUs that have too much to the SKUs that have too little. The total falls because the surplus released exceeds the shortfall topped up. Here we should state the limits of this calculation openly. The 5.90 million baht of inventory reduction in Benefit 1 is a gross figure that counts only the surplus side released. The top-up on the shortfall side cannot be estimated until you have per-SKU stockout records, so it is not deducted in this model. The actual net reduction will therefore be smaller than 5.90 million baht. With that understood, the reason counting inventory reduction and stockout reduction together is not double counting is that the two occur on different groups of SKUs. Conversely, be suspicious of any business case that adds both without being able to show the basis for that reallocation. When you rebuild this calculation for your own plant, always present the surplus released and the shortfall topped up as separate figures.

  • Baseline lost sales opportunity: THB 6.0 million per year
  • Reduction rate (planning assumption): 20%, so recovered sales 6.0 x 20% = THB 1.2 million per year
  • Profit effect at the 25% contribution margin: 1.2 x 25% = THB 0.30 million per year (approx. JPY 1.38 million)

McKinsey indicates that AI-enabled supply chain forecasting can reduce lost sales and stockouts by up to 65% in the right applications. That, however, is an upper bound. Here we adopt 20% as a planning figure, less than a third of it.

Benefit 3 – less effort on supply and demand balancing

  • Current state: 3 production control staff x 10 hours per week = 30 hours per week x 50 weeks = 1,500 hours per year
  • A 40% reduction gives 600 hours per year
  • At a fully loaded labour cost of 350 baht per hour, 600 x 350 = THB 0.21 million per year (approx. JPY 0.97 million)

This does not reduce cash out directly. It is a soft benefit in the sense that the freed-up hours can be redirected to other work. Unless you are assuming headcount reduction, this line should be discounted when you present it for approval.

Benefit 4 – fewer expedited shipments and emergency changeovers

  • Baseline annual spend on air freight expedites, express couriers and unplanned changeovers: THB 0.9 million
  • A 25% cut, the same rate as the error reduction: 0.9 x 25% = 0.225, so THB 0.23 million per year (approx. JPY 1.06 million)

Total annual benefit

Benefit lineAnnual amount (THB million)Reference in JPYNature
Inventory holding cost reduction1.06approx. JPY 4.88 millionP&L, recurring
Gross profit recovered from fewer stockouts0.30approx. JPY 1.38 millionP&L, recurring
Reduced supply and demand balancing effort0.21approx. JPY 0.97 millionP&L, recurring, soft
Fewer expedited shipments and changeovers0.23approx. JPY 1.06 millionP&L, recurring
Total gross annual benefit1.80approx. JPY 8.28 million
(Reference) working capital released5.90approx. JPY 27.14 millionOne time only. Not included in the annual benefit

Costs

Initial investment (one-time)

ItemAmount (THB million)Reference in JPY
Data preparation, master deduplication, stockout logging0.80approx. JPY 3.68 million
Model build, PoC, parameter design0.70approx. JPY 3.22 million
Integration with existing systems (production, inventory, purchasing interfaces)0.75approx. JPY 3.45 million
Standing up the operating structure (training, review design, procedures)0.35approx. JPY 1.61 million
Total2.60approx. JPY 11.96 million

Annual running cost (from year one)

ItemAmount (THB million per year)Reference in JPY
SaaS licence and cloud usage0.45approx. JPY 2.07 million
Maintenance and model retraining0.20approx. JPY 0.92 million
Internal data operations effort0.15approx. JPY 0.69 million
Total0.80approx. JPY 3.68 million

What payback looks like

The annual net benefit in steady state is:

Gross annual benefit 1.80 minus annual running cost 0.80 = THB 1.00 million per year (approx. JPY 4.6 million)

Dividing the 2.60 million baht initial investment by that steady-state net benefit gives a simple payback of 2.60 / 1.00 = 2.6 years.

Year one, however, does not deliver the full effect. Data preparation and integration take time, and it takes several months for inventory to settle at its new level. Assuming year one delivers 50% of the gross annual benefit, the cumulative position looks like this.

YearGross benefitCostNet for the yearCumulative
Year 10.903.40 (initial 2.60 plus running 0.80)-2.50-2.50
Year 21.800.80+1.00-1.50
Year 31.800.80+1.00-0.50
Year 41.800.80+1.00+0.50

Units are millions of baht. On a profit and loss basis, the cumulative position turns positive in year 4.

Viewed on a cash basis, assuming the 5.90 million baht of released working capital is realised as 2.95 million baht in each of years one and two:

  • Cumulative year 1: -2.50 + 2.95 = +0.45
  • Cumulative year 2: 0.45 + 1.00 + 2.95 = +4.40
  • Cumulative year 3: 4.40 + 1.00 = +5.40 (thereafter only +1.00 a year is added)

On a cash basis the cumulative position turns positive in the first year. But the inventory release happens once. From year three onward the only addition is 1.00 million baht a year, and this must not be re-read as generating 5.90 million baht of cash every year. That is the single most frequent error in internal investment proposals.

Sensitivity analysis – how the answer changes with the error reduction rate

The most uncertain assumption in the whole model is that error can be cut by 25%. Let us flex it. The dead stock reduction rate and the stockout reduction rate are scaled proportionally so that they equal 40% and 20% respectively at a 25% error reduction.

Error reductionInventory reductionGross annual benefitAnnual running costNet annual benefitSimple payback
10%2.360.840.800.04Effectively unrecoverable
25% (base case)5.901.800.801.002.6 years
40%9.442.750.801.95About 1.3 years

Units are millions of baht.

The conclusion this table points to is unambiguous. If error reduction stops at 10%, the annual running cost of 0.80 million baht eats almost all of the benefit and the investment does not pay back. The investment decision on demand forecasting AI is therefore not a binary of works or does not work. It comes down almost entirely to a single question: what level of error reduction can you realistically expect? Which is exactly why you need to measure FVA during the PoC and find out where your own data lands in that range.

For the broader framework of measuring benefits, our article on measuring the impact and designing the ROI of AI adoption sets out how to choose metrics and how to present them in an internal approval process.

Why you should not apply McKinsey’s 20 to 50% error reduction directly to your own plant

McKinsey indicates that AI-enabled supply chain forecasting can reduce forecast error by approximately 20 to 50% in the right applications. This figure appears in almost every demand forecasting AI proposal. Both the party quoting it and the party being quoted at should check the following three points.

1. It says a 20 to 50% reduction in forecast error, not a 20 to 50% improvement in accuracy.

These are confused constantly. At a plant with MAPE of 28%, a 25% cut in error takes MAPE to 21% (28 x 0.75 = 21). That is a 7 point improvement. Read it as accuracy improves by 25% and interpret that as MAPE of 28% becoming 3%, and your expectations are off by an order of magnitude.

2. The conditional clause, in the right applications, is the substance of the sentence.

As shown in the previous chapter, this effect does not appear for project-driven items, intermittent-demand items, or plants where lead time is shorter than the forecast horizon. The 20 to 50% range is the distribution across cases that met the applicability conditions.

3. It depends on your starting MAPE.

Cutting error by 50% at a plant with MAPE of 50% to reach 25% is a completely different exercise, in difficulty and in money, from cutting error by 20% at a plant already at 15% to reach 12%. The worse your starting point, the easier it is to post a large reduction rate. The better your starting point, the smaller the reduction rate will be.

Implementation case figures likewise need to be read with their context. LIXIL is reported to have deployed AI demand forecasting across roughly 1.2 million building material models and 2.3 million SKUs, achieving reduced transport costs and less manual intervention through better area-level forecast accuracy. But a scale of 2.3 million SKUs is precisely the condition under which machine learning performs at its best. It is unreasonable for a plant with a few hundred SKUs to expect the same rate of improvement.


Five issues specific to sites in Thailand and ASEAN

A model or an evaluation standard built at head office in Japan will not function if you carry it unchanged into a Thai or Vietnamese site. There are five local issues.

Issue 1 – in a year when tariffs structurally changed demand, down-weight the historical data

The 2025 front-loading that TPSO pointed to is poison for a forecasting model. The 2025 shipment record mixes real demand with volume pulled forward in anticipation of tariffs. Train on that 2025 data as it stands and the model memorises the pulled-forward volume as the normal level of real demand, then produces an over-forecast for 2026. And 2026 is the inventory drawdown that follows as the payback. Over-forecasting multiplied by contracting demand is the combination that inflates inventory the most.

Possible remedies include the following.

  • Flag the periods in which front-loading occurred and expose that flag to the model as an external variable
  • Rather than simply weighting recent data more heavily (time decay weighting), cut the training window at the structural break
  • Do not let the model learn the 2025 surges to the US and China (+30.6%, +13.6%) as seasonality

The important point is that a human has to make this judgement and tell the model. The fact that front-loading occurred is not written anywhere in the data.

Issue 2 – put the holiday calendar into your features

For demand forecasting at ASEAN sites, this is in fact the single highest-return improvement available.

  • Songkran (Thai New Year): mid-April. The number of shutdown days and how the surrounding days fall change from year to year. Because companies add their own extra closures, the number of working days in April varies annually.
  • Lunar New Year (Greater China): falls between late January and mid-February and therefore crosses month boundaries depending on the year. It affects both parts procurement from China and Taiwan and demand from Chinese-affiliated customers. Look at a simple year-on-year comparison for January and the comparison does not hold between a year when Lunar New Year fell in January and a year when it fell in February.
  • Tet (Vietnamese Lunar New Year): at Vietnamese sites, operations and logistics swing sharply for several weeks around Tet. On top of that, the wave of resignations and returns after Tet affects capacity on the supply side.

A model built at head office in Japan normally contains only the Japanese calendar. There are many cases where simply adding this produces a visible improvement in monthly forecast MAPE. In implementation terms it costs almost nothing, since it means adding working days, holiday flags and days-before-and-after-holiday counts as features. Put the calendar in before you make the model more sophisticated. That is the correct order.

Issue 3 – baht movements and export ratio

At plants with a high export ratio, part of the movement in demand originates in the exchange rate. If baht strength persists, export competitiveness falls and orders decline. Whether that relationship actually exists for your plant can be tested against historical data.

There is a caveat, however. If you use FX as an explanatory variable, you run into the problem that at the moment of forecasting, the future exchange rate is unknown. What you can use in a forecast is either past FX as an actual, or FX forwards and projections. If you use past FX, you need to verify the lag, typically several months, before FX movements show up in orders, and you can only use it within the span of that lag. Push FX into the model without verifying the lag and the model will pick up meaningless correlations.

Issue 4 – do not reuse the same model across two sites

As we saw at the start, Thai manufacturing in 2026 is in a slowdown, with MPI at -0.8% in May and automotive at -17.94%, while Vietnam is in an expansion, with first-half IIP at +10.8% and PMI at 51.8. Within a single company, demand is pointing in opposite directions.

When that is the case, these three things need to be separated by site.

  1. The model’s training data – do not train Vietnamese demand on Thai actuals
  2. Bias monitoring – a slowdown tends to produce over-forecasting and an expansion tends to produce under-forecasting, so the direction of correction is reversed
  3. Safety stock policy – in an expansion the cost of a stockout is relatively higher, while in a slowdown holding cost and obsolescence risk are relatively higher

There are, on the other hand, things you should standardise. The evaluation standards for forecast accuracy (the definitions and calculation methods for MAPE, bias and FVA), the review cadence and the meeting minutes format, and the data quality criteria. Standardise those and you can compare sites and hold a meaningful discussion about which site is running the better process. Remember it as separate the models, align the yardstick.

It is also worth noting that while Vietnam’s PMI shows 14 consecutive months of production expansion, employment has declined for four consecutive months. Raising output without adding people means the slack in capacity is shrinking. There is less room to absorb an upside surprise through extra production, so on the Vietnamese side it is practical to weight the upside risk of the forecast more heavily than the downside.

Issue 5 – how to treat BOI and digital investment

The Thailand Board of Investment (BOI) includes AI-enabled system development, cloud services and data centre operation among the activities eligible for investment promotion. There is also a tax incentive framework for the cost of developing and training digital talent. Investment applications for the digital industry in the first quarter of 2026 rose to 2.4 times the level of the same period a year earlier.

That said, whether incentives apply is judged by BOI case by case. There is no automatic entitlement simply because something is AI. How your demand forecasting AI deployment is positioned within your business activities and investment plan changes the treatment, so for projects where the investment amount is large, we suggest checking with BOI or a specialist on a case-by-case basis before you finalise the plan. Building that verification period into the schedule at the point where you plan the internal approval process reduces rework later.


A 90-day plan to get started

Demand Forecasting AI ROI for Thai Factories - MAPE and Payback - figure 3

As a general market norm, reaching a PoC takes two to three months, and including full rollout the range is six months to a year. Here we set out the first 90 days, that is, how to assemble the material for a decision through the PoC.

Days 0 to 30 – take stock of your data and fix the evaluation criteria

You do not build a model in this period. What you build is the foundation for a judgement.

  • Extract shipment history and assess its quality: how many years do you have, where are the gaps, when were item codes created or retired
  • Deduplicate the item master: is the same physical item present under multiple codes? If this is broken, everything downstream is meaningless
  • Separate inter-site transfers from real demand: are transfers from plant to warehouse being booked as shipments?
  • Fix the baseline: measure current MAPE (both SKU by week and product group by month), bias, days of inventory, stockout rate and write-off value
  • Build the naive forecast benchmark: compute MAPE for last month’s actual and for the same month last year. This becomes the reference line for FVA
  • Agree the evaluation criteria: put the PoC acceptance conditions in writing at this stage and get stakeholder agreement

The last item matters most. Decide the acceptance conditions afterwards and you will inevitably land on the fudged conclusion that it is not as good as we hoped but it is not unusable either. Write it as a number, for example product group by month MAPE must improve on the naive forecast by at least 5 points.

Days 31 to 60 – run the PoC

  • Narrow the scope: one line or one product group, roughly 30 to 80 SKUs. Do not do this company-wide
  • Backtest: split historical data into a training period and a validation period, and compare forecast against actual in the validation period. Always check how you cut the periods so that information from the validation period does not leak into the training period
  • Measure FVA by stage: measure MAPE after naive, statistical model, machine learning and human adjustment, and make visible where value is created and where it is destroyed
  • Test the effect of external variables: compare with and without the holiday calendar and working day counts
  • Check bias: is the mean error close to zero? Is it skewed toward particular item groups or particular months?

If at this point the result is that the machine learning model cannot beat the statistical model, that is not a failure. It is a valuable finding. In that case the conclusion is that what deserves investment is not a more advanced model but data preparation, or the vessel side, meaning your inventory policy and ordering rules.

Days 61 to 90 – embed it in the process and make the investment decision

Producing forecast numbers on its own achieves nothing yet.

  • Connect it to inventory policy: decide how the improved forecast error will be reflected in the safety stock formula. Without this decision, inventory will not fall by a single baht
  • Redesign the decision flow: who looks at the forecast, and what happens when which threshold is crossed. Make it a concrete rule, such as escalate to the production planning meeting when the forecast moves more than plus or minus 20% month on month
  • Guardrails on human adjustment: define the permitted adjustment band (for example up to plus or minus 15%) and make a written rationale mandatory beyond it
  • Design the integration with existing systems and estimate the effort: specify the interfaces to the production management, inventory and purchasing systems
  • Convert benefits into money and decide: take the error reduction rate you actually measured in the PoC, apply it to the sensitivity framework in this article, and produce a payback period

How difficult days 61 to 90 turn out to be depends heavily on the state of the receiving side in your production management system. For an overview of existing systems, our production management system comparison is worth reading alongside this.

The decision at day 90 has three options. First, proceed to full rollout. Second, change the scope and run another PoC. Third, stop and redirect the investment into building the vessel. Designing the PoC so that the third option remains available is what healthy project governance looks like.


Five common failure modes

Failure 1 – making MAPE the only acceptance condition

Because comparison against a naive forecast (FVA) was never measured, you end up paying 0.80 million baht a year, indefinitely, for a system that is no more accurate than a year-on-year comparison. Build the naive forecast benchmark on your own data before you sign.

Failure 2 – installing the engine with no vessel to receive the forecast

Safety stock still a fixed value from ten years ago, reorder points still manual, production planning still on fixed lot sizes. In that state, however much forecast accuracy improves, neither days of inventory nor stockout rate will move a millimetre. In terms of the model above, of the 1.80 million baht of gross annual benefit, the 1.06 million baht arising from inventory and the 0.23 million baht from reduced expediting, a combined 1.29 million baht or about 72%, disappear entirely. What remains is 0.51 million baht, which is below the 0.80 million baht annual running cost. On top of that, if reorder points remain manual, most of the 0.30 million baht from stockout reduction should also be lost. We have conservatively left it in here, but realistically what actually remains is closer to the 0.21 million baht of reduced balancing effort. Either way, in this configuration the investment runs at a loss.

Failure 3 – letting sales overwrite model output without limits

The experience of being blamed for a stockout stays with people more vividly than the experience of being blamed for excess inventory. As a result, human adjustment carries a structural upward bias. Measure FVA by stage and you can see in the numbers that this overwriting is degrading accuracy. The practical answer is not to ban adjustment but to require a band and a written rationale.

Failure 4 – starting at all SKUs by week

SKU-level MAPE for industrial goods is 20 to 40% as standard. Start at fine granularity and the stakeholders who see the first result will judge the system unusable, and the project stops. Start at product group by month, and only after it works take the class A items down to a finer grain.

Failure 5 – not recording demand during stockouts

Almost every plant has a history of what it sold, and almost none has a history of what it could not sell because there was no stock. Train on historical data in that state and the model structurally under-estimates demand. The low forecast then produces low inventory, which produces another stockout. This negative loop cannot be solved by making the model more sophisticated, however far you take it. Structurally, the only fix is on the data side. Starting to record stockouts, late deliveries and lost orders is the first piece of data preparation to undertake. Note also that it takes time from starting to record until you have accumulated enough volume to train on, so the sooner you begin, the better.


FAQ

How accurate is demand forecasting AI?

It varies widely by industry and granularity, and there is no single universal level. As a guide, at SKU level: 10 to 25% for stable-demand FMCG and consumer staples (deteriorating to 25 to 35% during promotional periods), 35 to 60% for apparel and short-lifecycle items, and 20 to 40% for industrial, B2B and project-driven items. The range generally regarded as acceptable is under 10 to 20%, but that depends on demand variability and lifecycle stage.

Be careful with claims such as accuracy of 95% guaranteed. An accuracy figure that does not state the underlying granularity and forecast horizon is useless for comparison. Also, McKinsey’s figure of a 20 to 50% reduction in forecast error does not mean accuracy improves by 20 to 50%. It means that at a plant with MAPE of 28%, cutting error by 25% takes MAPE to 21%.

How much does a demand forecasting system cost?

In the model case used in this article, a plant with 500 million baht of annual revenue, the initial investment is 2.60 million baht (approx. JPY 11.96 million) and the annual running cost is 0.80 million baht (approx. JPY 3.68 million). This assumes the reference conversion at 1 THB ≈ 4.6 JPY.

The breakdown of the initial investment is data preparation 0.80, model build and PoC 0.70, integration with existing systems 0.75, and standing up the operating structure 0.35. The annual running cost is licences 0.45, maintenance and retraining 0.20, and internal data operations effort 0.15.

One caution: what appears on a vendor quotation is normally only the model build and the licence. Build into your budget from the outset the fact that data preparation and integration with existing systems account for about 60% of the total (1.55 of the 2.60 initial investment).

What is the difference between production planning AI and demand forecasting AI?

Demand forecasting AI outputs how much will sell in the future and the uncertainty around it. It is an input to planning. Production planning AI, the production scheduler, takes required quantities and due dates and decides when to make what and in what sequence, under constraints of equipment capacity, changeover time and labour shifts. It is an output of planning. Between the two sits inventory optimization AI, which sets safety stock, reorder points and order quantities per SKU.

The sequence is demand forecasting, then inventory optimization, then production planning. Install demand forecasting AI alone and, if the inventory policy and production planning downstream stay fixed, your plant KPIs will not move.

Can we implement this with only three years of data?

Three years is 36 data points for a monthly forecast and roughly 156 for a weekly one. Capturing seasonality is generally said to require at least two years of monthly data and preferably three, so in terms of learning seasonality, three years just barely meets the minimum.

More important than volume, however, are quality and structure. Check the following.

  1. Item code changes. If the same product carried a different code three years ago, your history is effectively fragmented
  2. Structural breaks. If a one-off surge such as the 2025 front-loading is included, that period should not be learned as real demand
  3. Whether stockout periods can be identified. Learning a period when nothing sold because there was no stock as zero demand will under-estimate demand

The realistic approach is to start at product group level rather than SKU level. Aggregation increases the information content per data point, so three years works perfectly well. In addition, a machine learning model that learns across SKUs can sometimes fill in for a short individual history using the behaviour of similar SKUs.

Can we implement inventory optimization AI first, on its own?

Yes, and depending on the situation that may well offer a better return.

At a plant where the safety stock formula has been running on a fixed value for years, simply recalculating safety stock to reflect demand variability and supply lead time variability can reduce inventory without changing forecast accuracy at all. That is because the work involved is correcting the imbalance between SKUs that have too much and SKUs that have too little.

The guide for deciding is where the primary cause of stockouts sits. If the main cause is misreading demand, start with demand forecasting. If it is supply-side delivery variability, start with inventory optimization. Which of the two applies can be established in about a week by taking stock of 10 to 20 stockout events from the past year and classifying their causes. That stocktake costs nothing, so whichever system you are considering, it is worth doing first.


Summary

For 2026, the public sector in Thailand (TPSO) sees exports at -3.1 to +1.1%, while the Thai National Shippers’ Council sees +8 to 10%. In a year when the external demand forecasts are split by double-digit percentage points, spending money to buy a forecast that will be right is a poor investment. What deserves the investment is a mechanism that manages how you are wrong.

The key points of this article.

  1. The return on demand forecasting AI is not determined by how many points you improved MAPE. It is determined by two things: whether there is added value over a naive forecast, that is FVA, and whether you have a vessel of planning, inventory and purchasing that receives the forecast and acts on it.
  2. Read accuracy across three layers. MAPE (the size of the error), bias (the direction and persistence of the error) and FVA (the increment over a naive forecast). Look only at MAPE and you will fail to notice a systematic over-read while inventory climbs.
  3. Whether a MAPE is good or bad is determined by industry and granularity. For industrial and B2B at SKU level, 20 to 40% is a standard level. An accuracy comparison that does not state granularity is meaningless.
  4. Costs land in four buckets. Data preparation, model, integration with existing systems, and operating structure. Estimate only the model portion and you will miss more than 70% of the initial cost.
  5. Set exactly one baseline and calculate along a single line of reasoning. The inventory reduction itself is not profit. It is a one-time release of working capital. The only thing you can book annually to the P&L is the holding cost reduction. If you count inventory reduction and stockout reduction together, you need to show the basis for it, namely that inventory shifts from SKUs in surplus to SKUs in shortage.
  6. If error reduction stops at 10%, the running cost eats the benefit and it does not pay back. Which is exactly why measuring FVA during the PoC is the fork in the road for the investment decision.
  7. Do not miss the ASEAN-specific issues. The payback from front-loading, the Songkran, Lunar New Year and Tet calendars, the lag in baht movements, the asymmetry where demand points in opposite directions in Thailand and Vietnam, and case-by-case verification of BOI incentives.
  8. Ninety days is enough to assemble the material for a decision. Days 0 to 30 for the data stocktake and fixing evaluation criteria, days 31 to 60 for the PoC and FVA measurement, days 61 to 90 for embedding it in the process and making the investment decision. Design it so that the option to stop and redirect the investment into the vessel remains open.

Demand forecasting AI is not magic. But the more the external forecasts diverge, the more valuable it becomes to understand quantitatively how your own forecasts miss, and to reflect that in inventory and production planning. A good place to begin is by building the naive forecast benchmark for your own plant. It costs nothing, and it becomes the reference point for every judgement that follows.

TOMAS TECH works from Bangkok on IT and OT integration for Japanese-affiliated plants, and we support the building of production planning, inventory and purchasing systems that actually act on a demand forecast. We are happy to talk at an early stage, including conversations such as we have not decided to implement anything yet but we would like to know whether forecasting is even viable on our data, or we would like a third party to look at whether the accuracy figures in a vendor proposal are reasonable. We can also share examples of how to approach this when operations span both Thai and Vietnamese sites. You can reach us through our contact page.


Sources