“We went live across the whole company last month, and usage is running at 70 percent.” The report has been flowing smoothly right up to that sentence, and then a director asks a single question — “So how much money did we make?” — and the room goes quiet. In meeting rooms at Japanese-affiliated factories around Bangkok, that silence is not unusual. There is a genuine sense that the tool is being used. The shop floor is not complaining. And still, nobody can answer in currency. AI ROI measurement tends to get filed behind tool selection, and it is always the last thing you are asked about. This article organises the whole problem in the order you actually meet it: how to count benefits across four layers, the four patterns that make measurement fail, a full ROI calculation comparing two rollout scenarios, and a 90-day roadmap for designing measurement before you deploy.
Why companies that report real benefits still cannot explain them in money
The numbers around generative AI contain a combination that looks contradictory at first glance. A very large share of companies say they are getting benefits. A very small share can present an ROI figure solid enough to use in an investment decision. The right place to start is with the nature of that gap.
“Delivering business benefits”: 86.7 percent
Teikoku Databank’s survey on corporate trends in generative AI (March 2026) was conducted across 23,349 companies in Japan and collected 10,312 valid responses, a response rate of 44.2 percent. In that survey, 34.5 percent of companies said they are using generative AI. The breakdown is 4.4 percent using it extensively and 30.2 percent using it to some degree. (Sub-totals do not always add exactly to the headline figure because of rounding.)
Among the companies that are using it, the share reporting that it delivers benefits to their operations is 86.7 percent — 25.2 percent saying it delivers substantial benefits and 61.5 percent saying it delivers some benefit. Close to nine in ten perceive a benefit of some kind.
Pause here for a moment. The 86.7 percent figure is a share of subjective assessments. It is not the share of companies that have demonstrated ROI. It answers the question “do you feel this is helping?” It does not answer the question “how much of your investment have you recovered?” When that distinction gets blurred as the number is copied into an internal deck, the pain arrives later, and it always arrives.
For context, the same survey shows a substantial non-adopting side as well. 13.6 percent say they are not using it much and 23.3 percent say they are barely using it — 36.9 percent combined. A further 14.2 percent say they are considering it. The overall picture is a rough balance between adopters and non-adopters, with a considering group waiting in the wings.
Only a small fraction of pilots reach measurable ROI
A contrasting figure comes from research outside Japan. MIT NANDA’s “The GenAI Divide: State of AI in Business” reported that 95 percent of enterprise generative AI pilots have not delivered measurable ROI.
The part that deserves careful reading is the stated cause. The report does not attribute it to inadequate model performance. The causes it identifies are data that has not been prepared, a lack of integration into business processes, and the absence of any definition of success agreed before work began. In other words, the primary cause of failure sits on the design and measurement side, not the technology side.
Summarising this as “95 percent failed” changes the meaning. The precise statement is that they have not reached measurable ROI. It is not that there was no benefit; it is that the organisation is not in a position to say whether there was one. That difference points to a completely different remedy.
The cancellation forecast is over 40 percent
There is one more number with direct bearing on investment decisions. Gartner predicts that more than 40 percent of agentic AI projects will be cancelled by the end of 2027. The reasons cited are escalating costs, unclear business value, and inadequate risk controls.
This one also needs careful phrasing. It is not a past-tense fact that 40 percent failed; it is a forward-looking prediction that more than 40 percent will be cancelled. That said, two of the three cited reasons — escalating costs and unclear business value — are precisely what the absence of effect measurement produces. Nobody has the spend broken out by line item, and nobody can state in numbers what improved. A project in that condition gets cut the moment the business climate or the corporate agenda shifts even slightly.
The gap is a break between the language of experience and the language of finance
Line the three numbers up and the shape becomes visible. The shop floor experiences a benefit (86.7 percent). That experience has not been translated into measurable ROI (95 percent have not got there). And when the translation does not happen for long enough, the project is stopped on the grounds that its business value is unclear (the forecast of over 40 percent cancelled).
In Japanese-affiliated manufacturers, this break tends to run particularly deep. The shop floor speaks in the language of “it got easier”. The management meeting speaks in baht. Between those two you need a staircase. What this article proposes is a way to build that staircase in four steps.

Organising AI ROI measurement into four layers: L1 to L4
Look at the reports produced by companies that are struggling with effect measurement, and most of them are not failing to measure. They are measuring at too shallow a layer. To separate what is being measured and at what depth, use the following four layers.
| Layer | What it measures | Representative indicators | Cost to measure | Weight in a management meeting |
|---|---|---|---|---|
| L1 Usage | Is it being used at all? | Login rate, active rate, adoption rate | Low | Low |
| L2 Time | Did the work get faster? | Time per task, measured before and after | Medium | Medium |
| L3 Business KPI | Did the operational numbers move? | Outsourced translation spend, ticket volume, overtime hours, lead time | Medium to high | High |
| L4 Financial | Did the P&L move? | Actual spend variance by line item, payback period, ROI | High | Highest |
L1 Usage — easy to measure, useless in the boardroom
L1 looks at whether the tool is being used: login rate, monthly active rate, and adoption rate. All of it comes out of the admin console automatically, so the cost of measuring is effectively zero.
That is exactly why so many companies stop reporting here. “Usage is at 70 percent.” “Monthly active users grew again this month.” The statements are factually correct, and this is the layer that carries the least weight in a management meeting. From a director’s chair, the only available response is “and?” Utilisation rate is not an input to an investment decision.
That does not make L1 worthless. The variable that plays a decisive role in the ROI calculation later in this article — the adoption rate — is captured at exactly this layer. L1 is weak as a destination for reporting and indispensable as an input to the calculation.
L2 Time — turning “it got faster” into a number
L2 compares time spent per task before and after deployment. How many minutes did it take to translate a technical document from English into Thai? How many minutes to draft an answer to an internal enquiry? Measure the same way before and after, and take the difference.
Once you reach L2, the reporting becomes concrete, because you can produce a figure like “8 hours per person per month”. Multiply by an hourly rate and you can express it in money.
But L2 has a decisive weakness. Time savings do not appear in the P&L unless headcount falls. If 8 hours a month are freed up per person and nobody has decided what those 8 hours are for, company spending does not change by a single baht. The same salaries are paid. The same outsourcing invoices are paid. This is where more companies get stuck than anywhere else in effect measurement.
L3 Business KPI — connecting to line items where cash actually moves
L3 looks at whether the operational numbers moved. The indicators you choose here have to satisfy one condition: they must be line items where cash actually moves.
The usual candidates are these.
- Outsourcing spend: translation, document preparation or data entry previously sent outside and now handled internally
- Overtime hours: the portion of overtime that is actually paid out and has now stopped being paid
- Recruitment cost and frozen headcount: a planned position you were able to leave unfilled
- Enquiry volume: first-line responses on an internal help desk or customer desk that are now automated, reducing handling effort
- Lead time: faster responses to quotation requests or technical enquiries, feeding through to orders won or inventory held
Of these, the two that tend to bite hardest at a Thai site with a multilingual workforce are translation outsourcing spend and internal enquiry volume. In an organisation where Japanese, English and Thai all circulate, the volume of translation and first-line answering is structurally high. On the enquiry side, automating first-line responses through something like a multilingual chatbot deployment is the standard construction.
The important discipline when designing L3 is to decide which line items you will watch before you deploy. Go looking afterwards and it reads as though you cherry-picked whatever number happened to look good. Declare the line items in advance and the credibility of the report survives whether the result is good or bad.
L4 Financial — speaking in P&L terms
L4 speaks in the language of finance: actual spend variance by line item, payback period, ROI. It carries the most weight in a management meeting and costs the most to produce.
L4 is built on top of L3. How much did actual spend on the line items chosen at L3 fall year on year? Subtract the investment, and derive payback period and ROI. Put the other way round, you cannot skip L3 and construct L4. A deck that estimates financial impact directly from an L1 utilisation rate has assumptions nobody can inspect, and finance and accounting will always read it as a manufactured number.
The real point of the four layers is the step between L2 and L3
If there is one thing to take away from this model, it is the step that sits between L2 and L3.
Moving from L1 to L2 is a matter of effort alone. Put in the measurement work — a stopwatch will do. Moving from L2 to L3, however, requires changing how the work is organised. What will the freed-up hours be spent on? When exactly do you stop sending that work to an outside vendor? How do you change the overtime approval rule? None of these are measurement problems. They are management decisions.
When a company says it “cannot measure the effect”, what is usually true is not that the measurement instruments are missing. It is that these decisions have not been made.
Four patterns that make AI ROI measurement fail
Here are the four failures we see over and over on real projects. It is worth checking which one your own reporting resembles.
Pattern 1 — stopping at L1
This is the most common by a wide margin. The report is complete once it says “usage 70 percent” and “monthly active users up on last month”. There is no bad intent involved. Those are simply the numbers that are easiest to obtain, so those are the numbers that get reported.
What happens at the next budget round is entirely predictable. A utilisation rate cannot be compared with other investment proposals. A new production line or an equipment refresh is argued in money, while AI alone is argued in percentages. An item that never made it onto the comparison table is a natural first candidate for cuts — that is simply how approval processes behave.
The fix is straightforward: build the reporting template from day one with L3 line items already in it. Even if only L1 is populated, the presence of empty L3 fields leaves a shared understanding in the organisation that “next year’s budget depends on filling these in”.
Pattern 2 — never measuring the before
This is the pattern where someone says “let us measure the benefits” three months after go-live. At that point, no measured record of pre-deployment task times exists anywhere.
The conversation then goes like this. “It used to take about 30 minutes each, didn’t it?” “Actually I think it took longer than that.” And that memory-based number becomes the basis of the capital request. From a finance perspective, that is not a basis at all.
The awkward part is that the before can never be recovered retrospectively. The after can be measured at any time; the before can only be measured in the two weeks preceding deployment. The single highest-return activity in the whole of effect measurement design is that two weeks of pre-deployment measurement. A project that skipped it will never have grounds to claim a difference, no matter how carefully everything afterwards is measured.
Pattern 3 — forgetting to multiply by the adoption rate
The third pattern distorts the money more than any other. The estimate is built on the assumption that everyone uses the tool: “60 target users × 8 hours saved per month”.
Let us run the actual arithmetic using the assumptions from Scenario A below. The following is TOMAS TECH’s own estimate, not based on public statistics or third-party surveys.
Without an adoption rate, 60 people × 8 hours = 480 hours per month, or 5,760 hours per year. Multiplied by an hourly rate of 268 THB, that comes to 1,543,680 THB per year. Add the 288,000 THB reduction in outsourced translation spend and the annual benefit appears to be 1,831,680 THB. Against a second-year cost of 1,290,000 THB, that produces a surplus of +541,680 THB. The capital request sails through.
Now assume the actual adoption rate is 60 percent, so 36 people use it. 36 people × 8 hours = 288 hours per month, or 3,456 hours per year. In money that is 926,208 THB per year. Add the same 288,000 THB translation saving and the annual benefit is 1,214,208 THB. Against the same second-year cost of 1,290,000 THB, that becomes a deficit of −75,792 THB.
| Assumption | Annual hours saved | L2 value in money | Total annual benefit | Second-year result |
|---|---|---|---|---|
| No adoption rate applied (all 60 users) | 5,760 hours | 1,543,680 THB | 1,831,680 THB | +541,680 THB |
| Adoption rate 60% (36 users) | 3,456 hours | 926,208 THB | 1,214,208 THB | −75,792 THB |
Whether or not you multiply by the adoption rate is the only thing that decides whether year two shows a surplus or a deficit. The gap in benefit is 617,472 THB. Looking only at the L2 portion, 926,208 THB inflates to 1,543,680 THB — roughly 1.67 times.
And there is not a single lie in the optimistic version. The 8 hours saved per person is the same. The 268 THB hourly rate is the same. The only difference is how many people actually use the thing. That is precisely why the adoption rate is an L1 metric that determines the L4 conclusion. It is the reason you should not dismiss L1 as “the layer that does not fly in a management meeting” — you keep collecting it every month because the calculation depends on it.
Raising the adoption rate is, of course, a lever in its own right. Handing out licences and walking away is not enough; translating usage into the specific language of each job — through something like generative AI training for manufacturing teams — moves the numerator of the ROI directly.
Pattern 4 — converting time savings to labour cost and calling it done
The fourth pattern is stopping at L2 while believing you have reached L4.
“We saved 3,456 hours a year × 268 THB = 926,208 THB of benefit.” As arithmetic, the statement is correct. As a claim, it is incomplete, because that 926,208 THB has not come out of any line item. The same salaries were paid, and the headcount is unchanged. There is nowhere in the accounts where those 926,208 THB appear.
You are entitled to count a time saving as a benefit only under the assumption that the freed-up hours are redeployed onto other valuable work. And to satisfy that assumption, one of the following has to be decided in advance.
- Freeze a position you were planning to fill (recruitment cost and payroll actually fall)
- Bring outsourced work back in-house (outsourcing spend actually falls)
- Fit work that used to require overtime inside normal hours (overtime pay actually falls)
- Process more volume in the time released (revenue or capacity actually rises)
Report a monetary conversion of freed-up hours without having decided any of those four, and management will ask where those hours are going right now. If you cannot answer, you cannot complain when the benefit is judged to be zero.
In fact, if you look at Scenario B below through cash-moving line items only, the benefit from year two onwards is just the 288,000 THB translation saving, which falls short of the 417,000 THB annual cost (288,000 − 417,000 = −129,000 THB). In other words, unless you decide where the freed-up hours go, even a disciplined, tightly scoped deployment fails to pay back. What puts the ROI into surplus is not the capability of the tool. It is the management decision about how the time gets used.
Calculating generative AI ROI for real: Scenario A versus Scenario B
From here we compare in concrete money. Same company, same tool, same hourly rate — the only variable is how widely the deployment is spread.
The following is TOMAS TECH’s own estimate, not based on public statistics or third-party surveys. It is a modelled figure assuming procurement inside Thailand in 2026. Real numbers move with the target processes, the volume of documents, and the state of the existing systems.
Model company assumptions
| Item | Assumption |
|---|---|
| Business | Japanese-affiliated manufacturer in Thailand (400 employees) |
| Target users | White-collar staff (administration, engineering, management), 60 people |
| Monthly working hours | 8 hours × 21 days = 168 hours per month per person |
| Average monthly salary (target group) | 45,000 THB |
| Hourly rate | 45,000 ÷ 168 = 268 THB per hour (salary basis, rounded) |
| Existing outsourced translation spend | 480,000 THB per year (English to Thai and back, technical documents and manuals) |
One note on the hourly rate. This article uses a salary basis: salary divided by working hours, giving 268 THB per hour. There is an equally legitimate convention that uses the fully loaded employer cost including statutory benefits, which typically comes out around 1.2 times higher. Neither convention is wrong. What matters is that the organisation uses one convention consistently. Every calculation in this article is stated on the salary basis of 268 THB per hour. A capital request that mixes salary-basis and employer-cost-basis figures in the same document will come back from accounting.
Scenario A: build a company-wide platform straight away
Licences go to all 60 target users, and a RAG platform capable of searching across internal documents is built in the first year.
Costs
| Line item | Year 1 | Year 2 onwards (per year) |
|---|---|---|
| Licences (60 users × 900 THB/month × 12 months) | 648,000 | 648,000 |
| RAG platform build (year 1 only) | 1,200,000 | 0 |
| Operations, maintenance and improvement | 480,000 | 480,000 |
| Internal programme effort (0.3 person-month/month × 12 = 3.6 person-months × 45,000 THB) | 162,000 | 162,000 |
| Total | 2,490,000 | 1,290,000 |
On how a document search platform is actually assembled, our article on building RAG over factory technical documents covers how to select the target documents and how heavy the preprocessing load becomes. The main reason the RAG line item grows this large is not the model. It is the scope of documents you decide to prepare.
Benefits
- Adoption rate 60 percent, so 60 × 60% = 36 people actually use it
- Hours saved per adopting user: 8 hours per month
- Monthly saving 36 × 8 = 288 hours per month, or 3,456 hours per year
- L2 value in money = 3,456 × 268 = 926,208 THB per year
- L3 reduction in outsourced translation spend (60 percent of 480,000 THB) = 288,000 THB per year
- Total annual benefit = 926,208 + 288,000 = 1,214,208 THB
Result
| Benefit | Cost | Net | |
|---|---|---|---|
| Year 1 | 1,214,208 | 2,490,000 | −1,275,792 |
| Year 2 | 1,214,208 | 1,290,000 | −75,792 |
| Year 3 | 1,214,208 | 1,290,000 | −75,792 |
| Three-year total | 3,642,624 | 5,070,000 | −1,427,376 |
Three-year ROI = −1,427,376 ÷ 5,070,000 = approximately −28.2 percent (it does not pay back)
A first-year loss is expected. What deserves attention is that years two and three continue to lose 75,792 THB annually. The build cost has gone and it still does not turn positive. That is not a case of “the first-year investment was heavy”. It is a structure in which the benefit never reaches the level of the recurring cost. Carry that structure into years four and five and all you accumulate is more loss.

Scenario B: narrow the scope and deploy in stages
Here the scope is limited to two processes — technical document search and translation — and to the 20 people who actually perform them.
Costs
| Line item | Year 1 | Year 2 onwards (per year) |
|---|---|---|
| Licences (20 users × 900 THB/month × 12 months) | 216,000 | 216,000 |
| Minimal RAG build (restricted document set, year 1 only) | 350,000 | 0 |
| Operations, maintenance and improvement | 120,000 | 120,000 |
| Internal programme effort (0.15 person-month/month × 12 = 1.8 person-months × 45,000 THB) | 81,000 | 81,000 |
| Total | 767,000 | 417,000 |
Benefits
- Adoption rate 75 percent (higher because the scope is narrow), so 20 × 75% = 15 people
- Hours saved per adopting user: 10 hours per month (deeper usage because the processes are narrow)
- Monthly saving 15 × 10 = 150 hours per month, or 1,800 hours per year
- L2 value in money = 1,800 × 268 = 482,400 THB per year
- L3 reduction in outsourced translation spend = 288,000 THB per year (translation is inside the target scope)
- Total annual benefit = 482,400 + 288,000 = 770,400 THB
Result
| Benefit | Cost | Net | |
|---|---|---|---|
| Year 1 | 770,400 | 767,000 | +3,400 |
| Year 2 | 770,400 | 417,000 | +353,400 |
| Year 3 | 770,400 | 417,000 | +353,400 |
| Three-year total | 2,311,200 | 1,601,000 | +710,200 |
Three-year ROI = 710,200 ÷ 1,601,000 = approximately +44.4 percent
Year one is already marginally positive at +3,400 THB. From year two, with the build cost gone, +353,400 THB per year remains.
What separated the two scenarios
Put side by side, several things run against intuition.
First, the absolute annual benefit is larger in Scenario A. A delivers 1,214,208 THB against B’s 770,400 THB. Hours saved also favour A: 3,456 hours a year against 1,800. And yet the three-year ROI is −28.2 percent for A and +44.4 percent for B — a spread of 72.6 points. The conclusion is not driven by the absolute size of the benefit but by the ratio of benefit to cost.
Second, the gap in cost is wider than the gap in benefit. Three-year costs are 5,070,000 THB for A and 1,601,000 THB for B, a difference of 3,469,000 THB. The three-year benefit difference is only 3,642,624 − 2,311,200 = 1,331,424 THB. Spreading the deployment does increase benefit, but cost grows faster. That is the arithmetic reason “company-wide from day one” fails so often.
Third, both the adoption rate and the hours saved per person improve when the scope is narrowed. A runs at 60 percent adoption and 8 hours a month; B at 75 percent and 10 hours. Narrow the target processes and the training becomes concrete, the moment of use becomes obvious, and the tool gets used more deeply. Push it out to everyone and a certain proportion of users will always end up in the “I do not know what to use this for” category.
Fourth, the weight of L3 changes. The 288,000 THB translation saving accounts for about 23.7 percent of A’s annual benefit and about 37.4 percent of B’s. B carries a higher proportion of benefit in the part where cash genuinely moves. How easy the case is to explain in a management meeting is proportional to that ratio.
In short, what determines ROI is not the technology but the way the scope is set. And scope can only be set before the work begins.
The 90-day measurement design roadmap

Here is how to turn everything above into action across 90 days. Note that this is not a tool deployment schedule. It is a measurement design schedule.
| Period | Activities | Deliverable |
|---|---|---|
| Day 0–15 | Narrow to three target processes or fewer / measure the before (two weeks of stopwatch measurement of time per task) | Baseline task times |
| Day 16–30 | Inventory the L3 indicators (identify the line items where cash actually moves: outsourcing spend, overtime hours, enquiry volume) | Metric definition document |
| Day 31–60 | Go live with a limited user group / review usage logs and adoption rate weekly | Actual adoption rate |
| Day 61–75 | Measure the after (same method, same processes as the before) | Actual hours saved |
| Day 76–90 | Connect L2 to L3 to L4 and build the financial table / decide to expand, retry or stop | One page for the management meeting |
Day 0–15 — the two weeks of before measurement decide everything
The highest-value part of this roadmap is the first 15 days, for the reason already given: the before can only be measured before deployment.
The work itself is unglamorous. Narrow to three target processes or fewer, and ask the people doing them to record time per task for two weeks. A stopwatch works. So does a slightly modified daily work sheet. What matters is that the measurement method is written down so it can be reproduced exactly during the after measurement.
There are two reasons for capping the target processes at three. One is that the measurement load stays within realistic limits. The other is what Scenario B demonstrated: a narrower scope raises both the adoption rate and the hours saved per person. “Let us try it broadly first and find out which processes benefit” looks like reasonable exploration, but from a measurement perspective it is the worst available choice, because it creates a large number of processes with no before.
Day 16–30 — do the L3 inventory together with accounting
These two weeks are for identifying the line items you will use at L3. Do not let IT complete this alone; bring accounting and general affairs into the room.
The questions are specific. Which account code carries outsourced translation spend, and at what granularity? Can overtime pay be extracted by department? Where do dispatched-worker and subcontract costs sit? If a line item cannot be pulled out of the accounting system on a monthly basis, L4 reporting becomes a manual exercise every time, and manual exercises always stop.
At the same stage, document the decision criteria. Decide in advance which of “expand”, “change the scope and retry” or “stop” you will choose at day 90, and at what threshold. What adoption rate, if you fall below it, triggers a change of scope? What fraction of the expected hours saved, if you miss it, triggers a stop? If you have not decided, the project will continue out of inertia — every time. And projects that continue out of inertia are exactly the ones that get cancelled en masse a few years later under Gartner’s heading of unclear business value.
Day 31–60 — review the adoption rate weekly
Once the limited user group is live, review usage logs and adoption rate weekly. Monthly is too slow, because when adoption is falling, the causes are concentrated in the first two to three weeks.
The common causes sit in operations rather than performance: the login procedure is tedious, internal documents were never added to the search scope, the output language does not match the work, or the manager simply is not using it. Review weekly and you can act. Look at it all at once after 90 days and the habit of not using it has already set.
As Pattern 3 showed, the difference between 60 percent and 100 percent adoption is the difference between a surplus and a deficit in year two. Reviewing adoption weekly is itself an ROI improvement measure.
Day 61–75 — measure the after exactly as you measured the before
There is only one principle to protect during the after measurement: same method, same processes, and ideally the same people as the before.
Cases where the method quietly changes are astonishingly common. The before was measured directly but the after is substituted with a survey. The before covered three processes but the after measures only the one process people found convenient. Once that happens, no difference can be claimed. Preserving comparability is the single greatest requirement of after measurement.
Day 76–90 — connect L2 to L3 to L4 and get it onto one page
In the final 15 days you stack the numbers up, in this order: the L1 adoption rate, then the L2 hours saved, then the L3 actual spend variance by line item, then the L4 financial table and ROI.
One page is enough for the management meeting. The content narrows to five items.
- Target processes and target headcount (what, and who used it)
- Actual adoption rate (how many people actually used it)
- Actual hours saved (the before-and-after difference, with the measurement method stated)
- Actual spend variance on the L3 line items (the money that genuinely fell)
- The three-year cost and benefit table, and the ROI
Then state which of expand, retry or stop you are choosing, judged against the criteria agreed before you started. Keeping “stop” available as an option from the outset is what makes this design credible. A measurement exercise that can only ever conclude “expand” is not measurement. It is a capital request written backwards.
Regulatory and regional factors when designing AI ROI measurement in Thailand and ASEAN
Running this calculation at a site in Thailand or elsewhere in ASEAN introduces several variables you would not face in Japan. Three of them are worth covering.
Thailand’s BOI — what the investment mix tells you
The Thailand Board of Investment (BOI) approved 649 projects worth 330,132 million baht (about USD 10.3 billion) in the first quarter of 2026. Data centre investment accounted for approximately 86 percent of that value. Capital tied to AI and digital is flowing into Thailand in concentrated form.
Closer to the manufacturing floor is the Smart and Sustainable Industry measure, which supports upgrades to existing operations. It received 61 applications worth 7,071 million baht (USD 221 million) in Q1 2026. The BOI’s manufacturing focus for 2026 is described as Industry 4.0, covering smart factories, AI-enabled production and automation.
From an effect measurement perspective there is one implication. The denominator of the investment decision — the cost — changes depending on whether incentives apply. The three-year costs in this article were 5,070,000 THB for Scenario A and 1,601,000 THB for Scenario B. If a scheme reduces that denominator, the ROI conclusion can move with it. The rational sequence, therefore, is to check whether the investment could fall within a support measure for upgrading existing operations before you build the financial table. Eligibility depends on the content of the investment and the application requirements, so confirm the final determination with the BOI and with your own accounting and tax advisors.
Thailand’s AI policy — the legal framework is still developing
Thailand’s National AI Strategy covers 2022 to 2027. It sets a goal of entering the global top 50 on the AI Readiness Index and aims to improve the AI skills of more than 10 million Thai people by 2027.
At the same time, AI-related legislation in Thailand is still under development. It would be incorrect to say that an AI law is already in force in Thailand. This does not affect the design of effect measurement directly, but it does mean that when you write internal policies and guidelines, you cannot justify them on the basis that Thai law requires them. For now, data handling and the response to information leakage risk have to be designed as a matter of your own risk controls rather than legal compliance.
That maps onto the challenges reported in the Teikoku Databank survey. The challenges cited there for generative AI adoption were accuracy of information at 50.4 percent, shortage of specialist talent and know-how at 41.3 percent, deciding which processes to apply it to at 40.0 percent, information leakage risk at 33.5 percent, and establishing internal rules at 25.5 percent. The top three are all matters to be addressed through internal design rather than external regulation.
If you have a Vietnamese site — the AI Law took effect in March 2026
Companies with operations in Vietnam have an additional development to track. Vietnam’s first dedicated AI statute, the Law on Artificial Intelligence (Luật Trí tuệ nhân tạo), was passed by the National Assembly in December 2025 and took effect on 1 March 2026.
It runs to 8 chapters and 35 articles. Its orientation is “management for development”, seeking to balance risk control with the promotion of technological innovation. On the corporate support side, it provides for a national AI development fund, an AI Voucher scheme, and a regulatory sandbox.
If you are rolling a measurement framework designed in Thailand out to a Vietnamese site, country-specific variables enter on both the cost and the benefit side. Measuring in a common format remains valuable, but we recommend stating explicitly in any cross-site report that the regulatory premises differ by country. “ROI was +44.4 percent in Thailand, so Vietnam should be the same” does not hold, because neither the regulatory environment nor the wage level is the same.
Frequently asked questions
How do you calculate generative AI ROI?
The basic approach is to build up costs and benefits separately over a period of around three years and compare them. On the cost side, alongside licences, build cost and operations and maintenance, always include internal programme effort converted into person-months. The model estimate in this article books 162,000 THB per year for Scenario A (0.3 person-month per month) and 81,000 THB per year for Scenario B (0.15 person-month per month). Drop that line item and the ROI looks better than reality.
On the benefit side, record L2 (time savings converted to money) and L3 (line items where cash actually moves) separately. And always multiply by the adoption rate. In Scenario A, three-year benefit is 3,642,624 THB against costs of 5,070,000 THB, a net of −1,427,376 THB and a three-year ROI of approximately −28.2 percent. In the tightly scoped Scenario B, benefit is 2,311,200 THB against costs of 1,601,000 THB, a net of +710,200 THB and a three-year ROI of approximately +44.4 percent. Same company, same tool — and the way the scope is set moves the answer that far.
When should AI ROI measurement start?
Before deployment. The one thing in effect measurement that cannot be recovered is the measured baseline of pre-deployment task times. The after can be measured whenever you like; the before can only be measured in the two weeks before you deploy.
The 90-day roadmap in this article places before measurement in Day 0–15. If you have already deployed and have no before, the practical move is to restart by adding a new target process. Pick one process where generative AI is not yet in use, measure the before there, and expand from that. For processes already in scope, you can fall back on year-on-year comparisons of L3 line items such as outsourcing spend and overtime hours, but the precision is lower.
What should we suspect when the benefits are not showing up?
Suspect four things, in order. First, the adoption rate. If the tool is not being used, no benefit can appear. Check the logs weekly and see what proportion of the expected user count is genuinely active.
Second, the granularity of the target processes. Too broad a scope makes the per-person time saving shallow. In this article’s estimates, the tightly scoped Scenario B improved both the adoption rate (60 percent to 75 percent) and the hours saved per person (8 hours to 10 hours per month).
Third, where the freed-up hours go. If you have decided none of frozen headcount, insourcing, overtime reduction or higher processed volume, then time can be freed up without anything appearing in the P&L.
Fourth, the layer you are measuring at. If all you are looking at is the L1 utilisation rate, it may not be that the benefits are absent — it may be that you are not yet in a position to state them. MIT NANDA’s research likewise found that 95 percent of generative AI pilots have not delivered measurable ROI, and attributed that not to model performance but to unprepared data, a lack of integration into business processes, and the absence of a definition of success agreed before work began.
Is it better to build AI capability in-house or use a vendor’s hands-on support?
This cannot be settled on money alone, but it becomes far easier to judge if you book internal programme effort as a cost on both sides. Building in-house looks cheaper because external payments are smaller, but the effort simply relocates inside the company. This article’s model books internal programme effort at 0.3 person-month per month (162,000 THB per year) for Scenario A and 0.15 person-month per month (81,000 THB per year) for Scenario B. Treat that effort as free and in-house delivery looks unfairly attractive.
As a practical rule of thumb, the workable split is to use external support for measurement design and the initial launch, and to keep the weekly adoption review and monthly reporting in-house. The Teikoku Databank survey found 41.3 percent of companies citing a shortage of specialist talent and know-how as a challenge, which suggests this is where in-house delivery tends to become the bottleneck. Conversely, monitoring adoption and translating usage into the specific language of each job are jobs that people who know the work do faster and more reliably, and they are poor candidates for outsourcing.
Can Thailand’s BOI schemes be used for AI investment?
It depends on the content of the investment. The BOI approved 649 projects worth 330,132 million baht (about USD 10.3 billion) in Q1 2026, with data centre investment accounting for approximately 86 percent of the value. The Smart and Sustainable Industry measure, which supports upgrades to existing manufacturing operations, received 61 applications worth 7,071 million baht (USD 221 million) in the same quarter. The BOI’s manufacturing focus for 2026 is described as Industry 4.0, covering smart factories, AI-enabled production and automation.
That said, software subscription fees such as generative AI licences are not automatically in scope. Before you finalise an investment plan, confirm the potentially eligible scope with the BOI and with your own accounting and tax advisors. The sequence matters, because checking after the contract is signed sometimes leaves you unable to go back.
How many measurement indicators should we use?
Three target processes or fewer, and two to three L3 line items, is the realistic setting. Adding indicators inflates the effort of measurement itself, which raises the internal programme effort cost. That enlarges the denominator, which makes ROI worse.
There is one criterion for selecting an indicator: can the number be extracted automatically from the accounting or operational system on a monthly basis? Any indicator that requires somebody to compile it by hand each month will stop by the third month. An indicator that has stopped is the same as an indicator that never existed.
Does the choice of generative AI tool affect measurement?
It does. At selection time, the thing to check is less the features themselves and more whether the data you need for measurement can be extracted. Specifically: can per-user, per-month usage be retrieved from the admin console, can it be exported as CSV or similar, and can it be aggregated by department? Choose a tool where that data is not available and you lose the ability to measure adoption at L1, which drops you automatically into the Pattern 3 trap.
Check the fit with the target processes in advance as well. If, like the model company in this article, the core need is English-to-Thai technical document translation, then multilingual quality and the range of internal documents the tool can reference become the evaluation axes. That consideration is tied directly to the design of what goes into the search scope, so it is efficient to think about it alongside the deployment model.
Summary
Here are the key points of this article on AI ROI measurement.
- The 86.7 percent “delivering benefits” figure is a subjective assessment, not ROI. In the Teikoku Databank survey, 86.7 percent of companies using generative AI reported business benefits, while MIT NANDA’s research found 95 percent of generative AI pilots have not reached measurable ROI. That gap leads to the future Gartner predicts, in which more than 40 percent of agentic AI projects are cancelled by the end of 2027.
- Measure the benefit in four layers. L1 usage, L2 time, L3 business KPI, L4 financial. Most companies stop reporting at L1. Time savings at L2 do not appear in the P&L unless headcount falls, so they only become L4 once connected to L3 line items where cash actually moves.
- There are four failure patterns. Stopping at L1, never measuring the before, forgetting to multiply by the adoption rate, and converting time savings to labour cost and calling it done. The adoption rate in particular is the single variable that swings year two between a surplus (+541,680 THB) and a deficit (−75,792 THB).
- ROI is determined by how the scope is set. In our own estimates, the company-wide Scenario A returns approximately −28.2 percent over three years while the tightly scoped Scenario B returns approximately +44.4 percent — even though A produces the larger absolute annual benefit.
- Design the measurement in 90 days. Measure the before in Day 0–15, document the L3 line items and decision criteria in Day 16–30, review adoption weekly in Day 31–60, measure the after by the same method in Day 61–75, and assemble the financial table in Day 76–90 to decide whether to expand, retry or stop.
- In Thailand and ASEAN, regulation changes the denominator. BOI support for upgrading existing operations, Thailand’s still-developing AI legislation, and Vietnam’s Law on Artificial Intelligence which took effect on 1 March 2026. When rolling a framework across borders, state the differing regulatory premises explicitly.
Being unable to answer “so how much money did we make?” is, in most cases, neither a tool problem nor a failure of effort on the shop floor. It is simply that the measurement was not designed before the deployment. The design takes 90 days, and the most important part of it is the first 15.
Talk to us
If you are wrestling with effect measurement or ROI design for generative AI, get in touch through the TOMAS TECH contact form. There is no need to commission a tool deployment. “We have already deployed and just need help structuring the one page for the management meeting”, “we started without measuring the before and want to know how to reset”, or “we would like to talk through how to narrow the target processes” are all perfectly good starting points. Working from Bangkok with Japanese-affiliated manufacturers, we can help you organise the picture around your specific situation.
References
- Teikoku Databank, Survey on Corporate Trends Concerning Generative AI (March 2026)
- Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027
- MIT: 95% of enterprise AI pilots fail to deliver measurable ROI
- Thailand’s BOI investment applications surge past 1 trillion baht in Q1 2026
- Thailand BOI — Manufacturing
- Thailand AI Policy News
- 6 laws taking effect from 1 March 2026 (Vietnamese)
- Vietnam’s Law on Artificial Intelligence 2026 (Vietnamese)