Blog

2026.07.31

Enterprise generative AI comparison – payback 11.2 vs 3.4 years

Enterprise generative AI comparison - payback 11.2 vs 3.4 years

Nobody is still debating whether to use generative AI. The conversations we are asked into now sit squarely in the evaluation phase of an enterprise generative AI comparison — which product, for which people, and how many seats. Yet the projects that open with a feature matrix are the ones that stall at the final approval step. Using a Japanese-affiliated manufacturer in Thailand as the model (300 employees, 60 of them white-collar), this article shows how the payback period splits into 11.2 years and 3.4 years on identical benefit assumptions, with every formula on the page.

Why starting from a product comparison table fails – shadow AI at 68% and the wrong subject

Walk into an evaluation meeting and the first thing on the table is almost always a landscape-format matrix. Products across the columns, features down the rows, circles and triangles filling the grid. For the effort that goes into building it, that table contributes remarkably little to the actual decision. The reason is simple. The subject of the table is the product.

What the buying side actually has to decide is not which product is superior. It is which of your own people get which kind of tool, and how many seats of it. The subject is your staff and your work, not the software. A beautifully precise table about the wrong subject still produces no answer, so the meeting ends with “let us gather a bit more information.” Six months later the same table reappears with newer version numbers.

Adoption went from 33% to 65% in two years, so waiting is no longer a position

Start with the state of the market. Enterprise adoption of generative AI rose from 33% in 2023 to 65% in 2025. McKinsey, which ran the research, positions this as the fastest diffusion the firm has measured for any technology it tracks. Nearly doubling in two years is a curve you almost never see in the world of internal business systems.

Narrow it to manufacturing and AI adoption sits at 68% — a different measure from the shadow AI figure of 68% discussed below — with AI-related spending up 48% year on year, driven mainly by predictive maintenance and quality control. Manufacturing, in other words, is already firing live rounds at AI in general, not just at generative AI.

The lesson from those two numbers is not “hurry up.” It is that the phase in which you could differentiate simply by adopting is already over. Once a majority has it, having it is not an advantage. From here the difference comes from how you allocate it and how you get people to use it.

The more practical reason the matrix cannot answer the question – people are already using it

There is a second reason the matrix fails to describe reality. While you are evaluating, your staff are already using these tools.

  • Generative AI use the company does not know about, commonly called shadow AI, occurs in 68% of companies.
  • 58% of employees use publicly available AI tools rather than tools their company has approved.
  • The top category is code generation tools, at a 72% usage rate.

For most companies, therefore, an evaluation is not “what shall we introduce from zero.” It is “how do we bring usage that is already happening, unmanaged, under management.” That shift in framing changes the design of the comparison itself. A greenfield rollout might justify picking the single most capable product. Absorbing existing usage argues for picking whatever feels closest to what people are already using, because migration is faster. The evaluation criteria are not the same.

Risk perception is also misaligned inside the organisation

Push a little further and you find the distribution of risk perception is skewed as well.

  • 45% of manufacturing employees believe there is little or no risk in using shadow AI.
  • Meanwhile, 43% of large enterprises have no framework for managing AI risk.
  • And the average cost of a shadow-AI-related data breach is USD 4.2 million.

The shop floor thinks it is not a big deal, management has no operating model for control, and in that state an incident of that order of magnitude is possible. Put those three facts together and the conclusion is that policy and tool deployment have to move together. A policy with no usable tools behind it will not be followed. Tools with no boundaries drawn around them leave shadow AI in place. This is not a sequencing question but a structural one — either half on its own does not work.

Three things to settle before you compare anything

Following from that, there are exactly three things to settle internally before you enter a product comparison.

What to decideWhat happens if you compare without deciding it
Who gets a seat (scope and headcount)The quote comes back as “everyone, to be safe” and unit price times headcount blows up the total
What they will do with it (use-case types)You try to satisfy every use case with one product, and only the top-tier plan survives as a candidate
What may be entered into it (the information boundary)Legal and IT stop it after purchase, and you sit on licences nobody uses

With those three settled, the candidate list narrows naturally. Skip them and the product with the most features always looks like the safest choice, so you end up buying capability you do not need. Cost inflation comes from seat count and tier design, not from unit price. That, in one line, is the argument of this article. If you also want to revisit the overall rollout process, we cover that separately in our guide to running a generative AI implementation.

What the three major products actually cost – the price increase and the add-on trap

Enterprise generative AI comparison - payback 11.2 vs 3.4 years - figure 1

Now to pricing — but not to the question of which one is cheapest. This is about how far the published price sits from what you actually pay, and how that gap differs by product. Get this wrong and the figure in your approval request will not match the invoice.

Microsoft 365 Copilot – the USD 30 is an add-on price

Microsoft 365 Copilot lists at USD 30 per user per month on an annual commitment. That much is widely known. The complication comes next.

That USD 30 is an add-on price, and it assumes you already hold a qualifying Microsoft 365 licence. To use Copilot you are paying the underlying suite plus USD 30. A company that has already deployed Microsoft 365 company-wide only adds USD 30. A company that has not must stack the foundation first.

On top of that, Microsoft raised the price of its core Microsoft 365 suites on 1 July 2026. Reflecting that increase, the effective all-in figure is USD 69 per seat per month for E3 (with Teams) and USD 90 per seat per month for E5. Note that this total does not include Copilot Studio agent usage charges. If your plan involves building out agents, budget above these numbers.

There is one more change that hits negotiation directly. The volume discount ladder that used to exist — 15% at 10 seats and above (USD 25.50), 20% at 100 seats and above (USD 24), 30% at 300 seats and above (USD 21) and 40% at 1,000 seats and above (USD 18)ended on 30 June 2026. Only certain time-limited commitments remain.

What does that mean in practice? It means the assumption that consolidating seats lowers the unit price no longer holds. Previously there was a point to pushing for 100 seats to capture the 20% break. Now adding seats does not move the unit price. If the unit price will not come down, the only remaining lever on the total is redesigning the combination of seat count and per-tier price structure. The model company in this article has 60 seats, a size that would have qualified for the 15% band at 10 seats and above had the ladder survived. Because the ladder itself ended on 30 June 2026, we calculate a 60-seat estate at the list price of USD 30.

The product strengths are clear. It runs inside Word, Excel, Outlook and Teams. It reaches SharePoint, mail and internal documents through Microsoft Graph. And it carries data loss prevention and oversharing controls in the product itself, which genuinely reduces the effort of governance design. For people whose working day is centred on editing Office files, that “it happens inside the app” experience is hard to replicate with anything else.

ChatGPT Enterprise and Business – price transparency and low training cost

ChatGPT Enterprise does not publish a list price. It is quoted by sales. So when it appears in a side-by-side table, the only honest entry is “quote required.” Fill that cell with a guess and the basis of your approval request collapses.

ChatGPT Business, by contrast, is published at USD 25 per seat per month. For a mid-sized rollout, this is the one that gets evaluated first in practice.

The biggest strength of this product line is that it has the highest general consumer recognition, which keeps the training cost low. That gets dismissed as a soft factor, but at 60 users it matters. If many of your staff have already touched it privately, training can start from “how do we use this at work” rather than “what is this.” Turn the earlier statistic around — 58% of employees are already using unapproved tools — and deploying the product they already recognise is the fastest way to recover shadow AI usage into the sanctioned estate.

Two facts on data handling are worth fixing in writing.

  • Data residency is offered in Europe, the UK, the US, Canada, Japan, Korea, Singapore, India, Australia and the UAE. For a Southeast Asian entity, the presence of Singapore is a practical input to the decision.
  • Data from ChatGPT business plans and from the API is not used to train models unless the customer explicitly opts in.

The second point is the single most misunderstood item in internal discussions. A vague fear that “anything you put into AI gets learned” lingers as resistance long after the rollout. Simply writing the fact that treatment differs by contract plan into the policy document removes a large part of that resistance.

Claude Enterprise – seat fee and usage fee are separate

Claude Enterprise is USD 20 per seat per month on an annual commitment. The important part follows: model usage is billed separately at API rates. Judging it “cheapest” from the seat fee alone will diverge from the invoice.

It is more accurate to read this as a difference in design philosophy than as a drawback. Because you pay in proportion to what you use, holding seats that are used occasionally does not inflate fixed cost. Conversely, heavy users accumulate usage charges. The wider the variance in usage across your organisation, the better this structure works for you.

For technical staff there is a Claude Team Premium seat that includes Claude Code, repriced on annual commitment from USD 150 to USD 100. As noted above, code generation tools are the number one category of shadow AI use, at 72%. If you have a development or engineering function, giving that group a sanctioned channel does more real work than adding another line to the policy.

The strengths are long context length and ease of connecting your own systems and documents. For work that involves reading large volumes of internal documentation, or building something on top of internal data, those two points are what count. For the mechanics of making internal documents searchable and citable, see our article on building factory knowledge retrieval with RAG.

Benchmarks will not tell you which one is strongest

A word on performance, written carefully.

Claude Opus 4.6 records the highest scores on coding and specialist-domain benchmarks, and holds a slight lead on knowledge-work benchmarks as well. As a general pattern, frontier models in the GPT and Claude families tend to outperform the default model behind Copilot.

But the ranking changes depending on the task. “This one is always strongest” is not a conclusion the current benchmark picture supports. And from an operational standpoint there is something more important. Whether the tool runs inside your Office files, and whether it can reach your internal documents, are axes entirely separate from raw model quality. Being one or two points higher on a benchmark matters less to your accounting team than being usable right there in Excel.

If you build a comparison table, put the performance column in as reference information and keep it off the main decision axis.

Data location – what a Southeast Asian entity needs to know

With operations in Southeast Asia, where the data physically sits is unavoidable.

Azure offers default single-region data residency in its Southeast Asia region (Singapore). In addition, Microsoft opened Malaysia and Indonesia regions in 2025 and plans to add India and Taiwan regions in 2026.

What this movement means is that the constraint of “we cannot use it because the data cannot stay in region” is shrinking year by year. A few years ago this was a major fork in the design. The options keep expanding. Put the other way round, “we are deferring this because of data location” carries less weight as an argument than it used to.

The three products on one page – how the prices differ in kind

Here is the summary in a form you can take into a decision meeting. The absence of a performance row is deliberate, for the reason given above — the ranking changes by task.

ItemMicrosoft 365 CopilotChatGPT Business and EnterpriseClaude Enterprise
Published priceUSD 30 per user per monthBusiness USD 25 per seat per month, Enterprise by sales quoteUSD 20 per seat per month
Contract formAnnual commitmentDepends on planAnnual commitment
Pricing preconditionsAdd-on price. Assumes a qualifying Microsoft 365 licenceOnly Business is published. Enterprise list price is not disclosedModel usage billed separately at API rates on top of the seat fee
Effective all-in guideE3 (with Teams) USD 69, E5 USD 90, excluding Copilot Studio agent usageNot applicableSeat fee plus consumption, dependent on volume
Volume discountLadder discounts ended 30 June 2026, only certain time-limited commitments remainNot applicableNot applicable
Main strengthsRuns inside Word, Excel, Outlook and Teams. Reaches internal documents via Microsoft Graph. DLP and oversharing controls in the productHighest general recognition, so training cost is low. Data residency includes SingaporeLong context length. Easy to connect your own systems and documents
Watch out forCore suites repriced on 1 July 2026. The ladder ended, so even a 60-seat estate pays listEnterprise pricing cannot enter a comparison table without a quoteThe seat fee alone does not tell you the total
Additional noteNot applicableBusiness plan and API data are not used for training unless you explicitly opt inTeam Premium seats including Claude Code repriced from USD 150 to USD 100 on annual commitment

What we want you to notice in that table is that the three prices are not expressed in a form that can be compared on the same footing. One is an add-on, one is undisclosed, one is a seat fee plus consumption. Lining up the unit prices and sorting them cheapest first is a meaningless exercise. What is meaningful is asking which price structure fits which group of your own people. That is the next section.

Splitting usage into three tiers – who gets what

We now put the subject of the comparison back where it belongs, from product to people. This is the core of the article.

Why “the same thing for everyone” costs more

Giving every user the same single tool looks fair, easy to administer and simple to negotiate. Most requests for quotation arrive in exactly that shape. Sixty white-collar users times a unit price times twelve months, end of story.

Two assumptions are buried in that approach. The first is that everyone uses it the same way. The second is that consolidating seats lowers the unit price. As we saw in the previous section, the second one collapsed when the ladder discounts ended on 30 June 2026. With the ladder gone, even a 60-seat estate pays list price.

The first assumption was never true to begin with. What an accountant wants to do inside Excel, what a buyer wants when reading an English-language contract, and what a Thai line leader wants when checking an internal rule in their own language are completely different pieces of work. Give different work the same tool and it will be excessive for some people and insufficient for others. Excessive seats go unused. Under-served groups go back to shadow AI.

The three-tier model – cut by the nature of the work

So we cut into three, by the nature of the work. The key is to cut by what the work actually consists of, not by the org chart.

TierHeadcountWhat they getMonthly unit priceMonthly subtotal
Tier A – work centred on Office files (management and administration)15 peopleMicrosoft 365 CopilotUSD 30USD 450
Tier B – work centred on research, drafting and translation30 peopleGeneral-purpose chat (ChatGPT Business class)USD 25USD 750
Tier C – shop floor and multilingual enquiries15 peopleInternal chatbot and RAG (API consumption billing)Not applicableUSD 120
Total60 peopleUSD 1,320

(450 + 750 + 120 = USD 1,320 per month. Headcount is 15 + 30 + 15 = 60 people.)

Tier A (15 people) – the group for whom staying inside the app is the value

Management and administration, accounting, general affairs and part of production control. The group that spends most of its working hours inside Word, Excel, Outlook and Teams.

Give this group a general-purpose chat tool and you will almost certainly get copy-paste purgatory. Select figures in Excel, paste them into a chat window, paste the answer back. That round trip alone erodes a large share of the ten hours per month we are assuming as savings. For this tier, running inside the app is itself a feature — and that is what justifies the USD 30 unit price.

This is also the group with the strongest requirement to reference internal documents and mail. Being able to reach SharePoint and mail through Microsoft Graph is, in practice, a difference nothing else substitutes for.

We hold the headcount to 15 because we are counting only the people whose core work genuinely happens inside Office. Somebody who opens Excel a few times a month is not in this tier. Count generously here and it flows straight through to the total.

Tier B (30 people) – the largest group, and the one that needs the most generality

Sales, purchasing, quality assurance, engineering and HR. Research, drafting and translation are the core. The work starts in a browser and the output is prose.

What this group needs is not Office integration but raw language capability and responsiveness. Documents that cross Japanese, English and Thai, mail with overseas suppliers, working through standards documents. General-purpose chat is at its most effective here, and this is also the tier where deploying a product with high general recognition brings training cost down.

We place 30 of the 60 people here, half the total, because at a Japanese-affiliated subsidiary in the region this is genuinely the thickest layer. And the unit price you use to fill this tier moves the total more than anything else does.

Tier C (15 people) – the group that gets no seat at all

This is the crux of the three-tier model. Tier C gets no seats.

Line supervisors, production section leaders, quality inspectors. What they need is not an environment for free-form conversation with an AI. It is the ability to ask internal rules, work instructions and past defect-handling records a question in their own language and get the right answer back. When the use case is narrow, an internal chatbot with RAG on API consumption billing is cheaper than handing out seats, and the accuracy of the answers is easier to control.

We budget USD 120 per month for it in this model. Handing 15 people a per-seat licence would clearly cost more. And even if you did give them a free-conversation environment, actual usage in this tier tends to converge on “search the internal rules” anyway. When the use case is narrow, a narrow tool is cheaper and faster.

This is also, incidentally, the tier most exposed to the language quality gap, which we come to in a later section.

A side effect of tiering – policy and training become easier to design

The benefits of splitting into three are not only financial.

  • The policy gets easier to write. Instead of one list of prohibitions for everybody, you can vary “what may be entered” by tier. Tier C only ever touches internal documents, so the whole question of input boundaries shrinks for them.
  • Training gets easier to design. Tier A learns how to use it inside Excel, tier B learns patterns for research and drafting, tier C learns how to ask. Because the content differs, each session is shorter. That sticks better than pushing the same two-hour session at everybody. For how to construct these sessions, see our piece on designing generative AI training for manufacturers.
  • You can measure results by tier. It gets you out of the company-wide average utilisation rate, a metric no decision has ever been made on.

The model – why payback splits into 11.2 years and 3.4 years at a 60-person Japanese subsidiary in Thailand

Enterprise generative AI comparison - payback 11.2 vs 3.4 years - figure 2

Numbers from here. Everything below is a model calculation based on the assumptions of this article. Real figures move with each company’s contract terms, existing licence holdings and the nature of the work. This is not a promise that your outcome will match. Read the structure of where the difference arises rather than the figures themselves.

Assumptions

  • Model company – a Japanese-affiliated manufacturer in Thailand, 300 employees, of whom 60 are white-collar.
  • The exchange rate is assumed at USD 1 = THB 32 for the purposes of this article. It differs from prevailing market rates.
  • Average white-collar monthly salary of THB 45,000, with 176 working hours per month.
  • Hourly rate = 45,000 divided by 176 = 255.68 and so on, rounded to approximately THB 256 per hour (we have rounded).
  • Because the Microsoft 365 Copilot discount ladder ended on 30 June 2026, a 60-seat estate is calculated at the list price of USD 30.

Pattern A – the same thing for everyone (all 60 seats on Microsoft 365 Copilot)

The most common approach. The arithmetic is extremely simple.

  • Licences – 60 seats x USD 30 x 12 months = USD 21,600 per year = THB 691,200 per year
  • Annual maintenance and improvement – THB 150,000 per year
  • Total annual cost = 691,200 + 150,000 = THB 841,200 per year

THB 841,200 a year. Whether that is expensive or cheap can only be judged once you place it next to the benefit figure below.

Pattern B – split into three tiers by use case

Calculated on the three-tier allocation from the previous section (USD 1,320 per month).

  • Licences and equivalents – USD 1,320 x 12 = USD 15,840 per year = THB 506,880 per year
  • Annual maintenance and improvement – THB 150,000 per year (the same as pattern A)
  • Total annual cost = 506,880 + 150,000 = THB 656,880 per year

The gap against pattern A is 691,200 minus 506,880 = a reduction of THB 184,320 per year. Looking at licence cost alone, that is a 26.7% reduction (from USD 1,800 to USD 1,320 per month).

What we want to underline is that none of this reduction comes from negotiating a discount. Every unit price is still the list price. The only thing being changed is how the tools are allocated. As we saw, the ladder discounts are gone and the room to negotiate unit price down is narrower than it was. Seat design is the lever that is left.

Up-front cost – cut this and adoption does not stick

The item people miss when they judge on licence cost alone is the up-front investment. And projects that cut here almost never achieve durable adoption.

LayerContentsUp-front cost
1. Platform and connectivityTenant configuration, SSO, DLP policy, conditional accessTHB 120,000
2. Business fitInternal data connection, minimal RAG build, 30 prompt templatesTHB 350,000
3. GovernanceUsage policy, PDPA and AI act checks, log retention designTHB 180,000
4. Adoption and trainingTraining in four languages (Japanese, English, Thai, Vietnamese) plus 90 days of hands-on supportTHB 250,000
Up-front totalTHB 900,000

(120,000 + 350,000 + 180,000 + 250,000 = THB 900,000. This up-front cost is treated as identical for patterns A and B.)

The largest of the four layers is business fit at THB 350,000 — connecting internal data, a minimal RAG build, and 30 prompt templates. This is the money that prevents “we bought it and nobody used it,” and it is the kind of spend where cutting it wastes the entire licence budget. Licences are billed whether or not they are used. Templates only generate value when they are used.

Note also that the THB 250,000 adoption and training layer covers four languages. A Japanese-affiliated entity in Thailand mixes Japanese expatriates, Thai staff, Vietnamese staff and English speakers. Japanese-only training material will not reach tier C at all.

The THB 180,000 governance layer covers the policy and regulatory work discussed in the sections below. Many companies conclude this can wait. Recall that the average cost of a shadow-AI-related breach is USD 4.2 million and the comparison of magnitudes speaks for itself.

Benefits – hours saved, and the rate at which they convert to cash

Now the benefit side. This is the number to set most conservatively, so we open every assumption.

  • Expected saving – 10 hours per person per month. This is a deliberately conservative placement. Vendor material carries much larger numbers, but averaged across all work, this is roughly where it lands.
  • 60 people x 10 hours = 600 hours per month
  • 600 hours x THB 256 = THB 153,600 per month = THB 1,843,200 per year

You must not, however, write THB 1,843,200 into an approval request as the benefit. That is the theoretical value of the hours saved, and it is not cash.

So we apply a cash conversion rate of 50%. The reason is straightforward. Free time does not automatically become money. Unless headcount falls, saved hours only convert to cash through one of three routes.

  1. Overtime reduction – actual spending on overtime pay goes down.
  2. Outsourcing reduction – external spend on translation, document preparation and research falls.
  3. Opportunity capture – freed hours go into work you previously could not get to.

If the saved hours connect to none of these, they simply dissipate inside the organisation. The 50% coefficient is our allowance for that dissipation.

Which gives an effective benefit of THB 921,600 per year.

Payback period – this is the conclusion

We divide the up-front cost of THB 900,000 by the annual net gain (benefit minus annual cost). The benefit figure is held identical for both patterns at THB 921,600 per year. In other words, the difference arises purely on the cost side.

Pattern A (everyone on Copilot)

  • Annual net gain = 921,600 minus 841,200 = THB 80,400 per year
  • Payback period = 900,000 divided by 80,400 = approximately 11.2 years

Pattern B (split into three tiers)

  • Annual net gain = 921,600 minus 656,880 = THB 264,720 per year
  • Payback period = 900,000 divided by 264,720 = approximately 3.4 years
ItemPattern A (same thing for everyone)Pattern B (three tiers)
Licences and equivalents (year)THB 691,200THB 506,880
Maintenance and improvement (year)THB 150,000THB 150,000
Total annual costTHB 841,200THB 656,880
Effective benefit (year)THB 921,600THB 921,600
Annual net gainTHB 80,400THB 264,720
Up-front costTHB 900,000THB 900,000
Payback periodApproximately 11.2 yearsApproximately 3.4 years

What this contrast means

Internally, 11.2 years is effectively a rejection. At 3.4 years you are inside the investment criteria of most companies. Same benefit, same up-front cost, same maintenance cost, same list prices. The only difference is who received what.

Why does it open up so far? The answer is the smallness of the denominator. Pattern A’s annual net gain of THB 80,400 is a razor-thin margin against an annual cost of THB 841,200. Because benefit and cost sit so close together, a small movement in cost swings the payback period enormously. That structure is exactly why THB 184,320 of cost reduction moved the payback from 11.2 years to 3.4 years.

The same structure works in reverse. Place the benefit assumptions slightly more optimistically and even pattern A starts to look dramatically better. Inflate hours saved per person, raise the cash conversion rate, and you can draw a chart in which any allocation works. Which is precisely why you should pin the benefit assumptions down conservatively and compete on the cost design. That is why this article never moved off ten hours per month and a 50% conversion rate.

For how to design the measurement metrics and how to actually observe the hours saved, see our article on measuring the effect and ROI of an AI rollout. Agreeing on the measurement method before you put a model into an approval request is time well spent.

How to read this model

Finally, the limits of the model, stated openly.

  • It does not account for your existing Microsoft 365 licence holdings. If you hold none, the effective all-in figures of USD 69 for E3 (with Teams) and USD 90 for E5 come into play, and the tier A unit price changes substantially.
  • It does not include Copilot Studio agent usage charges.
  • Tier C’s API consumption billing varies with volume. USD 120 per month is an assumption of this article.
  • The exchange rate of USD 1 = THB 32 is an assumption of this article and differs from prevailing market rates.

None of these invalidate the model. They mark the places where you should substitute your own numbers. Once the structure is clear, substitution is not difficult.

Set expectations one notch lower for Thai and Vietnamese – what SEA-HELM shows

This is the point that deserves careful handling precisely because you are operating in Southeast Asia. It is also the point most often missing from rollout plans written at a Japanese head office.

The facts, with the scope of each metric kept separate

A caution first. Each of the figures below comes from a different metric. Reading them side by side as though larger means better is a mistake. Read them separately, with the scope stated.

  • On the SEA-HELM Thai-language leaderboard, Qwen 3 VL 32B leads at 59.73, with Qwen 3 Next 80B MoE second at 58.09.
  • On the SEA-HELM six-capability average, SIAMGPT-32B is highest at 63.59, ahead of Typhoon2.5 and OTG-R1.
  • On the all-language average, Gemma-SEA-LION-v3-9B-IT scores 69.35.

To repeat, 59.73, 63.59 and 69.35 are separate numbers with different evaluation scopes. You cannot read them as “69.35 is the highest, so that one is the strongest.”

The gap between languages – Thai is structurally disadvantaged

With that said, there is one fact that bears directly on operational decisions. The Thai score band sits below Indonesian and Vietnamese. Thai sits around 60, while Indonesian and Vietnamese are in the upper 60s.

Two reasons are given.

  1. The tonal complexity of Thai.
  2. The small size of publicly available Thai training corpora.

The second is the more consequential. It suggests this is not the kind of problem that resolves itself as model generations advance. A shortage of training data is not guaranteed to be solved by the passage of time.

The operational implication – vary the allocation by language

The design principle that follows is unambiguous. Assume Thai output quality is one notch below Japanese and English, and design accordingly.

And do not tell your Thai staff they will get the same accuracy as in Japanese. That is a matter of honesty, and it is equally a matter of adoption. People who start with inflated expectations stop using the tool at the first failure. Tell them up front that it will produce a draft but not something you can send as is, and the failure lands inside expectations.

The allocation splits like this.

LanguageIntended usageOperating rule
Japanese and EnglishMore autonomous. Output may be connected directly to work in defined situationsThe normal review process is sufficient
ThaiAssume draft plus human proofreadingAnything leaving the company must pass proofreading by Thai staff

This principle is consistent with the three-tier model above. Tier C handles shop floor and multilingual enquiries, so it is precisely the tier most exposed to Thai output quality. That is exactly why tier C receives an internal chatbot with RAG grounded in internal documents rather than a free-conversation seat. When the basis of an answer is pinned to internal documents, the surface area that depends on free generation by the language model shrinks. You compensate for the accuracy shortfall on the architecture side.

Split the training material by language too

The THB 250,000 allocated to four-language training connects directly to this point. Translating the same deck is not enough. People working in Thai need to be told explicitly that proofreading is a precondition, and people working in Japanese do not need that content at all. Expectations and operating rules differ by language — aligning that at the start avoids a great deal of confusion later.

When regulation removes options – Thailand’s draft AI act and Vietnam’s AI Law

Legal counsel sometimes enters halfway through a product selection and the options change. That is avoidable rework, so here is the position as of 31 July 2026. This field moves quickly, so make your actual decisions on current information and with professional confirmation.

Thailand – the draft AI act, not yet in force

Start with the premise. Thailand’s AI act is still at draft stage and is not in force. Explaining it internally without that caveat creates unnecessary alarm.

On 2 July 2026, ETDA (the Electronic Transactions Development Agency, under the Ministry of Digital Economy and Society, MDES) published a revised draft of the AI act and opened a public hearing, reported as running for around 30 days.

The draft is structured in three risk-based tiers.

CategoryContents
1. Prohibited AISystems performing subliminal manipulation or unjustified discrimination
2. High-risk AISystems affecting national security, health, the environment, telecommunications, transport or public utilities
3. Designated AI systemsSystems that may become subject to notification, registration or licensing under future announcements

The obligations placed on the deployer, meaning the party using a high-risk AI system, are listed as follows.

  • Operating a risk management framework
  • Complying with the provider’s instructions
  • Appointing a competent supervisor
  • Retaining operational logs for at least six months
  • Notifying the authorities of unanticipated risks
  • Maintaining a mechanism for human oversight

There are also transparency and labelling obligations. Publishing AI-generated content on high-impact subjects such as elections, impersonation or food safety requires disclosure of AI involvement. Developers are required to embed a machine-readable AI generation marking.

Extraterritorial reach deserves attention as well. If conduct affects people in Thailand, the rules apply even where the conduct occurs outside Thailand. Foreign providers are required to appoint a representative in Thailand. If a system built by a Japanese head office is used at the Thai entity, this point needs to be worked through.

Sanctions are administrative fines of THB 1 million to THB 5 million per violation, with service suspension orders and blocking orders through ISPs also possible.

Operationally, the most important point is the relationship with PDPA. The AI act and PDPA do not operate independently. An organisation that processes personal data with AI must comply with both at the same time. The argument that “the AI act is only a draft, so we need do nothing” does not survive the fact that PDPA already exists.

How to think about ordinary office use

Taking a step back, 60 white-collar staff using generative AI for drafting and translation does not fall squarely within the definition of high-risk AI above. It is a different kind of thing from a system affecting national security or public utilities.

That said, the deployer obligations listed for high-risk use double as an excellent checklist for good operational design. Name a supervisor, retain logs for a defined period, keep a human in the final review — these are, before they are legal requirements, preparations that let you explain what happened when something goes wrong. Given that 43% of large enterprises have no framework for managing AI risk, designing to the level of the draft is a cheap form of insurance.

Vietnam – the AI Law, already in force

Unlike Thailand, Vietnam’s AI Law is already in force. Keep the distinction clear.

The National Assembly passed the Law on Artificial Intelligence on 10 December 2025, and it has been in force since 1 March 2026. It applies to both domestic and foreign operators and covers research, development, provision, deployment and use.

There are transitional provisions. Providers and deployers of existing AI systems may continue operating as before until 1 March 2027 (until 1 September 2027 in the health, education and finance sectors), unless the authorities identify a serious risk.

The implementing decree, Decree No. 142/2026/ND-CP, was promulgated on 30 April 2026 and took effect on 1 May 2026.

On obligations, providers must apply machine-readable markings to AI-generated audio, images and video. Deployers must indicate whether that content risks creating a false impression about the truth of events or persons.

Sanctions include suspension of system operation, recall and administrative fines, with serious violations carrying up to 2% of annual revenue. Designing a penalty as a proportion of revenue lands harder the larger the company.

The practical conclusion when you put the two countries side by side

PointThailandVietnam
Current statusThe AI act is a draft (revised version published 2 July 2026, public hearing of around 30 days)The AI Law is in force (in force since 1 March 2026)
Implementing decreeNot applicableDecree No. 142/2026/ND-CP (promulgated 30 April 2026, effective 1 May 2026)
Transitional provisionsNot applicableExisting systems until 1 March 2027 (health, education and finance until 1 September 2027)
Extraterritorial reachYes. Applies to conduct outside the country where people inside it are affected. Foreign providers must appoint a local representativeApplies to both domestic and foreign operators
Labelling and markingDisclosure of AI involvement on high-impact subjects. Developers embed a machine-readable AI generation markingProviders mark audio, images and video in machine-readable form. Deployers indicate the risk of a false impression
SanctionsAdministrative fines of THB 1 million to THB 5 million per violation, service suspension orders, blocking orders through ISPsSuspension of operation, recall, administrative fines. Serious violations up to 2% of annual revenue

For a company with sites in both Thailand and Vietnam, the practical conclusion is to operate one rule set aligned to the stricter regime. Splitting policy by country breaks down in operation. Since Vietnam is already in force, aligning the group policy there is the natural choice. We also cover the Thai situation in our article on AI adoption in Thailand.

Seven items your generative AI usage policy must cover

Now we turn all of the above into an internal policy. Thick policies do not get read. At 60 people, a few A4 pages that everyone will actually read is more effective. The seven items below are limited to what follows directly from the facts in this article.

1. The list of approved tools, and how unapproved tools are handled

The starting point is to name the tools people are allowed to use, explicitly. Policies that open with prohibitions do not get followed. Given that shadow AI occurs in 68% of companies and 58% of employees use unapproved tools, the first job of the policy is to show the sanctioned channel.

Then add one line describing how to request approval for something not on the list. Without a route, people simply use it quietly.

2. What information may and may not be entered

Input containing personal data falls under PDPA. As covered above, an organisation processing personal data with AI must comply with both the AI act and PDPA at the same time.

Writing “do not enter confidential information” in the abstract does not work. Write it as specific document types — customer lists, employee personal data, unreleased drawings, pricing information. Note that the boundary may differ by tier. Tier C only touches internal documents, so the degree of freedom on input is small to begin with and their section of the policy can be short.

3. How input data is treated with respect to model training

Data from ChatGPT business plans and from the API is not used to train models unless the customer explicitly opts in. State this fact in the policy, and record that your company has not opted in.

Why put it in writing? Because it is the single most misunderstood point internally. The vague fear that “anything you put into AI gets learned” persists until it is contradicted on paper. And while it persists, practical usage does not spread.

4. Where the data is stored (data residency)

State which region the data sits in. ChatGPT data residency is offered in Europe, the UK, the US, Canada, Japan, Korea, Singapore, India, Australia and the UAE. Azure offers default single-region data residency in its Southeast Asia region (Singapore), with Malaysia and Indonesia regions opened in 2025 and India and Taiwan regions planned for 2026.

One or two lines recording which you chose and why makes responding to audits and head office enquiries dramatically easier.

5. Responsibility for verifying output, and human oversight

Put in writing that final responsibility for AI output rests with the human who used it. This follows the thinking behind the human oversight mechanism and the appointment of a competent supervisor that Thailand’s draft AI act lists as deployer obligations for high-risk AI.

As covered above, Thai output quality should be assumed to be one notch below Japanese and English. So it is practical to vary the strength of the verification requirement by language. Write it down to the level of specificity of “any external document produced in Thai must pass human proofreading.”

6. Labelling and disclosure of AI-generated material

Thailand’s draft AI act requires disclosure of AI involvement when publishing AI-generated content on high-impact subjects such as elections, impersonation and food safety. Under Vietnam’s AI Law, deployers must indicate whether the content risks creating a false impression about the truth of events or persons.

Day-to-day manufacturing work rarely hits these cases, but it is a live issue in public relations, marketing and recruitment. Add one line to the effect that use of AI-generated images, audio or video in externally published material requires prior sign-off from a named function.

7. Log retention and appointment of a supervisor

Thailand’s draft AI act requires deployers of high-risk AI to retain operational logs for at least six months. Ordinary office use does not automatically qualify as high risk, but the cost of building to that standard from the outset is small — that is the practical judgement.

And decide by name who reads the logs and who makes the call. With 43% of large enterprises having no framework for managing AI risk, what separates companies is not the elegance of the policy text but whether the responsible person actually exists.

Of these seven, items 1, 2, 3 and 4 can be settled during the initial rollout, while items 5, 6 and 7 are the kind that improve as you operate. Try to perfect all seven before deploying anything and you will never deploy. Think of the THB 180,000 governance layer in the up-front cost as the money that gets these seven items into their first working form.

A 90-day rollout roadmap

Enterprise generative AI comparison - payback 11.2 vs 3.4 years - figure 3

Here is everything above, laid out on a timeline. Read it as an answer to the question of how to distribute the four up-front layers (platform and connectivity, business fit, governance, adoption and training, totalling THB 900,000) across 90 days.

Phase 1 (days 1 to 30) – deciding, tiers and boundaries

What you do in these 30 days is not product selection. It is deciding who gets what.

  • Finalise the tiering. Assign real named employees to tiers A, B and C. Our model uses 15, 30 and 15, but the split varies by company. The actual work is judging, one person at a time, whether their core work genuinely happens inside Office.
  • Inventory existing licences. How many people hold Microsoft 365, and on which plans. Since Copilot is an add-on price, you cannot fix the effective tier A unit price without this.
  • Draw the boundary on what information may be entered (policy item 2). Bring legal and IT in from the start.
  • Map the reality of shadow AI. Who is using what today. Since it occurs in 68% of companies, assume it exists at yours and go and look. Whatever you find becomes direct evidence for the tiering.
  • Obtain vendor quotes. ChatGPT Enterprise does not publish a list price and is quoted by sales, so start early if you want it in the comparison.

What starts in this phase is the platform and connectivity layer (THB 120,000) and the first half of the governance layer (THB 180,000) — tenant configuration, SSO, DLP policy, conditional access. Until that is done you cannot hand out licences in the next phase.

Phase 2 (days 31 to 60) – building, fitting it to the work

Turn the decisions into something usable. The most expensive layer, business fit at THB 350,000, is concentrated here.

  • Internal data connection and a minimal RAG build. The tier C internal chatbot takes shape in this phase. Scope it to internal rules, work instructions, past defect-handling records and similar, starting with the documents that attract the most questions. Not trying to load everything is what “minimal” means.
  • Thirty prompt templates. This is the operational crux. Deploy without templates and usage stops at “I tried it once.” Vary the content by tier — spreadsheets and reports for tier A, research, translation and mail for tier B, question patterns for tier C.
  • Begin distributing licences. Tiers A and B. Rather than a company-wide switch-on, starting with a handful of pilot users from each tier gets the template improvement loop turning faster.
  • Finalise the wording of the usage policy (the seven items). Log retention design is settled here too.

Phase 3 (days 61 to 90) – embedding, delivered in four languages

The adoption and training layer (THB 250,000) runs in this phase.

  • Four-language training (Japanese, English, Thai, Vietnamese). Not a translation of one deck, but content with different expectations and operating rules per language. For Thai in particular, state explicitly that draft plus human proofreading is the assumption.
  • Tier-specific training. Tier A on using it inside Office, tier B on patterns for research and drafting, tier C on how to ask. Because the content differs, each is short.
  • Ninety days of hands-on support. End at the training session and usage falls the moment there is nobody left to ask. That is what the support period is for.
  • First measurement of results. Check whether the assumed ten hours per month per person actually holds at your company. If the measured figure is far below, the cause is almost always the templates and the tiering. Fix those before you change product.

What to watch after day 90

Ninety days is not the end. Three metrics are worth watching afterwards, and only three.

MetricWhy watch itWhat to do when it looks bad
Utilisation by tierThe company-wide average cannot support a decision. Skew such as tier A low or tier B high is what carries meaningReduce seats in the unused tier at the next renewal, or move them to another tier
Residual shadow AIIf unapproved usage persists after sanctioned tools are deployed, the policy is not workingAdd a sanctioned channel that covers the remaining use case
Cash conversion of hours savedFree time does not become cash unless it connects to overtime reduction, outsourcing reduction or opportunity captureHave line managers explicitly designate the conversion route

The third is the hardest and the most important. We assumed a 50% cash conversion rate precisely because this does not happen on its own. If nobody decides where the freed hours go, the benefit stays where it started, in the model.

Frequently asked questions

Which enterprise generative AI is actually cheapest?

The practical answer is that unit price alone does not produce an answer.

The published unit prices are USD 30 per user per month for Microsoft 365 Copilot, USD 25 per seat per month for ChatGPT Business and USD 20 per seat per month for Claude Enterprise. Rank those three alone and the order is obvious, but the preconditions differ. Copilot’s USD 30 is an add-on price that assumes a qualifying Microsoft 365 licence (after the price increase, the effective all-in figures are USD 69 for E3 with Teams and USD 90 for E5). Claude’s USD 20 is a seat fee, with model usage billed separately at API rates. ChatGPT Enterprise has no published list price and is quoted by sales.

And what moves the total most is not unit price but seat count. In this article’s model, with no discount on any unit price, allocation alone produces a reduction of THB 184,320 per year, a 26.7% cut in licence cost. Finding the seats you do not need to buy beats finding the cheapest product.

Is Microsoft Copilot worth deploying at a company that does not use Microsoft 365?

The cost structure changes fundamentally. Copilot’s USD 30 is an add-on price that assumes a qualifying Microsoft 365 licence. Without one you have to stack the underlying suite, and reflecting the 1 July 2026 price increase, the effective all-in figures are USD 69 per seat per month for E3 (with Teams) and USD 90 per seat per month for E5 (excluding Copilot Studio agent usage).

Copilot’s essential value lies in running inside Word, Excel, Outlook and Teams and in reaching SharePoint, mail and internal documents through Microsoft Graph. An organisation that does not use those applications realises very little of that value. If you are not on Microsoft 365, general-purpose chat will give you better value for money — that is what the structure implies.

Note also that the volume discount ladder ended on 30 June 2026 (only certain time-limited commitments remain). Any model built on “it gets cheaper if we buy in bulk” needs updating.

If we use ChatGPT for work, is the information we enter used for training?

Data from ChatGPT business plans and from the API is not used to train models unless the customer explicitly opts in.

The important part is that treatment differs by plan. Discussing free consumer use and business plans as though they were the same thing derails internal debate. And given that 58% of employees use publicly available AI tools their company has not approved, the concentration of risk is not in the tool you contracted for. It is in the tools individuals are using on their own.

As countermeasures, we recommend stating in the policy that your company has not opted in (item 3 of the seven above) and recording your data residency choice (item 4). ChatGPT data residency is offered in Europe, the UK, the US, Canada, Japan, Korea, Singapore, India, Australia and the UAE.

What is the minimum a generative AI usage policy should contain?

This article organises it into seven items. (1) The list of approved tools and how unapproved tools are handled, (2) what information may and may not be entered, (3) how input data is treated with respect to model training, (4) where the data is stored, (5) responsibility for verifying output and human oversight, (6) labelling and disclosure of AI-generated material, and (7) log retention and appointment of a supervisor.

Two pieces of advice. First, open with the list of approved tools rather than the prohibitions. If people cannot tell what they are allowed to use, they will use things quietly. Second, do not make it thick. At 60 people, a few A4 pages that everyone actually reads is more effective.

Note also that Thailand’s draft AI act and PDPA do not operate independently. An organisation processing personal data with AI must comply with both at the same time. You cannot defer the handling of personal data on the grounds that the AI act is still a draft.

We hear accuracy in Thai is low. Is there any point deploying to Thai staff?

There is, provided you separate the expectations and the usage pattern from Japanese and English.

As a matter of fact, on the SEA-HELM Thai-language leaderboard Qwen 3 VL 32B leads at 59.73 and Qwen 3 Next 80B MoE is second at 58.09. And the Thai score band sits below Indonesian and Vietnamese (Thai around 60, Indonesian and Vietnamese in the upper 60s). The reasons given are the tonal complexity of Thai and the small size of publicly available Thai training corpora.

So the design principle is an allocation of draft plus human proofreading for Thai, and more autonomous use for English and Japanese. And do not tell your Thai staff they will get the same accuracy as in Japanese. Start with inflated expectations and people stop at the first failure.

Placing an internal chatbot with RAG in tier C rather than free-conversation seats is also a response to this point. Pin the basis of an answer to internal documents and the surface area that depends on free generation by the language model shrinks. The accuracy shortfall can be compensated for on the architecture side.

Do we already have to comply with Thailand’s AI act?

As of 31 July 2026, Thailand’s AI act is still a draft and is not in force. On 2 July 2026 ETDA (the Electronic Transactions Development Agency, under the Ministry of Digital Economy and Society, MDES) published a revised draft and opened a public hearing of around 30 days.

That does not, however, mean there is nothing to do. Two reasons.

The first is PDPA. The AI act and PDPA do not operate independently, and an organisation processing personal data with AI must comply with both at the same time. PDPA already exists, so as long as you handle personal data, action is required now.

The second is that the deployer obligations the draft sets out are, in themselves, good operational design. Appointing a competent supervisor, retaining operational logs for at least six months and maintaining a mechanism for human oversight are obligations aimed at high-risk AI, but even in ordinary office use they are effective ways of getting to a state where you can explain an incident. With 43% of large enterprises having no framework for managing AI risk, there is value in building it early.

Note that Vietnam’s AI Law is already in force (passed 10 December 2025, in force since 1 March 2026). If you have sites in both Thailand and Vietnam, operating one policy aligned to the stricter regime is the practical route.

Summary

This has run long, so here is the part you can act on.

  • Put the subject of the comparison back from product to people. The decision is not which is superior but who gets what. The three things to settle are the scope and headcount of recipients, the types of use case, and the boundary on what may be entered.
  • The room to negotiate unit price has narrowed. The Microsoft 365 Copilot volume discount ladder ended on 30 June 2026, and the core suites were repriced on 1 July 2026. If the unit price will not come down, the only lever on the total is seat design.
  • Published prices are not on the same footing. Copilot’s USD 30 is an add-on price, ChatGPT Enterprise has no published list price and is quoted by sales, and Claude Enterprise’s USD 20 is a seat fee with model usage billed separately. Sorting unit prices cheapest first is a meaningless exercise.
  • Split into three tiers. Tier A of 15 people on Copilot for work that stays inside Office, tier B of 30 people on general-purpose chat for research and drafting, tier C of 15 people on an internal chatbot with RAG and no seats at all. That comes to USD 1,320 per month.
  • Hold the benefit identical and the payback still splits. With an effective benefit fixed at THB 921,600 per year for both patterns and an up-front cost fixed at THB 900,000, pattern A (everyone on Copilot) comes to approximately 11.2 years and pattern B (three tiers) to approximately 3.4 years. The difference arises purely on the cost side.
  • Keep the benefit assumptions conservative and fixed. Ten hours per month, 50% cash conversion. Freed time does not become cash unless it connects to overtime reduction, outsourcing reduction or opportunity capture.
  • Design Thai with expectations one notch lower. The Thai score band sits below Indonesian and Vietnamese, for reasons of tonal complexity and small training corpora. Thai gets draft plus human proofreading, English and Japanese get more autonomy.
  • On regulation, Thailand is draft and Vietnam is in force. With sites in both, run one policy aligned to the stricter regime. And Thailand’s draft AI act and PDPA apply at the same time.
  • Seven policy items, a few A4 pages. Thick policies do not get read, and a policy nobody reads does not stop shadow AI.

One last time. What moved the payback from 11.2 years to 3.4 years was not a discount and not a higher-performing product. It was splitting 60 people into 15, 30 and 15. It follows that the way you allocate your own evaluation time should probably change to match.

Simply working through your headcount and how you would cut the tiers changes how you read the quotes vendors put in front of you. The right numbers differ by company depending on your existing Microsoft 365 licence holdings and how much shadow AI is happening today, so an early stage where you only want to rerun the model on your own conditions is a perfectly good place to start. For the practical side of rolling out and writing policy in Thailand and Vietnam, feel free to reach us through our contact page.

References