The proof of concept worked, and then it never reached production. That is the most common ending for an AI agent project. Research from Anaconda and Forrester puts it plainly: 88% of agent PoCs never make it into production. The cause is not the model you picked or the platform you bought. It is that four things — authority, data scope, monitoring and ownership — were never decided before anyone started experimenting. This article is written for Japanese-owned manufacturing operations in Thailand, and works through the four decisions to make before you start, what implementation actually costs, and a 90-day sequence for getting to a real decision.
Why AI agent implementation stalls at the PoC stage
Somebody in the company tries it. It appears on a management meeting agenda. Vendors start sending proposals. Yet only a small minority of companies ever reach live operation. Before anything else, it is worth looking at the size of that gap in numbers.
88% never reach production
According to the Anaconda and Forrester research, 88% of agent PoCs never reach production. Similar levels have been reproduced in the a16z and MIT Sloan CIO panels. Gartner puts it slightly higher still, at 89% stalling at the PoC stage. Different researchers, different samples, and yet all of them converge on the same conclusion: close to nine out of ten never go live.
The same research also identifies what blocks the move to production.
| Blocking factor | Share reporting it | What it means |
|---|---|---|
| Inadequate evaluation criteria | 64% | There is no definition of “it worked”, so nobody can decide to move to production |
| Governance friction | 57% | Authority, data usage and accountability are unsettled, so approval never comes |
| Model reliability | 51% | Output variability does not meet the requirements of the work |
What deserves attention is that the top two have nothing to do with model performance. Inadequate evaluation criteria and governance friction are both consequences of decisions not taken, not of technology falling short. Even the third item, model reliability, reduces to the evaluation criteria problem once you accept that “what level counts as a pass?” is a question you were supposed to answer in advance. In other words, the 88% figure does not exist because the models are immature. It exists because organisations started experimenting without deciding what had to be decided before experimenting. That is the backbone of this article.
78% have a PoC, only 14% reach company-wide operation
Research published by Digital Applied in March 2026, based on 650 enterprise technology leaders surveyed in February and March, makes the structure more concrete.
- 78% of companies hold at least one PoC
- but only 14% have reached production operation at company-wide scale
- 64% of the companies holding a PoC have attempted to expand it and have been stuck for six months or more
- the average PoC duration before that stall is 4.7 months
The distance between 78% and 14% — a 64-point gap — is the scaling gap itself. The important nuance is that the 64% are described as stalled, not as failed. Something that works was built. But the moment anyone tries to widen it, all the questions the PoC never had to answer arrive at once.
The 4.7-month average is telling in its own right. Roughly half a year of running the PoC, then the start of an expansion attempt, then six months or more of no movement. Add it together and close to a year has passed. Throughout that period, the salaries and licence fees already committed generate no return.
Production rates differ by nearly threefold across industries
The same research also breaks the production rate down by industry.
| Industry | Reaching production |
|---|---|
| Financial services | 21% (highest) |
| Healthcare | 8% (lowest) |
Intuition says heavily regulated industries should move slowest. Yet financial services, one of the most heavily regulated industries there is, has the highest rate of reaching production.
Here is how to read that inversion. Heavily regulated industries already have a procedure for deciding who approves a new mechanism, what gets recorded, and how far a machine is allowed to act on its own. Audit response, access control and log retention already exist as business processes, so when something new like an agent arrives, it can be placed inside a framework that is already there. In an organisation without those procedures, every question has to be argued from zero. There is not even a list of what needs to be decided, so the discussion sprawls and approval never lands. Governance is not an obstacle to implementation. It is the mechanism that makes implementation faster.
The average sunk cost of a failure is about USD 2.1 million
Failed agent projects at Fortune 1000 companies are reported to carry an average sunk cost of about USD 2.1 million. That is obviously a large-enterprise figure, and it does not transfer directly to a subsidiary in Thailand. But break it into components and the same line items appear regardless of scale.
- Licence and API charges (billing continues throughout the PoC period)
- Fees paid to external vendors
- Internal labour cost (the hours put in by project members)
- Integration development (the work of connecting to existing systems)
- Data preparation (organising the documents the agent is meant to learn from or refer to)
Of these, the last two may survive as assets even if the PoC is cancelled. Making an existing system able to expose its data through an API, and structuring internal documentation, both have value whether or not you ever deploy an agent. Which also means the bulk of the sunk cost sits in licences and labour, and those grow monotonically with time. Leaving a project stalled for six months is far more expensive than it looks. The only way to break that pattern is to set the decision criteria before the PoC begins.
Starting from “let’s see if it works” guarantees a stall
Most projects begin with “let’s try it and see whether it is usable”. It looks like a healthy starting point. It has four structural defects.
First, there is no pass mark. Experiment without a definition of “usable” and you get one of two outcomes: satisfaction the moment a demo runs, or a permanent refrain of “the accuracy still is not good enough”. The first collapses in production. The second never reaches production at all.
Second, the authority question gets deferred. In a PoC the person running it uses their own account, so nothing about permissions surfaces. Only when the move to production begins does the question appear — under whose authority does this thing read the core system? — and that is where it stops.
Third, the data scope stays vague while the footprint expands. In production you have to decide which folders the agent may read, and whether customer drawings are included. That line is a management decision rather than a technical one, and management decisions take time to clear.
Fourth, nobody has decided whose mechanism it is. Who fixes it when the output goes wrong? Who updates it when the work changes? Leave that blank and degradation starts on day one of live operation.
Those four map exactly onto authority, data scope, monitoring and ownership, which the later sections work through in detail. The places where the 88% stop are almost entirely these four.
62% experimenting, 23% scaling: where Japan and Thailand actually stand
The previous section was general survey data. From here we look at where Japanese companies and Thai operations actually sit, in numbers. Use it as a ruler for measuring your own position.
McKinsey: 62% experimenting, 23% scaling
In McKinsey’s “The State of AI” (November 2025 edition), 62% of organisations are at the experimentation stage with AI agents, while only 23% have reached scaled operation of at least one workflow company-wide.
That pairing of 62% and 23% is structurally the same as the 78% and 14% in the previous section. The samples and definitions differ, so they cannot be compared directly, but the shape is consistent across multiple studies: experimenting is the majority position, running it in production is not.
Worth noting is how loose the 23% threshold is — it counts if even one workflow has scaled. Which means that as of 2026, running an agent reliably in a single workflow already places you in the top fifth. Read the other way: there is no need to aim at a company-wide rollout from the start. Getting one workflow to run dependably is, in itself, a competitive position.
Japan’s business usage rate for generative AI is 55.2%
Japan’s Ministry of Internal Affairs and Communications publishes an annual White Paper on Information and Communications, which includes an international comparison of how widely generative AI is used at work. The 2025 edition reports the following.
| Country | Business usage rate for generative AI |
|---|---|
| China | 95.8% |
| United States | 90.6% |
| Germany | 90.3% |
| Japan | 55.2% |
Japan alone sits nearly 40 points lower. At minimum, that tells you that in Japanese companies, using generative AI at work is not yet the default assumption.
This matters for anyone whose head office is in Japan. Agents sit one layer above generative AI. Dropping an agent that decides and acts autonomously into an organisation where even human-directed generative AI use has not taken hold is a large step up. The realistic sequence is to establish generative AI use in a specific workflow first, and then hand over to an agent the part of that workflow where a person is making the same judgement every time.
23% of Japanese companies say results greatly exceeded expectations
PwC Japan’s Spring 2025 survey on generative AI found that 23% of Japanese companies reported results that greatly exceeded expectations.
That number does not mean results are absent. The question is where the other 77% fall, and that depends heavily on the expectations set at the time of implementation. Set the expectation at “an entire job will be automated” and you will almost certainly fall short. Set it at “the time this specific task takes drops by 30%” and it lands inside achievable range. Setting expectations is not a matter of attitude — it is the design of the evaluation criteria. The 64% figure for inadequate evaluation criteria in the previous section connects straight to this point.
AI usage policies exist at 50% of large companies and 30% of SMEs
In Japan, the rate at which companies have established an AI usage policy is about 50% for large companies and about 30% for SMEs.
So even among large companies, half are using AI with no policy in place. Nothing is documented about what may be entered, how far the output can be trusted, or who is accountable. While use is confined to individuals typing into a chat window, that rarely produces a serious incident.
Agents are different. An agent reads from and writes to systems in places where no human is watching. Put one into production without a policy and, when something goes wrong, there is no way to reconstruct who authorised what. The 57% governance-friction figure is what that absence looks like when it finally surfaces, right at the point of moving to production.
For an operation in Thailand there are additional questions. Does the head office policy, written in Japanese, apply to the Thai entity? Is it available in a language local staff can read? How does it sit alongside Thailand’s Personal Data Protection Act (PDPA)? A site whose policy exists only in Japanese is not unusual — and for local staff, that is functionally the same as having no policy at all.
A USD 7.8 billion market in 2026, embedded in 40% of enterprise applications
The supply side is worth tracking too.
- The AI agent market reaches USD 7.8 billion in 2026, up 50% year on year
- Gartner forecasts that 40% of enterprise applications will embed AI agents by the end of 2026
The second point is the one that matters operationally. Agents are shifting from something you choose to adopt into something that arrives inside the business applications you already run. Production control, accounting, groupware, CRM — an update enables an agent feature, and there it is. Which means the design of authority and data scope can no longer be deferred. Even if you build nothing yourself, an agent embedded in an existing application will be reading your internal data.
For the state of agent technology itself and how far it is spreading into manufacturing applications, see AI agent trends 2026. This article is confined to what you have to decide when you actually implement, so it does not repeat the technology landscape.
The gap specific to operations in Thailand: 10.7% of organisations, 72% of employees
This is one of the core sections of the article. Thailand has a structural pattern that differs from both Japan and the West. Import a head office implementation plan without understanding it and you will either go nowhere or end up somewhere unmanageable.
The inversion: 10.7% of organisations, 72% of individual employees
Line up the Thai numbers and a clear fault line appears.
| Indicator | Figure |
|---|---|
| AI adoption rate among organisations in Thailand | 10.7% |
| AI usage at work among individual employees in Thailand | 72% |
| AI adoption rate among organisations in Singapore | 69% |
| AI adoption rate among organisations in Vietnam | 23% |
Organisational adoption is only 10.7%. That is more than a sixfold gap against Singapore’s 69%, and less than half of Vietnam’s 23%. And yet 72% of individual employees are using some form of AI in their work. That difference means a great deal of AI use is happening without organisational approval.
The Microsoft Work Trend Index 2026 reports the same shape: Thai workers lead the region in AI adoption while their organisations lag. The shop floor is running ahead and the company has not caught up. That is the reality in Thailand.
Three risks this structure creates
Individual usage running ahead is not in itself a bad thing. Fast adaptation to new tools is a strength. The problem is the organisation remaining unaware of it.
First, you cannot see the routes by which information leaves. Employees may be pasting drawing dimensions, customer part numbers, cost data and defect reports into generative AI tools on personal accounts. Not out of malice — simply because they want to finish the job faster. If the organisation offers neither a policy nor an alternative, that flow will not stop.
Second, nobody is verifying the quality of the output. As long as use stays individual, the only person checking correctness is the person using it. Research in Thailand finds that 61% of companies name the ability to interpret, verify and challenge AI output as a critical future skill. Read that in reverse and it says there is a widespread recognition that the ability is not yet sufficient. Unverified output finding its way into reports and customer replies is a quality risk that advances quietly.
Third, the work becomes person-dependent. When an individual improves their own throughput with AI, the know-how accumulates in that individual’s hands. Neither the prompts nor the instincts for checking output get shared. The day that person transfers or resigns, the work returns to its original speed and the organisation has accumulated nothing.
Thailand’s digital maturity sits at 2.12 out of 4.0
Thai companies average 2.12 out of 4.0 on digital maturity. That is up from 1.56 in 2025, so the direction of travel is positive.
It is a level just past the midpoint of a four-stage scale. The reasonable reading is that in the majority of organisations, the underlying business systems are digitised to some degree, but data is not yet joined across them and used for decisions.
That is a serious precondition for agent implementation. An agent only produces value once it can reach the data that informs a judgement. Where inventory, production plans and process progress are scattered across paper and Excel, there is nothing to give it.
An average of 2.12 tells you that a substantial number of sites need data foundations before they need agents. Skip that and move straight to agent evaluation, and you hit a wall in the PoC: there is nothing to feed it. The “insufficient domain training data — 41%” blocker mentioned earlier is exactly this state.
Three issues specific to Japanese-owned manufacturers in Thailand
On top of the above, Japanese-owned sites have their own circumstances.
Issue 1: decision-making is split between head office and the site
AI policy is usually written by the head office IT or DX department. Meanwhile, the reality of the work in Thailand is hard for head office to see. The result is a head-office-approved tool that does not fit the work on the ground, staff who use something else, and a policy that exists on paper only.
The countermeasure is for the site to specify which workflow it will use AI for, and how, and to present that to head office first. “We want to advance AI adoption” gives head office nothing to decide on. “Our delivery-date response desk handles 50 enquiries a day, averaging 12 minutes each” moves the discussion forward.
Issue 2: the language stack has three layers
In a Japanese-owned plant in Thailand, it is normal for Japanese managers to work in Japanese, local staff in Thai, and customer and head office reporting in English. The documents an agent has to work with span all three.
The problem is that the same content exists in three languages and the update timing drifts apart. A revised Japanese work standard with a stale Thai version is a common state. Unless you explicitly decide the priority order of source documents, the output will not be stable.
Issue 3: the design has to assume staff turnover
A certain amount of turnover occurs on Thai production floors. You have to assume the person you trained at go-live may not be there a year later. Getting good results out of an agent requires a feel for how to phrase an instruction, and if that feel lives only in one person’s head, usage drops the moment the role changes hands. Capturing instruction patterns as templates and writing the procedure in Thai is a precondition for it sticking.
Turning the head start in individual usage into an organisational asset
There is no need to read Thailand’s structure pessimistically. The fact that 72% of employees already use AI means resistance to implementation is low. In a Japanese head office, an invitation to “try using AI” often lands flat. In a Thai operation, the people already using it are the majority.
What is needed is not prohibition but a place to land. In practice, this order works.
- Establish the facts: find out who uses which tool for which task, with no penalty attached
- Show where the line sits on inputs: rather than listing prohibitions, give per-workflow examples of what may and may not be entered
- Provide an environment the organisation can stand behind: once the company supplies a place where work data can legitimately be handled, the reason to use a personal account disappears
- Spread the good techniques: collect what individuals have worked out, turn it into templates and share it
These four steps double as the preparation phase for agent implementation, because the exercise of mapping individual usage produces exactly the candidate list of workflows worth automating with an agent.
Note also that Thai government support frameworks are in motion. The Digital Economy Promotion Agency’s (DEPA) TH-AI Passport and tax incentives for workforce development are the relevant ones; these are covered later alongside the 90-day roadmap and budgeting.

What makes an AI agent different: the boundary against generative AI, RPA and chatbots
One of the things that leads implementation decisions astray is terminology. The phrase “AI agent” is used for things as different as a chatbot and an extension of RPA. Here we draw the boundary by asking where the judgement sits.
Divide by who holds the judgement
Classify by who holds the business judgement, rather than by underlying technology, and the practical differences become clear.
| Type | Who holds the judgement | Input | Nature of operation | Response to the unexpected |
|---|---|---|---|---|
| RPA | The human (who writes out every step in advance) | Fixed screens and files | Repeats exactly as specified | Stops or errors out |
| Chatbot | The human (who asks the question) | A human question | Answers the question | Replies “I do not know” |
| Generative AI (conversational use) | The human (instructing each time and deciding whether to accept the output) | A human prompt | Produces the requested text or analysis | The human judges and re-instructs |
| AI agent | The agent (given a goal, it determines the steps itself) | A goal plus available tools and data | Executes multiple steps autonomously until the goal is met | Tries a different approach itself, or escalates to a human |
The decisive difference is the bottom row, response to the unexpected. RPA stops when it meets something unexpected. An agent tries something else. A thing that stops is safe, because it does nothing. A thing that tries something else can carry the work forward without human involvement — and can also execute the wrong action if it picks the wrong approach. Acting autonomously and executing a wrong judgement are two faces of the same property. That is precisely why the scope of authority and the design of monitoring are preconditions for implementation. Discussions that RPA never required are mandatory for agents.
What to do with existing RPA investments
Plenty of Japanese-owned plants in Thailand have already implemented RPA. We are often asked whether adopting agents means throwing RPA away. The answer is no.
Laid out, the picture is as follows.
| Nature of the work | Suitable approach | Reason |
|---|---|---|
| Steps are fully fixed, exceptions are almost nonexistent | RPA | Cheap, fast, and the result is fully predictable |
| Steps are fixed but the input format varies | Generative AI plus RPA | Generative AI normalises the format, RPA executes |
| Judgement is required and the situation differs each time | AI agent | The steps cannot be fully written out in advance |
| Judgement is required and errors are unacceptable | Human plus an agent-produced draft | The agent proposes, the human executes |
Rows two and four give the best return on investment. Most RPA failures happen because the input does not match what was assumed, and inserting generative AI at that point stabilises it. You keep the existing RPA investment intact while reducing how often it breaks. Row four matters just as much. If you stop the agent short of final execution and have it produce a draft that a person checks and sends, you keep most of the labour saving while cutting the risk sharply.
When you evaluate vendor proposals, three questions will tell you what you are actually looking at: does the product require scenarios to be defined in advance, can it work across multiple systems, and what happens when it fails? A product that requires pre-defined scenarios is effectively a workflow tool, and judging it against agent-level expectations will not go well.
Also, if you intend the agent to reference internal documents or drawings, you will need RAG (retrieval-augmented generation) underneath it. Preparing the data to be referenced frequently takes more effort than the agent itself. That work and its costs are covered separately in what RAG implementation costs and how to approach it.
The four things to decide before you start: authority, data scope, monitoring, ownership
This is the centre of the article. As the earlier sections showed, the places where PoCs stop are concentrated in four spots. Deciding these four before the PoC begins changes the outcome substantially.
Why deciding them “before” is what works
The Digital Applied research contains a decisive figure. Organisations that put a dedicated AI operations structure in place *before* scaling up had a 5.7x higher success rate than those that decided it afterwards.
The same structure, built either before or after, produces a 5.7-fold difference. That difference comes from sequence, not from the presence of the structure. Build it afterwards and you end up designing the structure around a PoC that already exists, so the constraints of the PoC become the constraints of production. The top five barriers to scaling can be re-read in that light.
| Blocking factor | Share reporting it | Which of the four it maps to |
|---|---|---|
| Complexity of legacy integration | 63% | Data scope and authority (what gets read, from where) |
| Output quality degradation at volume | 58% | Monitoring (a mechanism to detect degradation) |
| Lack of monitoring and observability | 54% | Monitoring |
| Absence of organisational ownership | 49% | Ownership |
| Insufficient domain training data | 41% | Data scope |
All five of the top blockers map onto one of the four decisions. Nothing has stopped because it is technically hard. It has stopped because nothing was decided.
The four are covered in turn below.
Decision 1, authority: on whose behalf does the agent act?
The first thing to settle is which account’s authority the agent runs under, which systems it touches, and whether it only reads or also writes.
The question never surfaces in a PoC, because the person running it uses their own account. In production, though, the agent may run around the clock and be invoked by many users. Without settling “on whose behalf”, no permission design is possible.
Here is what has to be decided.
| Item to decide | Example options | What to weigh |
|---|---|---|
| Executing identity | A dedicated service account, or inheriting the user’s own account | If visibility differs by user, inheriting is mandatory |
| Read scope | By system, by table, or by record | Are there customer-specific or department-specific viewing restrictions? |
| Write permission | Read-only, draft creation, or full execution | Does a wrong execution affect money or delivery dates? |
| Approval requirement | Human approval on everything, only above a threshold, or none | Setting thresholds by amount, quantity or customer is the practical approach |
| Execution logging | The record format for who instructed what, and what was executed | Essential for audit and for reconstructing an incident |
The executing identity is the most important of these. Giving a dedicated account broad rights makes the design easier, but it opens a hole: data a given user is not supposed to see becomes visible through the agent. A salesperson restricted to their own accounts can simply ask the agent and get information on every account. That is a collapse of access control. The countermeasure is to design the agent to inherit the user’s own permissions. It is more work technically, and vastly cheaper than fixing it later.
On write permission, we recommend fixing the initial setting at “draft creation only”. There is no need to hand execution — placing orders, issuing shipping instructions, sending messages to customers — to an agent from the outset. Automating up to the draft still captures most of the labour saving, and the risk of a wrong execution is zero. Once accuracy is confirmed in real operation, execution rights can be opened up in stages, starting where the impact is smallest.
Decision 2, data scope: what it may read, and what it may not
Next comes the scope of data the agent references. This is not a technical question. It is a mix of management judgement and legal judgement, and it takes time to clear. Which is exactly why it has to be started early.
Start by sorting data into four categories.
| Category | Examples | How to handle it |
|---|---|---|
| Make available | Work standards, product specifications, inventory, production plans, past defect reports, internal rules | Actively organise and expose |
| Available with conditions | Customer drawings, cost data, quotations, contracts | Decide after checking the scope of the NDA with the customer |
| Do not make available | Performance reviews, payroll, health information, recruitment records | Exclude explicitly, and verify that the exclusion actually holds |
| Not organised in the first place | Person-dependent judgement criteria, rules passed on verbally | Either document them or place them out of scope |
Row two, the conditional category, is where the real arguments happen. Whether customer drawings may be read depends both on the clauses in the NDA with that customer and on the data handling terms of the service you use. If the contract requires that data not be used for training, you have to check the service’s contract tier as well.
For an operation in Thailand, PDPA — Thailand’s Personal Data Protection Act — adds another dimension. If documents containing employee information or contact details of customer personnel fall inside the reference scope, you need a clear position on the legal basis and the boundaries of that handling. A head office policy written for Japan is not sufficient on its own.
Row four is in fact the biggest wall. The reality behind “insufficient domain training data — 41%” is usually not that the data is missing, but that it was never written down. The criteria an experienced engineer applies in their head are recorded nowhere, and an agent cannot reference what does not exist in writing. What is needed here is the judgement to narrow the target workflow and document only what that workflow requires, rather than documenting everything before starting. Try to complete company-wide knowledge organisation first and a year disappears into it.
In practice, data scope design proceeds together with the plan for preparing reference documents.
Decision 3, monitoring: how you detect degradation
Third is monitoring. “Output quality degradation at volume — 58%” and “lack of monitoring — 54%” are effectively the same problem stated twice. Quality degrades. If you cannot detect it, you will not notice.
Why does it degrade? Three reasons. The input distribution changes. A PoC is tested on representative cases; production brings enquiry formats nobody anticipated, new products and new customers. Accuracy measured during a PoC is accuracy against the PoC’s input distribution and nothing more. The reference data goes stale. A work standard is revised and the reference document is not updated, and the agent answers confidently from the old standard. The failure arrives as a plausible answer rather than as “I do not know”, which is what makes this kind of degradation so awkward. User behaviour changes. As people get comfortable they start using it for purposes nobody planned for. That is a good thing in itself, but accuracy is not guaranteed outside the intended use.
Monitoring therefore has to be designed as business monitoring (is it correct?) rather than system monitoring (is it running?). At minimum, watch these four.
| Monitoring item | What you look at | Frequency |
|---|---|---|
| Output correctness | A human judges correctness on a sample | Weekly (daily in the early period) |
| Escalation rate | The share of cases the agent handed back to a human | Weekly |
| User edit rate | The share of outputs users modified rather than used as-is | Weekly |
| Unanticipated inputs | The content of enquiries with no match in the reference data | Monthly |
The user edit rate is the single most practical indicator. If users are heavily reworking the output every time, the agent is not functioning in that workflow. If it is used with almost no edits, you can start considering opening up execution rights. This metric can be captured automatically and is directly tied to business value.
Recording unanticipated inputs is equally important. Enquiries with no match in the reference data become, verbatim, the list of documents to prepare next. Monitoring is not only a detector of degradation; it is a generator of improvement material.
One caveat: sampling weekly and judging correctness takes somewhere between a few hours and a dozen or more hours per month, depending on volume. Unless you decide whose working hours this comes out of, monitoring will be a formality within three months. Which leads directly to the fourth decision.
Decision 4, ownership: whose mechanism is this?
Last comes ownership. “Absence of organisational ownership — 49%” says that in roughly half of organisations, nobody has settled whose mechanism it is.
Four roles need to be assigned.
| Role | Responsibility | Where it should sit |
|---|---|---|
| Business owner | Decides what it does and how far it is trusted. Accountable for results | Head of the department that owns the work |
| Operations lead | Daily monitoring, checking output, handling user questions | Inside the department that owns the work |
| Technical lead | Connection maintenance, permission management, model and tool updates | IT, or the vendor |
| Data lead | Updating reference documents, managing their freshness | The department that authors the documents |
The data lead is the one most often left out. Updating reference documents is the job of the department that writes the work standards — and that department is usually unaware the agent exists. Unless “when you revise a standard, update the reference source too” is built into the business process, freshness will decay without exception.
Making IT the business owner is another classic failure pattern. When a department that cannot judge whether the output is correct becomes the owner, monitoring stops functioning. The owner must always sit on the business side.
The relationship with external vendors is worth settling too. Even if you outsource the build, operations stay with you. Outsource operations wholesale and every change to the work costs money and time, and improvement stops. Try to insource everything from day one and you start too slowly. The realistic structure is to partner externally for the build and early operation while shifting operational ownership in stages to your own team. The three things usually transferred are monitoring, reference document updates and user support. That handover is broken down stage by stage in support for bringing AI in-house.
Getting all four onto a single page
The four decisions are not independent; they depend on each other. Settle authority and the data scope follows. Settle the data scope and the monitoring items follow. Settle the monitoring items and the required hours and owner follow. Before the PoC starts, we recommend producing the following single page. One A4 sheet is enough.
- Target workflow: what it is, who does it, how it is done today
- Objective: what improves by how many percent (the evaluation criteria)
- Executing identity: which account, which systems, read or write
- Data scope: what is available, what is not, what is conditional
- Approval: which operations require human approval, and at what threshold
- Monitoring: which metrics, at what frequency, judged by whom
- Roles: business owner, operations lead, technical lead, data lead, by name
- Stop criteria: what state ends the PoC
Including that last item, the stop criteria, is what pays off in practice. The pattern of spending 4.7 months and then stalling comes from an inability to decide to stop. Set a criterion in advance — “if the edit rate is still above 70% at the three-month mark, we pause” — and you prevent the sunk cost from growing.
Producing this single page takes two or three meetings once the right people are in the room. That is the scale of work behind the 5.7x difference.
What AI agent implementation costs, and what makes up that cost (2026)
Now to cost. The thing to watch here is that AI agent cost tends to be discussed as though it were licence cost. In reality, the visible cost is only part of the total.
Three price bands
Implementation broadly falls into three bands. The figures below are indicative market ranges, and they vary considerably with the target workflow, the number of systems to integrate and the state of the existing data. Treat them as a starting point for discussion, not as a quotation you can rely on for your own site.
| Band | Form | Indicative initial cost (THB) | Indicative monthly cost (THB) | Where it fits |
|---|---|---|---|---|
| A: configuring an off-the-shelf service | Configure and use the agent features bundled with an existing business application or SaaS | 100,000 to 500,000 | 20,000 to 100,000 | The work is generic and stays inside an existing application |
| B: building a workflow-specific agent | Narrow to one or two workflows and build with connections to internal data | 800,000 to 3,000,000 | 50,000 to 200,000 | The work is specific to your company and needs core system integration |
| C: multi-workflow rollout with core system integration | Spans several workflows with two-way connections to production control, accounting and similar | 3,000,000 to 15,000,000 | 150,000 to 600,000 | Company-wide rollout, multiple sites, modification of existing systems |
For readers who budget in yen — which includes most head office reporting lines — at roughly 4.5 yen to the baht, band A’s initial cost is about 450,000 to 2.25 million yen, band B about 3.6 to 13.5 million yen, and band C about 13.5 to 67.5 million yen. Exchange rates move, so use these only to get a sense of the order of magnitude.
There is no need to aim at band C from the start. As noted earlier, in 2026 scaling a single workflow already puts you in the top fifth. Getting one workflow running reliably in band B, then widening on the basis of what you learned, ends up faster and cheaper.
Three running costs that are easy to miss
Initial costs appear on the quotation. Running costs are the problem. Underestimate them at the quotation stage and the reaction after go-live is “this is more expensive than we thought”, which destabilises the decision to continue.
One: token and API charges
An agent calls the model multiple times for a single task. It decomposes the goal, searches for what it needs, evaluates the result, and tries another approach if required. It is normal for one request to consume several to more than ten times the tokens of a conversational generative AI exchange.
When estimating, work from this formula.
Monthly token cost = expected tokens per case x cases per month x unit price
The critical point is to set the per-case token figure from the PoC’s measured value. Use a catalogue figure or another company’s case study and the order of magnitude will be wrong. Note also that tokens per case rise as reference documents grow. Advancing data preparation raises the running cost, and that relationship has to be built into the budget plan.
Two: the labour cost of monitoring
As described above, monitoring is human work. Sample weekly, judge correctness, record the trend. Depending on the volume of the target workflow, allow a few hours to a dozen or more hours per month.
It never appears on a quotation, and it happens without fail. In organisations that have not allocated those hours as working time, monitoring stops within three months. Monitoring that has stopped creates a state in which degradation advances and nobody notices.
Three: retraining and reference data updates
Work standards change. Products change. Customer requirements change. Each time, reference documents need updating and output formats may need adjusting. The effort is smaller than the initial build but it is continuous. A common rule of thumb is to budget 10 to 20% of the initial build cost per year for maintenance and improvement. Skip this and a year later you have a system that answers from stale information and nobody uses.
Add all three and running costs come to a non-trivial annual proportion of the initial cost. Unless you compare on a three-year total cost of ownership, you will make the wrong call between band A and band B.
Put the cost of failure into the estimate
The USD 2.1 million average sunk cost mentioned earlier is a large-enterprise figure, but the way of thinking applies at any scale. Work out in advance how much disappears and how much remains if the PoC is cancelled.
| Cost item | Treatment on cancellation |
|---|---|
| Licence and API charges | Gone |
| Vendor build fees | Mostly gone |
| Internal labour | Gone (though the knowledge remains) |
| API enablement on existing systems | Remains (usable for other purposes) |
| Preparation of reference documents | Remains (valuable as work standardisation) |
The conclusion is straightforward. Do the investments that remain, first. Making it possible to extract data from existing systems, and organising business documentation, are not wasted even if the agent turns out to be useless. Start from tool selection instead, and the spending that disappears is what accumulates first. Changing nothing but the sequence changes what a failure costs you.
Three things to check when evaluating a quotation
Once you have a quotation from a vendor, check these three points. Is there a line item for data preparation? If it appears as a single line saying “to be prepared by the customer”, that effort stays with you — re-estimate the person-days yourself. Are the assumptions behind the running cost stated? A monthly figure with no stated assumption about cases per month and tokens per case has no basis. Are the conditions for operational handover written down? If it is vague about which tasks you take over, every minor change turns into a change order.

Where agents actually pay off on a manufacturing site
Leaving the abstractions behind, here are the areas where Japanese-owned manufacturers in Thailand see results most readily. What they share is work where a person makes a similar judgement every day, and the material for that judgement already sits inside a system.
Area 1: production control and delivery-date responses
Answering customer enquiries about delivery dates is the leading candidate for agent implementation, for three reasons. The volume is high — sites handling several dozen enquiries a day are not unusual. The material for the judgement is in the systems — inventory, production plans, process progress, procurement lead times. The procedure is routine — the flow of checking several systems, consolidating, and drafting a reply is the same every time.
What you hand to the agent is querying across inventory and plans and producing a draft reply. A person checks and sends. With that arrangement, the risk of giving a customer a wrong delivery date stays at zero while the time spent researching compresses dramatically.
The thing to watch is the difference in data freshness between systems. Where inventory is real-time, production results arrive in a daily batch and procurement updates weekly, the reliability of the consolidated answer drops. Implementing an agent is also an exercise in making freshness inconsistencies visible.
Area 2: quality and defect analysis
The initial investigation when a defect occurs suits agents well. Searching for similar past defects, extracting the manufacturing conditions for the lot concerned, checking the history of related corrective actions. Done by hand it takes hours, and all the material sits in the records. What you hand over is the research part — “for this defect mode, list similar records from the past year and the countermeasures taken at the time” — while determining the cause and deciding the countermeasure stays with people.
The precondition is that records can be joined. Inspection records with no equipment ID, clocks that differ between systems, lot definitions that vary by process — in that state, there is no dataset to hand the agent. This is where “insufficient domain training data — 41%” shows up most concretely on a factory floor.
Area 3: equipment maintenance
Here the use is presenting, across sources, the response history for similar past faults, the relevant section of the equipment manual and the stock position for spare parts. Maintenance work is especially prone to becoming person-dependent. An experienced technician can guess the cause from the symptom alone, and the reasoning behind that guess is nowhere in writing. The value of an agent here is in easing that dependence.
Realistically, though, past fault records often say little more than “attended to”. Unless symptom, cause, action and result are retained in a structured form, the agent has nothing to reference. In this area you often have to start by improving the record format.
Area 4: procurement and purchasing
Comparing quotations, supporting supplier selection, detecting early signs of delivery delays. Comparing quotations from several suppliers line by line and organising the differences is labour-intensive precisely because the formats vary. That is generative AI’s strong suit.
Do not hand order placement to the agent. Operations that move money stop at the draft, following the principle set out in the authority design. Automating up to the comparison table alone visibly reduces the buyer’s workload.
Area 5: multilingual documentation
This is the area where the benefit is clearest at a site in Thailand. Rolling out Japanese work standards in Thai, summarising Thai daily reports in Japanese, organising English customer requirement specifications in both Japanese and Thai. The value of an agent is not translation as such but translation that respects your internal terminology. If defect names, process names and equipment names can be translated consistently against your internal code system, translation quality stabilises. That is difficult to achieve with a general-purpose translation tool.
The operational essential is deciding which language is authoritative. Make two languages authoritative and you generate discrepancies with every update. The three-layer language issue raised earlier becomes a concrete design item right here.
How to choose where to start
Five areas have been laid out, but they should not all start at once. There are four criteria for selection.
| Criterion | What to look at |
|---|---|
| Volume | Occurrences per day. Too few and there is no effect |
| Where the material sits | Is the judgement material inside a system, or in someone’s head? |
| Impact of an error | Does a mistake reach money, delivery dates or safety? |
| Ease of cooperation | Is the owning department willing? |
Pick a first workflow with high volume, material inside the systems, low impact from errors, and a cooperative department. At most sites that means delivery-date responses or multilingual documentation. What matters is not choosing your most painful workflow first. The first objective is validating the mechanism and the organisation, not maximising the return. Choose a difficult workflow and the points you are trying to validate get tangled up with difficulties specific to that work, leaving you unable to tell what actually caused the stall.
A 90-day roadmap for Japanese-owned manufacturers in Thailand
Here is everything above, reduced to an executable order. The assumption is a site where individual use of generative AI has already begun but nothing has been implemented at organisational level. Given Thailand’s structure, that is the most common pattern.
| Period | Main activities | Deliverables | Departments involved | Where it stumbles |
|---|---|---|---|---|
| Day 0-30 | Establishing the facts and making the four decisions | Inventory of actual usage, choice of target workflow, the one-page design, evaluation criteria, stop criteria | Owning department, IT, management | The target workflow is not narrowed down and several are started at once |
| Day 31-60 | A scope-limited PoC and a trial run of monitoring | A working agent, monitoring records, a measured edit rate, a list of gaps in reference data | Owning department, IT, vendor | Monitoring hours are not secured, so no measured values are obtained |
| Day 61-90 | The production decision and confirming the structure | Go or no-go on production, roles confirmed, re-estimate based on measured costs, rollout plan | Owning department, management, IT | With no decision criteria, the project is simply extended |
Day 0-30: decide everything without touching a tool
Most of what happens in these 30 days involves no tools at all.
- Interview people about actual individual usage: with no penalty attached, find out who uses what for which task. This becomes the candidate list of target workflows
- Narrow to a single target workflow: choose using the four criteria above. Do not start two at once
- Measure the current effort: time the volume and the minutes per case with a stopwatch. This becomes the baseline for the later ROI calculation
- Put the four decisions on one page: authority, data scope, monitoring, ownership
- Set the evaluation criteria and the stop criteria: what improvement by what percentage counts as a pass, and what state means stopping
- Take stock of reference data: for the target workflow, which documents exist, where, in which language, and when they were last updated
The deliverable is a document, not a working system. Because it also gives you a package to put in front of vendors, quotation accuracy improves.
Where it stumbles is the failure to narrow the target workflow. Bring the relevant departments together and each will bring its own problem, and the scope swells. Stating explicitly that “this cycle covers one workflow; the next workflow comes in the next cycle” is what brings it back to a point. Note also that these 30 days can be done entirely with your own staff. Completing the most important decisions before any significant spending occurs is the single biggest lever for containing sunk cost.
Day 31-60: run it small and actually do the monitoring
Build the PoC narrowly around the target workflow and use it in the real work. The objective here is not to confirm that it runs but to obtain measured values.
- Operate with a limited set of users: have a handful of people use it in real work. Do not roll out to everyone
- Actually perform the monitoring: sample weekly and judge correctness. Measure the effort this takes, too
- Record the edit rate: the share of outputs used as-is versus the share modified
- Measure token consumption: get a real per-case figure. It becomes the basis of the running cost estimate
- Record gaps in the reference data: enquiries the agent could not answer become the preparation list verbatim
- Decide how exceptions are handled: when the agent cannot make a judgement, who does it go to and how
Where it stumbles is monitoring. Two consecutive weeks of “we are too busy, skip it this week” and it never happens again. The monitoring owner and the time slot have to be secured as work during Day 0-30. The other trap is widening the user group too far. As more people gain access, feature requests pile up and the thing you are supposed to be evaluating shifts underneath you. Day 31-60 should freeze the feature set and concentrate on measurement.
Day 61-90: decide, and lock in the structure
Once the measured values are in, make the production decision. There are four inputs.
- Were the evaluation criteria met? How did the measurements compare with the criteria set in Day 0-30
- Is the edit rate acceptable? If users are heavily reworking every output, it is too early to go to production
- What is the measured running cost? Derive the annual cost from measured token usage and compare it against the savings
- Is there a structure that can sustain monitoring? Are the people and the hours available
If even one of the four is not satisfied, treat pausing rather than extending as a live option. That is the point of setting the stop criteria in advance.
If they are satisfied, lock in the structure.
- Name the business owner, operations lead, technical lead and data lead individually
- Build the procedure for updating reference documents into the existing document management process
- Produce a Thai-language operating procedure
- Decide the record format and storage location for monitoring
- Choose one candidate for the next target workflow
What you have at the end of 90 days is not a finished system but a basis for decision and a structure. Following that order ends up faster than aiming at a company-wide rollout from the outset. The 5.7x difference lives in how these 90 days are used.
Building DEPA and BOI schemes into the budget plan
As a Thailand-specific addition, it is worth building a check on public support schemes into the plan.
DEPA is running the TH-AI Passport (a THB 1.6 billion programme, with registration opening on 5 June 2026), and launched Coding Thailand 2026 on 16 March 2026 at Siam Square SiamScape. On the workforce side, corporate training expenditure related to human resource development carries a tax deduction of up to 250%. Because the bottleneck in agent implementation is usually people rather than tools, it is worth checking the tax treatment of training expenditure while confirming the operating structure in Day 61-90.
If your site holds BOI privileges, the treatment of digital-related investment is worth confirming as well. Eligibility conditions, coverage and application procedures follow the latest guidance from the responsible agencies, so this is listed here only as a checkpoint at the budgeting stage. The practical sequence is not to build the plan around a scheme, but to decide the plan and then check which schemes apply to it.

Measuring the effect and calculating ROI
Here is how to work out the numbers behind the investment decision. Given that inadequate evaluation criteria, at 64%, is the single largest blocker, deciding the calculation in advance is a direct countermeasure.
The basic approach: four items
There is no need to make this complex. Estimate the annual benefit from the following four items and compare it against cost.
| Item | Formula | Ease of quantification |
|---|---|---|
| 1. Labour reduction | (current time per case − time per case after) x cases x labour rate | High |
| 2. Faster response | Reduced opportunity loss from shorter response times | Moderate |
| 3. Quality stability | Less rework from reduced person-to-person variation and fewer omissions | Low |
| 4. Knowledge accumulation | Reduced person-dependence, less effort in handovers | Low |
Put item 1 at the centre of the investment decision. Items 2 to 4 carry real value, but they rest on many assumptions and quantifying them weakens the credibility of the estimate. The healthy approach is to start from a scope where item 1 alone gives you a plausible payback, and present items 2 to 4 as additional effects you can also expect.
A worked example: delivery-date responses
Take area 1 from the previous section and run the numbers. This is an illustrative calculation to show the method, not a claim about any particular result. Substitute your own measured values.
Assumptions
- Target workflow: responding to customer delivery-date enquiries
- Volume: 50 cases per day x 240 working days per year = 12,000 cases per year
- Current time: 12 minutes per case on average (querying several systems plus drafting the reply)
- Time after implementation: 4 minutes per case on average (the agent drafts, a person checks and sends)
- Labour rate: THB 400 per hour (fully loaded cost for administrative staff)
Calculating the benefit
- Current annual effort: 12 minutes x 12,000 cases ÷ 60 = 2,400 hours
- Annual effort after: 4 minutes x 12,000 cases ÷ 60 = 800 hours
- Effort saved: 1,600 hours per year
- In money: 1,600 hours x THB 400 = THB 640,000 per year
Calculating the cost
- Initial cost: THB 900,000 (near the bottom of band B, including connection to existing systems)
- Annual running cost: THB 258,400
– Token and API charges: approximately THB 140,000
– Monitoring labour: 8 hours per month x 12 months x THB 400 = THB 38,400
– Maintenance and reference data updates: approximately THB 80,000
Calculating the payback
- Annual net benefit: 640,000 − 258,400 = THB 381,600 per year
- Simple payback period: 900,000 ÷ 381,600 = about 2.4 years
For shop floor systems in manufacturing, the general sense is that a payback of two to four years passes an investment decision without difficulty. This estimate sits inside that range.
Three common errors in the estimate
The calculation above is simple, but three mistakes recur in practice.
Error 1: leaving running costs out of the cost side
Compute the payback from the initial cost alone and you get 900,000 ÷ 640,000 = about 1.4 years, a long way from the actual 2.4. Token costs and monitoring labour in particular never appear on a quotation, so they have to be added deliberately.
Error 2: assuming an optimistic post-implementation time
It is tempting to assume that 12 minutes becomes 1 minute, but the time a person spends checking never goes away. Assume it on paper instead of using the Day 31-60 measurements and this part comes out too low.
Error 3: not deciding what the freed-up time is for
Save 1,600 hours and do nothing with them and there is no financial effect. Are you absorbing increased production without adding headcount, or redirecting people to higher-value work? Leave that vague and the post-go-live verdict is “it is easier now, but the numbers have not changed”.
Write down the evaluation criteria before you implement
More important than the formula itself is writing this calculation down during Day 0-30.
Try to measure the effect after implementation and you find you have no baseline to compare against. All you can produce is “it feels a bit faster”, which gives neither the investment decision nor the continuation decision anything to stand on. Measuring current volume and time per case is something that can only be done before implementation.
The evaluation criteria document should contain at least these four things. The metrics you will measure (time per case, volume, edit rate). How you will measure them (who, when, by what method). The pass mark (at what level you move to production). The stop line (at what level you stop).
The countermeasure to the single largest blocker — inadequate evaluation criteria at 64% — is writing those four lines. Not technology, not budget. Documentation.
Frequently asked questions
How much does AI agent implementation cost?
It varies widely with the form of implementation. As general market ranges, configuring an off-the-shelf service runs around THB 100,000 to 500,000 initially, building a workflow-specific agent narrowed to one or two workflows around THB 800,000 to 3,000,000, and a multi-workflow rollout with core system integration around THB 3,000,000 to 15,000,000. These vary considerably with the target workflow, the number of systems to integrate and the state of your existing data.
When evaluating a quotation, checking the running cost matters more than checking the initial cost. Token and API charges, monitoring labour and reference data updates are the three that rarely appear on a quotation and that recur continuously. Compare on a three-year total or you will decide wrongly. If cost is a constraint, follow the 90-day roadmap and start with a single workflow. Fixing the design before requesting quotations also narrows the spread of the figures, because the scope is clear.
What is the difference between an AI agent and RPA?
The difference is who holds the judgement. With RPA, a person writes out every step and the tool executes them exactly. Unexpected input makes it stop. An AI agent is given a goal and determines the steps itself. Faced with an unexpected situation, it tries a different approach.
That difference is the value and simultaneously the risk. A thing that stops is safe because it does nothing, but the work does not progress until a person intervenes. A thing that judges for itself moves the work forward, and can also execute a wrong judgement. Which is why the scope of authority and the design of monitoring are mandatory.
In practical terms, there is no need to discard RPA and replace it with agents. For work whose steps are completely fixed, RPA is cheaper and more certain. And since most RPA failures come from variation in input format, inserting generative AI in front of it to normalise the format is often the highest-return arrangement available.
How long does implementation take?
Narrowed to a single workflow, 90 days is the guide for reaching a decision. That breaks down as 30 days for establishing the facts and design, 30 days for a scope-limited PoC and monitoring, and 30 days for the production decision and confirming the structure.
Bear in mind that the research puts the average PoC duration at 4.7 months, with 64% stalled for six months or more afterwards. The main driver of that extension is not technology; it is repeating “let’s watch it a little longer” with no decision criteria in place. Setting the stop criteria at the start prevents that stall. Note also that if the documents to be referenced are not yet prepared, a separate period for that work is required. Narrowing the target workflow narrows the preparation scope too, which is another reason the “start with one workflow” principle works.
Can a small site implement this?
Yes — and in some respects a small site has the advantage. With fewer stakeholders, agreement on authority and data scope comes faster. Ownership is easier to assign. The leading causes of the 88% stalling at PoC are governance friction and the absence of organisational ownership, and both get worse as the organisation gets larger.
Where size does work against you is volume. ROI is determined by time saved per case multiplied by the number of cases, so workflows with few cases produce little effect. The right test is the case volume of the target workflow, not the headcount of the site. For work occurring a few times a day, simplifying the procedure will do more than agent implementation. At a small site, the realistic sequence is to start with band A, configuring an off-the-shelf service, confirm the effect, and then consider a purpose-built agent.
Can we discuss this in Japanese from a site in Thailand?
Yes. TOMAS TECH is based in Bangkok and works in Japanese, Thai and English. In Japanese-owned manufacturing in Thailand, it is normal for Japanese managers to discuss in Japanese while local staff operate in Thai, and the design has to assume that split.
What matters operationally is not letting everything end in Japanese. If the policy, the operating procedures and the reference documents exist only in Japanese, they functionally do not exist for local staff. Given that 72% of Thai employees already use AI at work, preparing material in a language local staff can read is a precondition for adoption sticking.
The same applies to the relationship with head office in Japan. A site that specifies which workflow it will use AI for, and how, and presents that upward, moves the discussion further than an abstract head-office-led policy.
Do we have to replace our core systems to implement this?
No replacement is needed. What is needed is making it possible to extract data from the systems you already have. The fact that the top scaling blocker is “complexity of legacy integration” at 63% tells you this is the biggest technical question.
The realistic approach is to leave existing systems untouched and create a read-only path. Check database read permissions, scheduled data extraction and the availability of APIs, and build a configuration that only needs to read. Integrations that involve writing have a much wider blast radius, so leave them to a later stage. Note that this work remains as an asset even if the agent turns out to be unusable. This is exactly the kind of investment to do first.
How do we manage the risk of information leakage?
There are two things to manage: leakage from the agent the organisation implemented, and leakage from individual use the organisation is unaware of.
The first is manageable by design. Explicitly categorise the data the agent may reference, exclude HR, payroll and health information, and actually verify that the exclusion holds. Designing the agent to inherit each user’s own permissions closes the hole where otherwise-invisible data becomes visible. Retaining execution logs makes after-the-fact tracing possible.
The second is the more serious in practice. In Thailand, 72% of individual employees use AI at work while organisational adoption sits at 10.7%. Prohibition has little practical effect. Providing an environment the organisation stands behind is the most realistic leakage countermeasure available. Alongside it, establish a policy, provide it in a language local staff can read, and clarify how it sits with Thailand’s PDPA. All three are worth starting on whether or not you implement an agent.
Conclusion
What separates a successful AI agent implementation from a failed one is neither model selection nor tool selection. It is whether authority, data scope, monitoring and ownership were decided before the PoC began. The 88% stall at PoC because those four are deferred and the project starts from “let’s see if it works”.
Here are the main points of this article.
- 88% never reach production. The leading blockers are inadequate evaluation criteria at 64%, governance friction at 57% and model reliability at 51%. The top two have nothing to do with model performance and stem from decisions not taken
- 78% hold a PoC and only 14% reach company-wide operation. 64% have been stalled for six months or more, with an average PoC duration of 4.7 months before the stall. The average sunk cost of a failed project is about USD 2.1 million
- Japan’s business usage rate for generative AI is 55.2%, behind China at 95.8%, the United States at 90.6% and Germany at 90.3%. McKinsey finds 62% experimenting and 23% having reached scale. Running even one workflow reliably puts you in the top fifth
- Thailand shows an inversion: 10.7% of organisations against 72% of individual employees. Organisational adoption trails Singapore at 69% and Vietnam at 23%, but low resistance on the floor is a genuine strength. A maturity level of 2.12 out of 4.0 indicates that many sites need data foundations before they need agents
- The four decisions are the substance. Authority means on whose behalf it acts, data scope means what it may and may not read, monitoring means how degradation is detected, ownership means whose mechanism it is. Organisations that put a dedicated structure in place before scaling had a 5.7x higher success rate
- Judge cost on the running side. Token and API charges, monitoring labour and reference data updates rarely appear on a quotation. Do the investments that survive a failure first — data extraction from existing systems and preparation of business documentation
- Use 90 days to produce a basis for decision. Spend the first 30 days touching no tools and writing the design and criteria down, the next 30 measuring a scope-limited PoC with real monitoring, and the last 30 deciding on production and confirming the structure
- Build ROI around labour reduction. Subtract the running cost, keep the human checking time in the estimate, and decide what the freed-up time is for. Evaluation criteria can only be created before implementation
On the external side, Gartner forecasts that 40% of enterprise applications will embed AI agents by the end of 2026, and the market is expected to reach USD 7.8 billion in 2026, up 50% year on year. Even if you build nothing, agents will arrive inside the applications you already run. The window in which the design of authority and data scope can be deferred is not long.
One last time: decide the four before you experiment. That is the whole of this article.
Working through the four decisions can begin before you sign for any tool, simply by laying out the current state of a target workflow next to the data you can actually reach. In Thailand, TOMAS TECH has worked on the question of where to extract factory data from and how to connect it to operational judgement, through building shop floor systems including the PEGASUS production and energy management system. If you only want to talk through design-stage questions — which workflow suits your operation as a starting point, or how much data your existing core systems can realistically expose — that is entirely fine, and it is no problem at all if you are still at the very beginning of your consideration. We will listen to the situation at your site and lay out the options for how to proceed. You are welcome to get in touch through our contact form.