Generative AI implementation failure is usually discussed as a technology problem, but when you take a stalled project apart, almost all of the causes sit in decisions made before the work started. MIT NANDA found that around 95% of corporate generative AI pilots produced little or no measurable impact on profit. This article sets out what separates that 95% from the remaining 5%, then covers the failure patterns that occur specifically at Japanese-affiliated sites in Thailand and the wider ASEAN region, and the decisions worth settling before you begin.
Why does generative AI implementation fail – start with the numbers
Companies often open a conversation on the assumption that their own proof of concept is the unusual one that got stuck. In practice, being stuck is the majority position, not the exception. It helps to line up the published research first and see where the average sits.
NANDA, a research initiative under MIT Media Lab, published The GenAI Divide – State of AI in Business 2025. It reports that roughly 95% of the generative AI pilots at the companies surveyed had stalled, producing little or no measurable impact on profit. By contrast, only around 5% achieved rapid revenue acceleration. The study was released in 2025 and is still widely cited in 2026 whenever corporate AI investment decisions are debated.
RAND Corporation research puts the failure rate for AI projects at over 80%, roughly twice the failure rate of conventional IT projects.
Gartner predicts that by 2026, 60% of AI projects unsupported by AI-ready data will be abandoned. Its observations show project failure rates jumping from 17% in 2024 to 42% in 2025, with 72% of companies likely to halt AI proofs of concept because their data is not ready.
Domestic research in Japan shows the same pattern. A 2026 survey covering 1,008 managers found that only one user in four had actually reduced working hours using generative AI. The main causes of failure cited were vague implementation objectives, undefined success metrics, and excessive expectations set without understanding the technical limits of AI.
| Study | Scope | Key figures | What it indicates |
|---|---|---|---|
| MIT NANDA | Corporate generative AI pilots | About 95% stalled, about 5% accelerating revenue | The distribution is sharply polarised |
| RAND Corporation | AI projects generally | Over 80% fail, roughly twice conventional IT | AI adds its own layer of difficulty |
| Gartner | Forecast abandonment of AI projects | 60% abandoned by 2026, failure rate from 17% to 42% | Data readiness is the dividing line |
| Japan domestic survey | 1,008 managers and others | Only one in four cut working hours | Adopting and actually using are different things |
Line these up and a common thread appears. None of the studies says the failure was caused by the model not being clever enough. MIT NANDA points to enterprise integration problems and misallocated budget. Gartner points to inadequate data preparation. The Japanese survey points to missing objectives and metrics. All of these sit outside the model, in the design work done by the organisation adopting it.

That changes what a management meeting should be spending time on. Comparing which model to use or which tool to sign for has almost no effect on moving from the 95% side to the 5% side. What determines the move is how the problem is chosen, how the partner is chosen, and what condition the data is in.
Generative AI implementation failure breaks down into four structures
Looked at case by case, the causes appear different every time. Organised, they converge on about four types.
The objective was vague from the start
This is the most common. A proof of concept that began as let us just try it has no defined ending. With no defined ending there is nothing to judge whether it should continue or stop, so it quietly dies once the champion loses enthusiasm. When the Japanese survey lists vague objectives and undefined success metrics as the leading causes of failure, this is the structure it is describing.
Worth noting is that an objective written at the granularity of improve efficiency or drive DX is, in practice, the same as having no objective at all. The only version usable for a decision is one specified down to which department, which task, how much time, by when, and by how much.
The learning gap – general purpose tools do not learn your workflow
The enterprise integration problem MIT NANDA identifies is this learning gap. General purpose generative AI tools are strong for individual use, but they do not learn from and adapt to enterprise workflows. The expectation that a tool will gradually settle into your business the more you use it simply does not hold.
On site it surfaces like this. The way your quotations are written has to be explained again every time. The tool does not know your part numbering scheme, so answers stay generic. A correction given last week is repeated next week. At the individual trial stage these are small inconveniences. The moment you try to embed the tool into a whole department, the same inconvenience is multiplied by the number of users.
Misallocated budget – the ROI is in the back office
MIT NANDA also cites budget misallocation. Generative AI budgets at many companies are weighted toward sales and marketing, while the actual return on investment sits in back office automation.
Applied to Japanese-affiliated manufacturers in ASEAN, that observation lands. Sales support and customer-facing chat are areas where effect is hard to measure, and where a local subsidiary’s revenue is largely determined by its relationship with headquarters and trading partners. Back office work is different. Accounting voucher processing, purchasing quotation comparison, organising quality defect reports, answering staff questions about internal rules – all of these can be counted as a number of cases and a time per case. Investing where you can count makes the next budget request far easier.
The data is not AI-ready
The data readiness problem Gartner identifies has two meanings in a generative AI context. One is that the necessary data was never digitised. The other is that it was digitised but is not in a searchable form.
Five years of daily work reports sitting in a shared folder as scanned PDF images are stored, but from the AI’s point of view they do not exist. The same applies when Excel files are scattered across individual local folders and the same part appears under three different spellings.
| Failure structure | Typical symptom | What to settle before starting |
|---|---|---|
| Vague objective | Nobody can state the exit condition for the pilot | Write the target task, hours to be cut and decision date as numbers |
| Learning gap | The same background is re-explained every time | Include a mechanism for referencing your own data in the design |
| Misallocated budget | Budget sits where effect cannot be measured | Start with work countable by case volume and time per case |
| Data not AI-ready | Material scattered across scans and personal folders | Inventory the location and format of the data first |
These four are not independent. A vague objective means the target task is never fixed, and without a target task there is no way to decide which data to prepare. So the order of work is objective, target task, data, then tool. Projects that begin with tool comparison stall because that order has been reversed. Our step by step view of sequencing and cost is set out in How to approach generative AI implementation and what it costs.
Failure patterns specific to sites in Thailand and ASEAN
Everything above applies worldwide. What follows applies to Japanese-affiliated companies operating in Thailand and Vietnam. When an initiative running smoothly at headquarters stalls only at the local site, the cause is almost always in this section.

Headquarters generative AI guidelines do not work locally
Most Japanese-affiliated groups have their headquarters IT department maintain a generative AI usage policy. That policy is almost always written around a Japanese domestic working environment and Japanese language documents.
By the time it reaches the local site, gaps appear. The approved internal tool has a Japanese-only interface that Thai staff cannot operate. Confidentiality classes are defined against the headquarters document taxonomy, so nobody can tell which class the import and export paperwork or local regulatory filings handled daily on site belong to. The approval desk runs on Japan hours, so answers do not come back inside local working time. The result is a policy that formally exists but that nobody uses in practice.
That state is dangerous because staff do not stop using AI – they move to personal accounts. Once unmanaged usage spreads, the next security incident triggers a blanket ban on generative AI across the whole site. That is the most expensive form of failure.
Accuracy degradation in Thai and Vietnamese was never verified
Accuracy that looked good in Japanese and English does not automatically hold in Thai or Vietnamese. It drops visibly in documents that mix the abbreviations local staff use daily, process names that settled into the local language on the shop floor, and Japanese technical terms borrowed phonetically into local spelling.
Despite this, a proof of concept is frequently evaluated only through Japanese expatriates working in Japanese. When it then rolls out and local staff start using it, the reaction is that it is not what was expected, and usage stops. Evaluation has to be done in the language that will actually be used, on documents that will actually be used.
Misreading the distribution of AI literacy among local staff
The assumption that local staff will not be used to AI no longer matches the facts in Thailand. Microsoft research puts AI usage among Thai data workers at 32%, twice the global average. Frontier Professionals, the advanced user segment, make up 32% of Thai employees, again twice the global average of 16%.
In other words, some local staff may be more fluent than the expatriates. The issue is not the average but the distribution. Employees who already use these tools daily sit in the same department as employees who have never touched them at work, so a single uniform group training session is boring for the top segment and too fast for the bottom. We break down why usage decays after training in Supporting the adoption and retention of AI use.
The data is still split between Excel and paper
At manufacturing sites across Thailand and Vietnam it remains common for the core system to cover only accounting and inventory, while production results, quality records and maintenance history stay in shop floor Excel files and paper daily reports. Ask generative AI to explain defect trends in that state and it has no data to reference, so it returns generalities.
As Gartner’s forecast indicates, projects without AI-ready data behind them are the ones that get abandoned. The important point is that you do not need to build a company-wide data platform first. You need only the data relating to the single task you chose first, in a form that can be referenced.
The success criteria for the pilot were never set
A frequent pattern at local sites is a pilot whose success criterion is how the shop floor feels about it. Sentiment cannot be measured, so no conclusion arrives. A pilot without a conclusion gets vaguely extended at each reporting cycle and disappears when the owner is transferred.
| Site-specific failure pattern | What is happening | What to settle first |
|---|---|---|
| Headquarters policy does not function | Unusable, so staff drift to personal accounts | A local reading of the policy matched to local document types and hours |
| Local language accuracy unverified | Evaluated in Japanese only, fails in production | State the evaluation language and real documents in the pilot plan |
| Literacy distribution misread | Uniform training suits neither end | Survey current usage first and design by segment |
| Data still scattered | Only generalities come back, no value emerges | Prepare the data for one target task only |
| No success criteria | No decision possible, the pilot fades | Document the decision date, metric and exit condition before starting |
None of these five is technically difficult. They are matters that could be settled in a discussion lasting from thirty minutes to a few hours, and they turn into failures because months pass without that discussion happening.
AI diffusion in Thailand is faster than expected – waiting is also a failure mode
Delaying the start in order to avoid failure looks like the safe choice. The Thai market numbers show the cost of that choice rising year on year.
According to the Global AI Diffusion Report presented by Microsoft at the Microsoft AI Tour Bangkok on 10 June 2026, meaningful AI usage in Thailand rose from 9.1% at the start of 2025 to 12.4% in the first quarter of 2026, placing the country second worldwide by growth rate. Korea leads at 43.2% and Japan is third at 34.1%. The report also states that 51% of executives have set out a strategic AI direction. Alongside this, Microsoft has announced plans to invest one billion US dollars in Thailand from 2026 to 2028 and to train 150,000 Thai workers.
Manufacturing-specific forecasts exist as well. AI use in Thai manufacturing is projected to rise by up to 15% by 2030, with demand forecasting expected to improve service levels by 65%, AI-driven industrial automation expected to lift productivity by 20%, and predictive maintenance expected to cut machine downtime by 53% and unsafe behaviour by 90%. It is also reported that 73% of Thai companies are preparing to adopt AI.
| Indicator | Figure | Source |
|---|---|---|
| Meaningful AI usage in Thailand | From 9.1% in early 2025 to 12.4% in Q1 2026 | Microsoft Global AI Diffusion Report |
| Ranking by AI adoption growth | Thailand second worldwide, Korea 43.2%, Japan 34.1% | As above |
| AI usage among data workers | Thailand 32%, twice the global average | As above |
| Machine downtime under predictive maintenance | Expected reduction of 53% | Thai manufacturing AI outlook |
| Companies preparing to adopt AI | 73% in Thailand | As above |
These numbers need careful reading. Rising adoption does not mean that adopting will succeed. They should be read together with the 95% from MIT NANDA. The number of companies starting is rising quickly, and most of them are not reaching a result.
The realistic position for a Japanese-affiliated operation in ASEAN is therefore neither rushing in nor continuing to wait. It is narrowing the scope and getting one case to work properly.
What did the successful 5% do differently
The MIT NANDA study also describes the tendencies of the roughly 5% that succeeded. In summary, they select one problem, execute on it, and build partnerships intelligently.

They pick one problem and execute on it
What successful cases have in common is that they did not widen the scope. Rather than starting from a company-wide rollout plan, they limited themselves to one task in one department and produced numbers there.
For a site in Thailand or Vietnam, tasks that make good first candidates include summarising and consolidating daily production reports, extracting line items from supplier quotations into a comparison table, translating quality defect reports between the local language and Japanese with a summary of key points, and answering staff questions about internal rules and work instructions. Each of these has a countable case volume, a measurable time per case, and an identifiable owner.
What should not be chosen first is work whose effect only appears indirectly. Improving meeting quality or stimulating idea generation are not worthless, but they cannot be judged, so they do not suit a pilot.
Do not insist on building in house – 67% for bought tools against 33% for internal builds
The clearest difference in the MIT NANDA study was in how tools were procured. Tools purchased from specialist vendors had a success rate of around 67%, while internally built tools succeeded around 33% of the time. Internal builds succeeded at less than half the rate of purchased tools.
The gap is not about the ability of in-house engineers. Building internally means owning model selection, business data integration, permission design, operational monitoring and continuous accuracy evaluation all at once, and if any one of them thins out, the whole thing stops. At a local site in particular, only one or two people can cover that work, and the project halts the moment that person transfers or resigns.
| Procurement route | Success rate in the study | What matters at a local site |
|---|---|---|
| Purchased from a specialist vendor | About 67% | Operation and accuracy evaluation can be carried externally over time |
| Built in house | About 33% | Dependence on one person, stopping when they move on |
This is not an argument against building in house. If you start by purchasing, establish an operating pattern, and then progressively widen what your own team runs, the success rate for internal work rises. We set out how to draw that line in Support for bringing AI capability in house.
They do not defer data preparation
The third point is the reverse of Gartner’s observation. Successful organisations check where their data is immediately after choosing the target task. It is the work of listing which folder holds what format, covering which period, updated by whom.
That inventory has a secondary benefit. It sometimes reveals that the data for this task was never retained at all, in which case the pilot target can be swapped out early. Compared with discovering the same thing after several months, it is a very cheap failure.
How to read generative AI use cases and adoption examples
A great deal of evaluation time goes into collecting other companies’ use cases. The cases themselves are useful, but reading them wrongly leads to wrong decisions.
First, the effect figures in a case study assume that company’s data condition and work volume. Put the same tool on the same task and the same numbers will not appear if your records are still on paper. When reading a case, the description of what state the company was in before starting is more useful than the effect figure.
Second, published cases are biased toward successes. As MIT NANDA shows, around 95% of pilots are actually stalled, so the distribution of published cases differs sharply from the distribution of the population. An area where many cases can be found is not therefore an area where success is more likely.
Third, the areas in manufacturing where effects show up as numbers are already reasonably well known, generative AI or not. The Thai manufacturing outlook figures – 65% service level improvement in demand forecasting, 20% productivity gain from industrial automation, 53% reduction in machine downtime from predictive maintenance – all come from areas where the input data is captured automatically from equipment. Where people judge and people record, the first task is to start capturing the record.
| How to read a case | What to look for | Easy mistake |
|---|---|---|
| Effect figures | Data condition and work volume before starting | Assuming the same figure applies to you |
| Industry featured | Whether the nature of the work resembles yours | Assuming same industry means reproducible |
| Tool used | Operating structure and headcount | Copying only the tool name |
Rather than spending time collecting cases, listing three candidate tasks of your own and checking where the data for each of them sits will settle the decision faster. How to design the metric that measures the effect is covered in Measuring the effect of AI adoption and designing for ROI.
What does generative AI implementation support actually cover
Generative AI implementation support is a broad term, and it is not obvious what can be asked of it. Taking the analysis above, the areas with the highest value from outside help are, in order, as follows.
The highest value is problem selection before starting. Listing candidate tasks, ranking them by where the data sits and how measurable the effect is, and choosing the first one. Get this wrong and every later stage is wasted.
Next is designing the success criteria. Documenting the decision date, the metric, the target value and the exit condition. As noted, a pilot without these produces no conclusion.
Third is data preparation, limited to the data relating to the target task, kept separate from any company-wide data platform programme.
Fourth is operational design and adoption in the local language, covering accuracy verification in Thai or Vietnamese, segment-based support for local staff, and visibility over actual usage.
| Scope of support | Main work | Value of external help |
|---|---|---|
| Problem selection | Listing and ranking candidate tasks | Highest, since a mistake here wastes everything after |
| Success criteria design | Documenting decision date, metric and exit condition | High, as internal-only efforts stay vague |
| Data preparation | Locating and shaping data for the target task | Medium, feasible internally if scope is limited |
| Operational design and adoption | Local language accuracy checks, segmented enablement | High, especially at local sites |
Even with external support, the final decision on which task to target should stay inside the company, because the people on site are the ones who know where it hurts. What an outside party can supply is the axis for choosing and the shapes that failure takes elsewhere. How to structure that engagement, including the available contract formats, is covered in Generative AI consulting.
Ten items to confirm before starting
The following checklist condenses everything above. If even one row is blank, starting a pilot in that state will not produce a conclusion.
| Item to confirm | What should be filled in |
|---|---|
| Target task | One department and one task, specifically named |
| Current effort | Case volume and time per case, as numbers |
| Target | Hours or cases to be reduced, as a number |
| Decision date | A specific calendar date |
| Exit condition | Written statement of what triggers stopping |
| Data location | A list of where the referenced data sits and in what format |
| Data format | Dependence on scanned images and personal folders resolved |
| Language of use | Evaluation language matches actual usage language |
| Users | Who and how many will use it, decided |
| Procurement route | Purchase or internal build, decided with a reason |
Of these ten, the three that most often trip people up are current effort, exit condition and language of use. Current effort is usually not measured, exit conditions are uncomfortable to write, and evaluation tends to default to Japanese.
Common questions
Our pilot has been stuck for months. Should we shut it down
First check whether the pilot ever had a decision date and an exit condition. If it did not, the situation is closer to never having properly started than to having failed. In that case, rather than deciding to withdraw, it is faster to narrow back down to one target task and write the conditions from scratch. If the conditions existed and were not met, withdraw, and record why the target was missed. That record raises the accuracy of the next attempt.
How much budget should we set aside for generative AI implementation
The order of magnitude changes with scope, so no single figure applies. The way to think about budget does have an order, though. For the first case, selecting the target task and preparing the data take more effort than the tool subscription. Budgeting only for tool fees leaves you short against the real workload. The stage by stage breakdown is in How to approach generative AI implementation and what it costs.
How should we evaluate Thai language capability
Pull 20 to 30 evaluation documents from real operations. It matters that these are real documents containing abbreviations and internal names, not tidied-up samples. Have a person write the expected output first, then compare it against the AI output. Skipping this and relying on a vendor statement that Thai is supported means discovering the accuracy shortfall only after rollout.
What should we do when the headquarters policy does not fit local practice
Rather than changing the policy, build a local mapping table. Map the document types handled daily at the site against the headquarters confidentiality classes, then have the headquarters IT department confirm it. With that table in place, local staff stop getting stuck on judgement calls.
Summary
Generative AI implementation failure happens as a design problem, not a model performance problem. MIT NANDA found around 95% of pilots stalled without producing measurable profit impact, with about 5% succeeding. RAND puts AI project failure above 80%, roughly twice the rate for conventional IT, and Gartner predicts 60% of projects without AI-ready data will be abandoned by 2026. Japanese domestic research found only one user in four had actually cut working hours.
The structure of failure breaks into four parts – vague objectives, the learning gap, misallocated budget, and unprepared data. At sites in Thailand and ASEAN, five more are added – headquarters policy that does not fit, unverified local language accuracy, misread literacy distribution, data split between Excel and paper, and missing success criteria.
Meanwhile AI diffusion in Thailand is fast, with meaningful usage rising from 9.1% in early 2025 to 12.4% in the first quarter of 2026 and ranking second worldwide for growth. The cost of deciding to wait is going up.
The direction to move is narrower, not wider. The successful 5% selected one problem, executed on it, and built partnerships intelligently. On procurement, purchased tools succeeded around 67% of the time against around 33% for internal builds. Filling in the target task, current effort, target, decision date, exit condition, data location and language of use with numbers and proper nouns before you start is the shortest route out of the 95%.
Even just working out where the first case should be and where that data currently sits will move the situation forward. If a decision is stuck between headquarters policy and local practice at your site in Thailand or Vietnam, or a pilot has stalled with no agreed next step, we are happy to talk at the evaluation stage before requirements are settled.
References
- MIT report – 95% of generative AI pilots at companies are failing – The 95% and 5% distribution from MIT NANDA, the learning gap, and success rates of 67% for purchased and 33% for internally built tools
- AI Project Failure Statistics 2026 – RAND Corporation on the failure rate above 80% and the comparison with conventional IT
- AI-Ready Data – Gartner on 60% abandonment by 2026, the shift in failure rate from 17% to 42%, and 72% of companies likely to halt proofs of concept
- AI adoption failure cases in Japan – The one in four result from a survey of 1,008 managers, vague objectives and undefined success metrics
- Thailand ranks second worldwide for AI adoption growth – Thailand figures and investment plans from the Microsoft Global AI Diffusion Report
- Thailand AI in manufacturing outlook – AI outlook for Thai manufacturing, 65% service level improvement in demand forecasting, 20% productivity gain, 53% downtime reduction from predictive maintenance