Blog

2026.08.08

Generative AI Data Analysis in Manufacturing | 4 Types, 3 Conditions

Generative AI Data Analysis in Manufacturing | 4 Types, 3 Conditions

Generative AI data analysis has still not reached the stage where you hand over a dataset and get back an answer you can act on. Even in 2026, only a small slice of the companies already using generative AI has managed to embed data analysis into daily operations. The reason is not that the models are not smart enough. Whether an analysis result is usable is decided on the data side, by three conditions that have nothing to do with which model you license. This article splits generative AI data analysis into four distinct types, sets out the three conditions that actually govern accuracy, and works through an original three-scenario cost estimate for a Japanese-owned factory in Thailand that produces a counter-intuitive result: the wider you spread the data preparation, the further away the payback moves.

The one generative AI use case that is not growing

Teikoku Databank surveyed 10,312 valid respondents between 17 and 31 March 2026. Of those companies, 34.5% said they were using generative AI in their business. That headline number gets quoted constantly. The breakdown underneath it does not.

Among the companies that said they were using generative AI, the most common use case was writing, summarising and proofreading text, at 45.1%. Information gathering followed at 21.8%, and idea generation at 11.0%. Data aggregation and analysis came in at 7.4%. That is one sixth of the leading use case, and barely ahead of programming assistance at 5.9%.

There is an important reading caveat here. That 7.4% is not 7.4% of all companies. It is the share of use cases among companies that reported using generative AI at all, and the base for that group is the 34.5%. Once you account for that, the proportion of all surveyed companies actually doing data aggregation and analysis with generative AI is smaller still.

So why is this the one area that will not grow? Not for lack of demand. In corporate planning and production control, the requests are more urgent than they are for text generation. Nobody enjoys losing days every month to assembling the same defect report. Everyone would like to see the defect trend daily rather than monthly. The demand is there. What is missing is a version that survives contact with the data. People tried it, and it did not hold up.

The way it fails is unusual, and that matters. If text generation produces bad output, you know the moment you read it. Data analysis does not work like that. Whether the number that came back is correct is invisible unless you already hold an independent way to verify it. And the sites most eager to hand analysis to AI are precisely the sites that lack that independent verification. That is the structural twist at the centre of the problem.

Generative AI data analysis splits into four types

Generative AI Data Analysis in Manufacturing | 4 Types, 3 Conditions - figure 1

The single phrase “using generative AI for data analysis” covers four jobs that have almost nothing in common in practice. They require different preparation, they fail in different ways, and they demand very different levels of investment. Most internal discussions go in circles because all four are being argued about at once.

TypeWhat it doesWhat it presupposesMain failure mode
1. Paste-and-ask analysisPaste a CSV or table each time and ask questionsAlmost nothingRedone every time, not reproducible
2. Connected analysisQuery a database or BI layer directlyVocabulary, granularity, boundaryA correct answer in an unusable shape
3. Agentic analysisRun autonomously across multiple sourcesType 2 plus permissions and audit trailIntermediate assumptions never verified
4. Report generationTurn a fixed set of figures into proseThe aggregation already existsBlurred line between fact and commentary

These four are not stages. You do not begin at Type 1 and graduate to Type 4, and that point is worth dwelling on. Type 4 is easier than Type 1, and Type 3 simply cannot function unless Type 2 is already stable. The right sequence is to decide which of the four your organisation actually needs, then work backwards to the preparation that type requires.

Type 1 — Paste-and-ask analysis is fast, but nothing survives

You paste a CSV or a spreadsheet extract into a generative AI tool and ask which process is generating the most defects. Preparation is close to zero and you can start today. In practice, when a company says it has “tried generative AI for data analysis”, this is almost always what it tried.

The value of this type is exploration. With no hypothesis in hand, scanning a dataset to notice that one process behaves differently from the rest is a legitimate and useful task. As a way of getting your bearings before building a pivot table, it is genuinely good.

The problem is that nothing survives. There is no record of who pasted which file, covering which range, on which date. Ask the same question next month and if the extraction filter on the pasted file shifted by a single day, you get a different answer, with no mechanism for deciding which one is right. Somebody quotes a figure in a management meeting, attributes it to the AI, and then cannot reproduce it afterwards. The usual ending is that an analyst rebuilds the whole thing by hand.

So Type 1 should be scoped deliberately as the place where hypotheses are born and nowhere else. Anything discovered here gets rebuilt in reproducible form, either through Type 2 or through your existing BI layer. Where that division of labour holds, Type 1 pays for itself easily. Where it does not, Type 1 is simply an activity that adds hours.

Type 2 — Connected analysis is the real prize, and the heaviest lift

Here you connect generative AI to a production control system or a quality database and let it translate natural language questions into SQL or an equivalent query, returning real figures from live data. When most people picture generative AI data analysis, this is the picture.

When it works, the payoff is large. It also demands by far the most preparation, and the preparation has nothing to do with model performance. What it needs is the three conditions covered in the next section: vocabulary, granularity and boundary. Connect without them and the AI will calmly, plausibly and confidently count a different population from the one you asked about.

The characteristic Type 2 failure is not that a wrong number appears. It is that a correct number appears in a shape you cannot use. That distinction is the heart of this article, so it gets its own section further down.

If your requirement also includes searching and summarising internal documents, work instructions or standards, that is a separate design problem from connecting to a database, and it needs its own retrieval architecture. The issues overlap substantially with those covered in our article on building RAG for factory knowledge. Connecting to numeric data and searching over documents look similar from the outside and require completely different groundwork.

Type 3 — Agentic analysis pays for autonomy with hidden assumptions

In this pattern the AI works across production records, quality inspection logs, purchasing systems and equipment uptime data, assembling its own sequence of steps and driving through to a conclusion. Vendor proposals in this space have multiplied since the start of 2026.

Technically it is an extension of Type 2. Operationally the risk changes in kind, not degree. An agent will always make judgement calls along the way. Drop rows with missing values. Convert because the units differ. Where a serial number appears twice, keep the newer record. A human analyst would footnote every one of those decisions. An agent’s final output shows none of them.

The result is that Type 3 fails in a specific shape: the conclusion looks reasonable, but nobody in the building can explain which population supports it. If Type 2 gives you a correct answer in an unusable shape, Type 3 gives you a correct answer nobody can verify.

If you intend to adopt Type 3, you need Type 2 running stably first, and on top of that you need permissions designed up front — which tables, within which scope — and an audit trail that records which intermediate queries were actually executed. Skip that order and, the first time something goes wrong, you will have no way to locate the cause.

Type 4 — Report generation is the most realistic and the most underestimated

The aggregation already exists in your BI tool or your spreadsheets. All you delegate to the AI is turning those figures into prose. Monthly quality reports, uptime summaries, executive briefing packs. This is Type 4.

A large share of the 45.1% who told Teikoku Databank they use generative AI for writing, summarising and proofreading almost certainly sits here. In other words, companies have already handed part of their data reporting to AI. They have simply done it under the condition that the AI does not perform the aggregation.

The Type 4 failure is a blurred line between fact and commentary. “The defect rate was 0.35%” is a fact. Unprompted, the AI will add “showing an improving trend versus last month”. That clause is not a result of aggregation. It is generated content. What you end up with is a report where the numbers are correct and the interpretation is invented.

The fix is simple. Split the template into fact paragraphs and assessment paragraphs, and require a named human author on every assessment paragraph. Deciding exactly where machine output stops and human ownership starts follows precisely the same logic as the transcription boundary discussed in our article on automating Excel work with AI.

Accuracy is decided by three conditions on the data side, not by the model

Generative AI Data Analysis in Manufacturing | 4 Types, 3 Conditions - figure 2

Whether Type 2 connected analysis succeeds is not determined by which model you pick. It is determined by whether the data you connect to satisfies three conditions.

Vocabulary — is it written down anywhere what NG_CD=07 means?

The first condition is whether column names, code values and abbreviations are defined in a form a human can read.

Suppose the defect code column is named NG_CD and a row holds the value 07. Whether that means a solder defect, a cosmetic scratch or a dimensional deviation is not discoverable by looking at the table. In most factories that lookup table lives in a veteran supervisor’s head, or on one sheet of a spreadsheet somebody built ten years ago. Without it, generative AI has no idea what 07 is.

The dangerous part is that the AI does not say it does not know. It infers meaning from the column name and returns a plausible-looking aggregation. It decides on its own whether a column called QTY means good units or input units, and computes a yield from that assumption. When the assumption is wrong, the output still looks entirely normal.

Building vocabulary means writing down three things. For each table, what every column represents. For every column holding coded values, the complete value-to-meaning mapping. And the expansion of every internal abbreviation that only makes sense inside your company, such as WIP, ADJ, RTN or SCRAP. Any area where you cannot write those three things down should be excluded from connected analysis, full stop.

Granularity — can you define what one row represents, including the exceptions?

The second condition is what a single row in the table actually represents. One unit, one lot, one shift, one inspection record? Tables that mix these are not rare in practice; they are normal.

The harder half is the exception definitions. If the following four are not defined, aggregation will drift, every time.

  • Duplicates: when a re-measured inspection produces two rows, do you count both, or only the newer one?
  • Cancellations: does a mis-entry correction arrive as a negative quantity row, or as a deletion of the original row?
  • Re-entry: when the same serial number flows through again after rework, do you count input as two or as one?
  • Missing periods: when equipment is stopped, does no row appear at all, or does a row appear with a value of zero?

Missing periods cause the most accidents. A row not existing and a value being zero mean different things. Ask generative AI to compute an average and the intervals with no rows drop out of the denominator. You wanted an uptime figure, and the only thing that disappeared from the calculation was the downtime.

How to structure records so that granularity is explicit is covered in our article on the three-layer design of a quality data management system. Ideally this is settled before generative AI enters the picture. In most real programmes the two run in parallel.

Boundary — over what period does the data mean the same thing?

The third condition is a boundary in time. From when to when has data been recorded under a single consistent definition?

Factory master data changes. A process was renamed in October 2025. The defect code scheme was consolidated and old codes 07 and 09 were folded into new code 12. The part numbering rule changed. Equipment IDs were reassigned when a line was added. These revisions are routine.

The problem is that the revision history is usually not retained. When master records are updated in place, historical data ends up being interpreted through today’s master. Ask generative AI to analyse defect trends over the past three years and you get a result that sums the old-code period and the new-code period under a single label. The trend appears to shift. Nothing on the shop floor shifted; the coding scheme did.

Building the boundary means holding three things: a history of master revision dates and their content, a mapping between old and new codes, and an explicit declaration written into the query or the prompt stating that this analysis covers data from a specified date onward only.

These three conditions do not substitute for one another. With vocabulary but no granularity definition, you count the wrong quantity under the right name. With granularity but no boundary, you count correctly across the wrong span of time. Only when all three are in place can connected analysis output be used for a business decision.

What benchmarks say about the limits on real databases

Raise the three conditions and someone will reply that surely this gets solved as models improve. Public benchmarks answer that reasonably well.

BIRD is the leading benchmark for natural language to SQL accuracy. It comprises 12,751 question and SQL pairs across 95 databases totalling 33.4GB and spanning more than 37 domains, deliberately built to resemble real, messy production databases rather than clean textbook schemas.

On the BIRD leaderboard, the best execution accuracy on the test set is 81.95%. On the dev set the figure is 77.64%, achieved by an AskData plus GPT-4o configuration. Human performance on the same task is recorded at 92.96.

That 81.95% needs careful reading. It is the highest benchmark score for natural language to SQL conversion against realistic databases. It is not a basis for claiming that generative AI analyses data with 80% accuracy. Benchmark questions are solved with the target schema and the relevant external knowledge already supplied to the model. In other words, the vocabulary and granularity described in the previous section are handed over for free by the benchmark itself. That is the condition under which 81.95% was achieved.

Move to settings that involve dialogue and real interaction and the numbers fall sharply. BIRD-Interact, the interactive variant of the same benchmark, tops out at 24.4%, and under some configurations sits in the 17.78% range. The closer the setup gets to what a human analyst actually does — asking clarifying questions, inspecting intermediate results, tightening conditions as they go — the worse the scores become.

The conclusion to draw is not that generative AI is unusable. It is that even with schema and vocabulary fully prepared, it misses close to one time in five, and without that preparation the conversation does not start. Model progress is not a substitute for the three conditions.

The failure mode where the answer is right but unusable

The failure that actually occurs on site is more troublesome than a wrong number.

AIMultiple published a 2026 evaluation comparing 36 models across 759 questions. Measured by strict execution match, meaning the returned result set must match exactly, the best model reached only 0.551. That looks low next to the BIRD leaderboard, and the gap is a difference in evaluation strictness rather than a contradiction.

The most operationally important finding from the same evaluation is this. Some models lose 22.7 points from projection differences alone. A projection difference means an extra column is present, the column order differs, or the aggregation unit is not the one expected. The requested number itself is correct. The shape of the returned table is not what was expected. That alone scores as incorrect.

On the shop floor, this is not treated as incorrect. It is treated in a worse way. A human receives the output, visually picks out the columns they need, reorders them, normalises the units, and then uses the result. Manual work has been created. The system installed to automate reporting has generated a recurring post-processing step. And because the person doing that post-processing believes the AI answered correctly, it never gets reported as a problem.

Here is what it looks like concretely in a factory. You ask for the top five processes by defect cost last month. Back comes a two-column table, process code and defect cost, five rows, sorted descending by cost. The numbers are right. The table cannot go into the meeting, because it shows process codes and no process names, and nobody attending the meeting has the code list memorised. So you ask again, this time requesting names. Now you get five columns — process name, process code, defect cost, defect count and period covered — and the row order has changed to process code order. The numbers are still right. But the column layout no longer matches last month’s deck, so the two cannot be placed side by side.

This round trip happens every month. It usually costs ten to fifteen minutes each time, which is why nobody escalates it. But if the project was justified on reporting hours saved, that round trip is being subtracted from the savings. The estimate later in this article deliberately puts Scenario A’s saving at a modest 12 hours per month precisely to absorb it.

Aggregation unit drift takes the same form. When you ask for the defect rate, whether the AI divides defect count by input quantity, defect cost by production value, or produces a defective-lot rate at inspection lot level is undetermined unless you specify it. All three are called the defect rate. All three are computed correctly. And naturally, the three do not agree. Building the vocabulary described earlier is, among other things, the work of eliminating that ambiguity.

The same evaluation also reports that 31.1% of BIRD’s gold queries — the SQL treated as the correct answer — contain problems. A CIDR 2026 paper analysing annotation errors in text-to-SQL benchmarks examines the same issue. The definition of correctness is itself unstable. That is an academic point, and it is also a practical one. Unless your organisation defines what correct means internally, correctness cannot be measured at all.

Three fixes you build in by design

To reduce correct-but-unusable answers, you do not change the model. You place constraints on the output. Three of them earn their keep in practice.

First, fix the output shape before you ask the question. Decide in advance which columns, in which order, at which aggregation unit. For frequently repeated questions, template it and instruct the AI to fill the template. Free-form output is reserved for exploration only.

Second, always install one reconciliation layer. Prepare a fixed query that checks against a known value. Pick a figure already confirmed through a separate channel, such as total input quantity for last month from the monthly closing report, have the AI derive the same figure from the population it used, and compare. If they do not match, discard the analysis. The presence or absence of this single check changes the safety of the whole operation.

Third, require the AI to output not a number but a number plus the definition of the range that produced it. Every output should append the period covered, the processes included, the conditions under which rows were excluded, and the version of the master data used. With that, anyone reading the number can judge whether the population was appropriate. Without it, the only person who can judge is the person who ran the query.

None of these three depends on model selection. All three are needed whichever model you use, and all three work whichever model you use. Put differently, running a model comparison before you have installed them is not a meaningful exercise.

Four more things that break at an ASEAN site

Everything above applies to a factory anywhere. Operate manufacturing sites across Thailand, Vietnam, Malaysia or Indonesia and four further factors stack on top.

Mixed languages is the first. Tables where column names are in English, category descriptions are in Japanese, and the free-text field operators actually type into is in Thai or Vietnamese are entirely normal. The same underlying event — a cosmetic defect — is recorded three or four different ways. Generative AI counts each variant as a separate thing. Before you can build an aggregation unit, you need a normalisation mapping.

Shift calendars and public holidays is the second. Public holiday calendars differ by country, are set annually by each government, and carry substitution days and one-off special holidays that shift year to year. Thailand is a clear case, and a company running plants in three countries has three calendars that do not line up. At a site running three shifts around the clock, the date boundary and the shift boundary do not coincide. When you calculate this month’s defect rate, whether the final shift of the month falls into this month or the next is a site-level rule. If that rule is not held as master data, the AI will cut on calendar days. Production records are cut on shift calendars. The two will always disagree. And because the disagreement is typically only a few percent, nobody spots it by eye.

Manually maintained master data is the third. Part number masters, process masters and defect code masters kept in one engineer’s spreadsheet and overwritten on each update are extremely common. In that state the boundary condition described earlier is impossible in principle. Before introducing generative AI, at minimum change the operating practice so that revision dates and revision content are appended to a separate sheet rather than overwritten.

The data source itself is the fourth. If equipment output reaches the system once a day via a paper shift report typed in by hand, daily tracking is not available at any price. That is a data frequency problem that sits upstream of generative AI entirely. The investment profile for capturing production data automatically from equipment is broken down in our article decomposing factory IoT cost into five layers.

These conditions show up in national statistics too. ETDA, Thailand’s Electronic Transactions Development Agency, reports in Thailand Digital Outlook 2026 that the digital maturity score of Thai enterprises stands at 2.12 out of 4, up from 1.56 the previous year, based on a survey of 834 organisations. Note carefully that 2.12 is a digital maturity score, not an AI adoption rate. It is improving, and it is only marginally above half of the four-point scale.

On AI adoption specifically, a 2024 ETDA survey of 580 Thai organisations found 18% had adopted AI and 73% were considering it. That survey identified concerns about data quality as a barrier to adoption. The substance of that finding lines up exactly with the three conditions set out above.

An original cost and payback estimate — three scenarios for a Japanese-owned factory in Thailand

Generative AI Data Analysis in Manufacturing | 4 Types, 3 Conditions - figure 3

From here the discussion turns to money. Every figure below is built up specifically for this article. None of it is a market rate. Substitute your own conditions as you read.

The assumptions are as follows.

  • A Japanese-owned factory in Thailand, around 300 employees, annual output of 1,200,000 units
  • Two production control staff each spend 20 hours per month producing the monthly defect and uptime reports, for a combined 40 hours per month
  • Staff salary of 25,000 THB per month against 160 working hours. The hourly rate is 25,000 divided by 160, which is 156.25, rounded down to 156 THB
  • The loss per defective unit, combining material, rework labour and shipment adjustment, is set at 180 THB
  • The current defect rate is 0.35%

One caveat about that 156 THB hourly rate before we proceed. It is a floor, not a justification for the investment. All it does is convert reporting hours into wages. In reality there is an opportunity cost — the improvement work those hours could have funded instead — which is very likely larger. But opportunity cost estimates vary widely, so this article deliberately uses only the floor. Read it this way: if a scenario pays back on wage conversion alone, that scenario is solid.

On those assumptions, here are three scenarios that differ only in how wide the data preparation goes.

ScenarioPreparation scopeInitial (THB)Annual running (THB)Annual benefit (THB)Net benefit (THB)Payback
A. Paste-and-ask company-wideVocabulary sheet and templates only60,00030,00022,464-7,536Never pays back
B. Connected analysis over all dataThree conditions across all processes450,000120,000160,41640,416About 11.1 years
C. Connected analysis on one themeOnly the top three defect processes180,00060,000181,152121,15217.8 months

The table alone does not explain why it comes out this way. The arithmetic is broken down scenario by scenario below, in a form you can follow on a calculator.

Scenario A — paste-and-ask analysis company-wide

Every department gets a generative AI account, and you prepare a vocabulary sheet and a set of prompt templates. Nothing is connected to a database. Each person pastes their own CSV extracts.

The initial cost of 60,000 THB covers writing the vocabulary sheet, preparing templates, and the internal briefing sessions. Annual running cost of 30,000 THB covers licences and maintaining the vocabulary sheet.

The benefit is calculated as follows. Of the 40 hours a month spent on reporting, paste-and-ask analysis is assumed to remove 12 hours. The data preparation for the aggregation remains manual, so what is actually saved is the writing and formatting portion.

12 hours per month multiplied by 12 months equals 144 hours. 144 hours multiplied by 156 THB equals 22,464 THB.

That is the annual benefit, and it sits below the annual running cost of 30,000 THB. Net benefit is 22,464 minus 30,000, which is -7,536 THB. Before there is any question of recovering the 60,000 THB initial outlay, the arrangement accumulates a loss every year.

To be clear, this is not an argument that Scenario A is worthless. For text generation and summarisation it pays for itself comfortably. The point of the estimate is that measured strictly against the data analysis use case, it runs at a loss. Part of the reason data aggregation and analysis is stuck at 7.4% in the Teikoku Databank figures quoted at the start of this article lies exactly here.

Scenario B — connected analysis across all data

You prepare vocabulary, granularity and boundary across the production control system, the quality database and the equipment uptime logs in full, and connect generative AI to all of it. This is close to what is usually pitched as a company-wide data platform.

The initial cost of 450,000 THB covers master data cleanup across every process, building the code mapping tables, defining granularity, establishing master revision history, and constructing the connection and the reconciliation layer. Annual running cost of 120,000 THB covers licences, platform operation, and the work of reflecting master revisions.

The benefit has two components.

First, hours saved. Of the 40 hours per month, 28 are assumed to be removed. Because the aggregation itself can now be pulled through connected analysis, the saving is substantially larger than in Scenario A.

28 hours per month multiplied by 12 months equals 336 hours. 336 hours multiplied by 156 THB equals 52,416 THB.

Second, defect reduction. With the data visible, the defect rate is assumed to fall from 0.35% to 0.30%. The difference is 0.05 points.

1,200,000 units multiplied by 0.0005 equals 600 units. 600 units multiplied by 180 THB equals 108,000 THB.

Combined, 52,416 plus 108,000 gives 160,416 THB of annual benefit. Subtracting the 120,000 THB annual running cost, 160,416 minus 120,000 leaves 40,416 THB of annual net benefit.

Payback on the initial cost is 450,000 divided by 40,416, which is approximately 11.1 years. That is not a number that clears any capital investment committee.

Scenario C — connected analysis limited to one theme

You narrow the scope to the top three processes by defect cost, prepare vocabulary, granularity and boundary for those three only, and connect. Every other process stays on Type 1 paste-and-ask for the time being.

The initial cost of 180,000 THB covers master data cleanup and connection for three processes, plus building one reconciliation layer. Annual running cost of 60,000 THB covers licences and maintaining master data for those three processes.

Hours saved come out lower than in Scenario B. With only three processes in scope, the reporting saving is assumed to be 16 hours per month.

16 hours per month multiplied by 12 months equals 192 hours. 192 hours multiplied by 156 THB equals 29,952 THB.

Against Scenario B’s 52,416 THB, that is a little under 60%. On this measure alone Scenario B still looks stronger.

Defect reduction reverses the ranking. Because scope is limited to three processes, they can be tracked daily. Somebody actually looks every morning, so the time between an anomaly appearing and corrective action starting shortens by an assumed two weeks on average. As a result the defect rate is assumed to fall from 0.35% to 0.28%. The difference is 0.07 points.

1,200,000 units multiplied by 0.0007 equals 840 units. 840 units multiplied by 180 THB equals 151,200 THB.

Combined, 29,952 plus 151,200 gives 181,152 THB. Subtracting the 60,000 THB annual running cost, 181,152 minus 60,000 leaves 121,152 THB of annual net benefit.

Payback is 180,000 divided by 121,152, or approximately 1.49 years, which expressed in months is 17.8 months.

The wider the preparation, the further away the payback

Line the three scenarios up and the result runs against intuition.

Scenario B saves the most hours yet takes 11.1 years to pay back, while Scenario C saves under 60% of B’s hours and pays back in 17.8 months.

Hours are not what creates the gap. On hours saved, B delivers 52,416 THB against C’s 29,952 THB, and B is genuinely ahead. The reversal comes from defect reduction, where B delivers 108,000 THB against C’s 151,200 THB.

Why does narrowing the scope produce more defect reduction? Because faster corrective action only materialises when the scope is narrow.

Making all process data visible does not increase the number of indicators a human being looks at every day. A dashboard refreshing 20 indicators daily becomes, in practice, something people glance at once a week. Checking the defect rate of three processes every morning, by contrast, is a habit that holds. Because it holds, anomalies are noticed sooner. Because they are noticed sooner, correction starts sooner. Because correction starts sooner, fewer defective units are produced. That chain only closes when the number of things being watched is small.

Put another way, the primary driver of benefit is not that data became visible but that behaviour changed. And whether behaviour changes is inversely correlated with how wide the preparation scope is.

Scenario B carries a second structural burden. A large share of its 120,000 THB annual running cost is the work of reflecting master data revisions. The more processes you have prepared, the more places need fixing every time a master record changes. So Scenario B has a benefit curve that flattens while its running cost grows in proportion to scope. That asymmetry is what produces the 11.1-year figure.

The investment decision for generative AI data analysis is therefore not a question of how far to prepare. It is a question of where you decide not to prepare. A project that cannot explicitly name the scope it is leaving alone will follow Scenario B’s path.

How to design the measurement itself — what counts as a benefit and what deliberately does not — is set out in our framework for measuring the effect of AI adoption. That is a useful reference when you substitute your own assumptions into the estimate above.

A decision table for choosing where to start

Everything above condenses into an actual decision, and the decision branches on only two questions.

The first question. Have you narrowed down to one indicator that is worth tracking daily? Narrowed down means it is already settled who does what when that indicator deteriorates. Indicators that people would merely like to see do not count.

The second question. Can you write down the vocabulary, granularity and boundary for that indicator today? Specifically: the meaning of every column that feeds the indicator, the code value mapping, what one row represents, how the exceptions are handled, and from which date onward the definition has been consistent. Can all of that be documented before the end of today?

Those two answers determine which type to begin with.

Indicator narrowed to one?Three conditions writable today?What to start
YesYesType 2 connected analysis, scoped to that indicator only
YesNoPrepare the three conditions first, and use Type 4 in parallel to cut reporting effort
NoYesNarrow down first. Use Type 1 to explore candidate indicators and pick one
NoNoType 4 and Type 1 only. Do not begin Type 2 or Type 3

Companies that land on the first row are not common. Most fall into the second or the fourth. That is not a failure; it is an accurate reading of where you actually stand. Commissioning Type 2 while the three conditions cannot be written down is how a programme enters the Scenario B path.

A note on how to proceed if you are in the second row. Preparing the three conditions across the whole company is exactly Scenario B. To avoid that, limit the unit of preparation to the tables that your one narrowed indicator actually reads. In practice, list only the columns used directly in computing that indicator, and build meaning and code mappings for those columns alone. Leave the other columns in the same table untouched. Whether your organisation can accept that degree of pragmatism is what separates 17.8 months from 11.1 years.

Also, keep the output of that preparation as documentation. If it lives only inside a generative AI prompt, it disappears the moment you change models or switch tools. The vocabulary sheet, the granularity definition and the boundary declaration are assets independent of any tool. Tools turn over every few years. These three do not. The substance of the investment is here, not in the tool.

Type 3 agentic analysis does not appear anywhere in that table. It is something to evaluate after Type 2 runs stably, the reconciliation layer functions, and permissions and audit design are complete. If you are being pitched Type 3 today, make a proven Type 2 track record a precondition.

One further boundary. Analysis with generative AI and forecasting with statistical or machine learning models are different disciplines. Predicting future values is covered in our article on demand forecasting with AI, where both the data requirements and the evaluation methods differ from everything in this article. Conflating the two breaks requirements definition.

Frequently asked questions

What data do we need in order to have generative AI analyse it?

Not volume, definitions. Three things specifically: a description of what each column in the target tables represents, a value-to-meaning mapping for every column holding coded values, and a definition of what one row represents together with how exceptions are handled, covering duplicates, cancellations, re-entries and missing periods. On top of that, set the boundary date from which the data has been recorded under a consistent definition. Any scope where you cannot write those four things down is not a candidate for connected analysis regardless of how much data sits there. Conversely, where those four exist, connected analysis works even if the scope is only three processes. The first step is not acquiring more data, it is documenting the definitions of the data you already hold.

Will generative AI data analysis replace spreadsheets?

It will not replace the aggregation, but it will replace what sits on either side of it. Before the aggregation is exploration: getting a feel for which angles are worth examining is faster than rebuilding pivot tables. After the aggregation is prose: turning confirmed figures into a report is delegable once the template is fixed. The aggregation itself, which produces the numbers people commit to, demands reproducibility and verifiability, so it is safest left in your existing systems. Scenario C in this article is built on that basis too, layering querying and prose generation on top of the aggregation platform rather than discarding it.

How much of report writing can we hand over?

Statements of fact, and no further. Keep assessment and judgement separate. “The defect rate was 0.35%” is a transcription of an aggregation result and can be delegated. “Showing an improving trend versus last month” is generated content. Even with an identical movement in the numbers, whether that movement counts as improvement depends on the target and the context, and the AI does not know either. The practical measure is to split the template into fact paragraphs and assessment paragraphs and require the author’s name on the assessment paragraphs. One more point: always require the output to append the period covered, the processes included, the exclusion conditions, and the master data version used. A report without those cannot be verified by whoever reads it.

Should we introduce a BI tool or generative AI first?

The sequence is the wrong frame. What matters is that both require the same preparation. Neither a BI tool nor generative AI will produce correct figures without vocabulary, granularity and boundary defined. The difference is in how errors surface. BI reproduces the same wrong definition identically every time, so eventually somebody notices. Generative AI has room to interpret each question slightly differently, so the error does not reproduce and detection is delayed. Given that, a realistic division is to keep committed figures in BI or your existing reporting, and put generative AI on exploration and prose. Before deciding what to buy, start writing down the three conditions. That work is required either way.

Can we analyse data that mixes Thai, Vietnamese and English?

You can, but you need a normalisation mapping first. The practical obstacle is not whether the model can read Thai or Vietnamese. It is that the same event is recorded in three or four different languages and gets aggregated as three or four different things. Build a mapping that collapses notation variants onto one canonical value, at minimum for defect categories, downtime reasons and process names. Avoid aggregating free-text fields directly; adding a coded category column with a fixed pick list turns out to be the faster route. On top of that, hold shift calendars and public holiday calendars as master data, per country. Public holidays are announced annually through official gazettes and carry substitution days, so substituting plain calendar dates will put your aggregation period out of step with your production records.

Summary

The points made in this article, in order.

  • Even among companies using generative AI in their business, data aggregation and analysis accounts for only 7.4% of use cases, per Teikoku Databank, March 2026, 10,312 valid responses. That is one sixth of the leading use case, writing, summarising and proofreading, at 45.1%. The 7.4% is a share of use cases among adopting companies, not a proportion of all companies
  • The reason it is not growing is not weak demand. People tried it and it did not hold up. Unlike text generation, a failed data analysis is invisible on inspection
  • Generative AI data analysis splits into four types: paste-and-ask, connected, agentic, and report generation. They differ in required preparation and in failure mode, so debating them as one topic goes nowhere
  • Accuracy is governed by three conditions on the data side, not by the model. Vocabulary, meaning whether NG_CD=07 is documented anywhere. Granularity, meaning what one row represents and how duplicates, cancellations, re-entries and missing periods are treated. Boundary, meaning over what period the definition is consistent and whether master revision history survives
  • Natural language to SQL against realistic databases tops out at 81.95% on the BIRD test set, short of the 92.96 recorded for human performance. On BIRD-Interact, the setting that involves dialogue and real interaction, the best result is 24.4% and some configurations sit in the 17.78% range
  • What happens on site is less often a wrong number than a correct number in an unusable shape. Some models lose 22.7 points from projection differences alone, meaning extra columns, different column order or a different aggregation unit, per AIMultiple, 2026, 36 models across 759 questions
  • The countermeasures are to fix the output shape in advance, install one reconciliation layer, and require the definition of the population alongside every number. All three work independently of model selection
  • ASEAN sites add four further problems: mixed languages, shift calendars and public holidays, manually overwritten master data, and the frequency at which data reaches the system. ETDA’s Thailand Digital Outlook 2026 reports Thai enterprise digital maturity at 2.12 out of 4, up from 1.56, across 834 organisations. That is a maturity score, not an AI adoption rate
  • In the three-scenario estimate, connecting all data pays back in 11.1 years while narrowing to one theme pays back in 17.8 months. The gap is created not by hours saved but by the defects avoided through faster correction, and that effect only grows large when scope is narrow. Every figure is built up for this article and is not a market rate
  • The investment decision is therefore not how far to prepare but where you decide not to prepare. A project that cannot name the scope it is leaving alone will take the long payback path

Generative AI data analysis does not fail because model performance is insufficient. It fails because the meaning of the data being handed over has never been written down, not even for internal readers. And writing it down company-wide pushes the payback further away. Narrow it to a single indicator and 17.8 months comes into view.

TOMAS TECH builds production control and quality data foundations for manufacturers operating across Thailand and the wider ASEAN region. We can help at the stage before any connection is discussed: taking stock of whether you can narrow to one indicator worth tracking daily, and whether the vocabulary, granularity and boundary for that indicator can be written down today. It is entirely fine to start before any product has been selected, or while it is still unclear which data is usable at all. If you want to check whether your own data can withstand connected analysis, get in touch through our contact page.

References