“If that person leaves, this process stops.” At Japanese-owned plants in Thailand it is the problem everybody half-knows about and nobody has actually taken on. Then generative AI arrived, and the phrase skill transfer AI suddenly started to sound like something you could buy. What has also arrived, in roughly equal numbers, are the failures: the videos filmed and dropped on a server, the manual project that stalled the moment a veteran was told to write it. This article is not a general argument about whether AI can or cannot transfer skill. It is about a separation: splitting tacit knowledge into three tiers, deciding how much of it goes onto AI, and where the boundary is beyond which the answer is training design rather than technology.
Why skill transfer AI projects fail
Failure pattern 1: film the veteran, drop the file on a server
This is by far the most common. Somebody records the veteran’s work on video and saves it to a shared folder, a NAS, or cloud storage. There is nothing wrong with that in itself. The problem is that nothing happens afterwards.
Six months later, what exists is a folder full of filenames along the lines of 20260312_grinding_kato.mp4. Each clip runs thirty minutes to an hour. When a junior operator wonders “what do I do when deburring leaves the surface rough?”, there is no way to know which file, at which minute and second, holds the answer. So they turn round and ask the veteran standing next to them. Nothing has changed from the day before the camera was switched on.
The reason video goes unused is not mysterious. Video is not a searchable format. Information that has not been cut to a granularity a person can search will not be accessed, however much of it accumulates. And there is a second problem: video captures almost none of the *why*. It shows the hands. It does not show what the person was looking at, what they were comparing it against, or how they decided. Veterans do those things unconsciously, so they do not say them out loud either. You end up with a high-fidelity recording of the surface of the work and nothing at all of the reasoning underneath it.
Failure pattern 2: ask the veterans to write it up, and watch it stall
The next most common pattern is the decision to “properly document everything,” followed by handing the writing assignment to the veterans themselves. This stops almost without exception. There are three reasons.
First, writing is not their job. Structuring knowledge into prose is itself a skill. Somebody who has stood on a line for forty years is not necessarily able to produce a hierarchical work instruction in a word processor, and asking them to is asking them to be bad at something in front of colleagues.
Second, nothing has been taken off their plate. The more senior the person, the more often they are pulled into firefighting on the floor. A writing task with no deadline attached will always lose to a machine that has stopped, every time.
Third — and this is the real wall — they are not aware of what they know. Judgement that has become automatic feels obvious to the person making it, and obvious things do not register as information worth writing down. What actually comes back, when something comes back at all, is a document that reads almost identically to the work standard you already had: correct, complete, and entirely free of the knowledge you were trying to capture.
Failure pattern 3: assume that installing AI will take care of it
And then there is the pattern that has multiplied through 2026. Generative AI can read internal documents now, the thinking goes, so if we point it at everything we have, skill transfer follows.
What this misses is the plain fact that the knowledge you want to keep was never written down in the first place. AI can read what exists. Decision criteria that live only inside a veteran’s head cannot be retrieved by any model, however capable. Index your existing folders and make them searchable and what surfaces is a work instruction that was superseded ten years ago, somebody’s personal notes, and an unapproved draft. That is worse than useless: you have built a system that returns wrong information confidently, with a document reference attached to make it look authoritative.
What the three have in common: nobody separated out what to keep
The thread running through all three is that none of them is a problem with AI performance. Each of them comes from failing to sort the knowledge worth keeping into types. Procedures, decision criteria, and physical feel all get lumped together under “tacit knowledge” and addressed as one thing. But those three differ completely in the work required to capture them, in how far capture is even possible, and in what they cost.
The difficulty shows up plainly in the statistics. Commentary on Japan’s 2026 White Paper on Manufacturing Industries (METI) reports that, among the obstacles to skill transfer, “difficulty in converting veterans’ knowledge and experience into explicit knowledge” was cited by 68.6%. Alongside it sit “we do not know how to make it explicit” at 28.6% and a shortage of management resources at 28.1%. In other words, a large share of manufacturers are in the position of knowing exactly what has to be done and having no idea how to do it.
The same commentary reports that 71.8% of firms that have implemented data integration are not using AI. Data has begun to connect, and the step from there into AI has not been taken. The gap in the middle is precisely the missing separation. Between “we have the data” and “we have deployed AI” sits a design step nobody scheduled: deciding which knowledge, in which form, for whom.
Split tacit knowledge into three tiers
The three-tier model
The first task in designing skill transfer AI is classification. The split that works in practice is this one.
| Tier | What it covers | Concrete examples | Can it be made explicit? | How it goes onto AI |
|---|---|---|---|---|
| Tier 1: Procedure | “What to do, in what order, at what settings.” Already articulated, or articulable immediately | Changeover steps, inspection items, setting values, jig mounting, start-up and shutdown sequences | High. Most of it already exists as documents — though usually out of date | Chunk the instructions and videos and make them searchable. Multilingual versions pay off here |
| Tier 2: Decision criteria | “Under which conditions do you choose which option.” Held in the person’s head, but retrievable if you ask properly | Whether to stop or continue when the sound changes, correcting for season and humidity, the order in which faults are isolated, where the line falls between outsourcing a repair and doing it in-house | Moderate. Requires elicitation through interviews (externalisation) | Structure it as Q&A and decision trees and make it retrievable through RAG. This is the core of skill transfer AI |
| Tier 3: Feel and embodied skill | “How hard to press,” “how the sound differs,” “how the smell changes” — discrimination learned through the body | Grinding pressure, torch angle and travel speed in welding, distinguishing abnormal noise, the smell of scorching resin | Low. In principle it does not transmit through documents or data | Do not try to put it onto AI. Route it to training design: OJT, hands-on practice, limit samples |
Why the split is necessary
Because without it you cannot make an investment decision.
Tier 1 has a clear cost-benefit profile. Take stock of existing documents and update them, cut the video, produce the other language versions. You can estimate the volume of work and you can measure whether it was done. But fixing Tier 1 alone does not solve the skill transfer problem. The situation stopping your junior staff is rarely “I do not know the procedure.” It is “I followed the procedure and it did not come out right.”
Tier 2 costs more, because there is an elicitation step in front of it — but this is where the return is. Once “why did you choose that just now” has been captured, the range of situations a junior can resolve without walking over to a veteran expands. In practical terms, the value of an investment in skill transfer AI is decided almost entirely by how thick you can make Tier 2.
Tier 3 cannot be captured, and accepting that is the single most important design decision in the project. If you let marketing language about “reproducing the craftsman’s touch with AI” pull Tier 3 into scope, the project will inflate and then stall. Tier 3 is absolutely a target for transfer — it is not a target for AI. It belongs to training curricula, limit samples and practical assessment.
Giving up on Tier 3 is what protects Tiers 1 and 2
When this is explained on the floor, a common reaction is: “so the part that actually matters is the part we can’t keep.” That is half right and half a misreading.
In most processes, what is actually holding a junior up is not Tier 3 but Tier 2. Telling by ear whether a sound is abnormal or normal is Tier 3. Knowing what to check first, and in what order, *once you have established that it is abnormal*, is Tier 2. If the latter survives, the junior can get part of the way on their own and call for help at the point they get stuck. The objective is to reduce the number of times a veteran gets called, not to reduce it to zero. A design that aims at zero is a design that will fail, and a shop floor where nobody ever needs a second opinion is not a shop floor anybody would want.
Separating Tier 3 out has a second benefit: it settles the priorities for training. Once you can say “in this process, there are exactly three things that must be learned through the body,” you can concentrate practical training time on those three. Training periods stretch to years precisely because Tiers 1, 2 and 3 are all being taught through the same channel — a senior standing next to a junior, saying everything out loud.
Breaking Tier 2 down further
Tier 2 covers a lot of ground, so one more level of division makes it easier to design against.
| Sub-category of Tier 2 | The shape of the question | How to elicit it | How to store it |
|---|---|---|---|
| Conditional branching | “What do you do when it’s A? When it’s B?” | Walk through past work records and confirm the branch points one at a time | Decision trees, condition tables |
| Exception handling | “When does following the procedure not work?” | Work backwards from trouble that actually happened | Case library (symptom → candidate causes → order of checks) |
| Prohibitions | “What do you never, ever do?” | Always push through to “why don’t you do it?” | Prohibition list, each item with its reason |
| Rules of thumb | “It isn’t in the textbook, but experience says do it this way” | Ask directly: “where do the manual and reality diverge?” | Attach as supplementary notes against the relevant procedure step |
These four categories map directly onto interview question design, which is covered further down.

Why skill transfer is more urgent at Japanese-owned plants in Thailand
Expatriate managers rotate every three to five years
In Japan, skill transfer is discussed on the timeline of retirement: veterans reach the end of their working lives, and the clock runs from there. At a Japanese-owned plant in Thailand, a second clock runs on top of it — the assignment length of the Japanese expatriate managers.
At most Japanese companies an overseas posting runs somewhere between three and five years. Over that period the expatriate learns the local equipment, the product mix, the customers, and the quirks of each supplier, and builds a working set of decision criteria on top of that learning. “This lot goes to a Japanese customer, so we stop it at this threshold.” “This machine behaves differently once the rainy season starts.” “Lot-to-lot variation from this supplier is wide, so we run heavier incoming inspection.” None of that is written in the standards issued by head office in Japan. It was created here, it is used here, and it is tacit knowledge in the fullest sense of the term.
And when the assignment ends, all of it disappears at once. The successor arrives from Japan and rebuilds it from scratch. Plants that have been repeating this cycle every three to five years are not unusual — they are the norm. If skill transfer in Japan is the problem of “forty years of accumulation lost once,” the Japanese-owned plant in Thailand has the problem of “three to five years of accumulation lost repeatedly.” Measured by total volume lost, the second can be the larger of the two.
The case for looking at skill transfer AI in Thailand rests on this double structure. The scope is not only the skilled trades on the floor; it includes the decision criteria held by expatriate managers. That is a consideration which simply does not appear in the domestic Japanese version of this discussion, and it is one of the reasons a design copied straight from a Japanese parent-company project tends to miss here.
The audience is Thai staff — multilingual is a requirement, not an option
There is a second condition, specific to Thailand and decisive. The people who will use the knowledge you capture are not Japanese.
Work instructions written in Japanese. Videos with Japanese narration. Q&A structured in Japanese. To the person who created them, they look complete. But the people who actually need them are Thai operators, technicians and supervisors. Material that stays in Japanese might as well not exist as far as they are concerned.
The point most often underestimated here is that translation and localisation are not the same thing. Run a Japanese work instruction through machine translation and you will get Thai. What you will not get is the vocabulary actually used on that shop floor — jig names, the local nickname for a machine, in-house abbreviations — and the result generates more confusion than the original. What is required is to build the glossary first, fix those terms, and translate on top of a fixed vocabulary.
This is, in fact, one of the places generative AI genuinely helps in a skill transfer context. The cost of producing and maintaining multiple language versions has fallen far enough that “write it in Japanese and translate it later” can be replaced by “operate in Japanese, Thai and English from day one.” What has not been automated is terminology control and review by the people on the floor. Skip those and low-quality translation destroys trust in the system, after which the system stops being used regardless of how good the underlying content is. The broader question of how to make something actually take root locally overlaps heavily with the ground covered in four reasons generative AI training fails to stick.
Thailand’s squeeze on mid-level skilled staff
Behind all of this sits the structure of the local labour market. In Thai manufacturing, the difficulty of securing mid-level skilled staff — technicians, supervisors, maintenance personnel — is a persistently reported problem. Fresh engineering graduates can be hired and inexperienced operators can be hired; the thin layer is the one in between, people with enough accumulated experience to make their own decisions.
For scale, Thailand’s manufacturing workforce stands at approximately 6.24 million people (as of December 2024). The Board of Investment (BOI) has been reported to have set out support measures for upskilling on the order of 5 billion baht covering roughly 100,000 people, which indicates that the shortage of skilled personnel is recognised as a national-level issue rather than a company-level inconvenience.
Seen from inside a plant, that squeeze presents itself as a sequence. A mid-level person is poached and leaves. No replacement of equivalent experience can be recruited. A junior is promoted into the gap out of necessity, cannot yet make the judgement calls, and comes to the Japanese manager to ask. The Japanese manager’s load rises. And a few years later, that Japanese manager returns to Japan.
Skill transfer AI is a way of breaking that chain at one of its links. It is not a tool that resolves the whole sequence. It works when it is pointed at a limited but genuinely useful objective: reducing the number of times somebody has to walk over and ask for a decision.
The elicitation step — do not make veterans write; interview them and structure it for them
Externalisation is listening work, not writing work
How do you get Tier 2 — the decision criteria — out? This is the operational heart of skill transfer AI.
As set out above, the write-it-yourself model breaks down. The alternative is to interview, and have the listener do the structuring. The person who holds the knowledge only has to talk; the documentation burden sits with somebody else. Whether that division of labour is actually established is the single strongest predictor of whether the project moves.
Conceptually this is the “externalisation” step of the SECI model, the classic knowledge management framework — the stage where tacit knowledge is converted into explicit knowledge. The theory has been around since the 1990s. What has kept it stuck in practice is a headcount problem: who does the interviewing, and who does the writing up. That is exactly where generative AI now changes the arithmetic.
Putting AI on the asking side — NTT DATA’s interview agent
As a concrete reference point, NTT DATA has published a proof of concept on the transfer of tacit knowledge. Targeting design work at Kawasaki Heavy Industries, the approach uses an “interview agent” that draws knowledge out of experienced engineers, supporting the externalisation process of the SECI model.
The significant thing is that AI is placed not on the answering side but on the asking side. Historically, eliciting tacit knowledge required an interviewer who understood the work; without the right questions, an expert never says their decision criteria out loud, because to them there is nothing to say. Putting AI in the questioner’s seat loosens that constraint on interviewer availability. The expert can work at their own pace, in dialogue, answering as they go.
The implication for how you design a skill transfer AI programme is substantial. AI’s first job is not to answer. It is to ask. Building a searchable knowledge base comes afterwards — and a knowledge base built before the asking has happened is a search index over documents that do not contain the answer.
Question patterns — what to ask to make decision criteria surface
Interviews live or die on question design. Ask “could you tell me the knack of this process?” and you will get “well, you get used to it,” and the session is effectively over. Getting decision criteria out requires patterns. Four of them work reliably in practice.
| Question pattern | How to actually phrase it | What it surfaces | The follow-up to chase |
|---|---|---|---|
| The branch point | “You stopped there and looked at something. What were you looking at, and what did you conclude?” | Conditional branching rules; what the person treats as an object of observation | “And if it had been the other way round?” — fills in the opposite side of the branch |
| The exception | “When does doing it exactly as the instruction says fail to work?” | Practical knowledge that lives outside the standard | “How often does that happen — a few times a year?” — pins down frequency |
| The near miss | “What’s the closest you’ve come to a serious problem, or the worst rework you’ve had?” | The precursors to failure, and the avoidance habits learned afterwards | “Was there any warning sign beforehand?” — produces early-detection points |
| The thing you never do | “What do new people tend to do that you would never do?” | Prohibitions, and the reasons behind them | Always ask “why not?” A prohibition without a reason decays into a rule nobody follows |
Of these, asking what somebody never does is disproportionately effective. People do not explain what they do, because it is too obvious to mention. What they avoid, by contrast, is held consciously, and it comes out with surprising fluency. Better still, an avoidance is almost always attached to a reason grounded in a specific past failure — which means one question yields a prohibition, a rationale, and a case study at the same time.
The other technique that matters is to ask on the floor, while the work is happening. Interview somebody in a meeting room and you get the generalisations they can recall in a meeting room. Stand beside them while their hands are moving and ask “what did you just look at?” and judgements they were not conscious of get put into words for the first time. Reviewing video together works well too: “you paused for a moment there — what were you checking?” That single question pulls part of what looked like Tier 3 down into Tier 2, which is the most valuable move available in the whole exercise.
Designing and running the interviews
As a practical sequence, the following slicing is easy to work with.
- Narrow the target processes. Do not start with everything. Pick roughly three processes that would stop if a specific individual left. Select on attribution risk and business impact.
- Read the existing documents first. Go into the interview having read the work instructions, standards and defect reports that already exist. Asking about things that are already written down wastes the session and damages the veteran’s willingness to continue.
- Run 30 to 60 minutes at a time, across several sessions. A single long interview produces fatigue rather than depth. Splitting it lets you open each session by confirming the previous one and then going deeper.
- Structure it on the spot and have the person confirm it. A raw transcript is not usable. Organise it into condition tables and case entries, then go back and ask “have I understood this correctly?” That confirmation step is what determines the quality of the output — it is where misinterpretations get caught, and it is also where the veteran sees their own knowledge take a shape they recognise.
- Have Thai staff sit in. Involving the eventual audience from the start means the places where an explanation does not land get identified in the room, before translation. You fix the granularity while it is still cheap to fix.
The people to interview are not only the skilled trades
At a Japanese-owned plant in Thailand, make sure Japanese expatriate managers are on the interview list. As described above, that is where decision criteria are sitting that will vanish with the assignment.
For expatriates, the question patterns shift slightly. “Is this judgement a head-office standard, or a rule you created locally?” “Why does the local threshold differ from the standard?” “Can you distinguish between what you inherited from your predecessor and what you decided yourself?” Those three questions alone produce something of real value as handover material — and they surface the uncomfortable but useful category of rules nobody now remembers the reason for.
And the expatriate’s decision criteria are exactly the material that needs to exist in more than one language. Capturing them in a form that local management can read, not just the next Japanese arrival, is what stops the shop floor being reset every time the posting rotates.

How to build skill transfer AI — implementation tier by tier
Tier 1 implementation: organise the procedures and produce the language versions
Tier 1 is weighted far more towards clean-up work than towards AI. The tasks are these.
- Take stock. List where each work instruction lives and which version it is. Shared folders, individual PCs, paper, notices on the wall — all of it counts.
- Discard and consolidate. Remove superseded versions, unapproved drafts and duplicates. Skip this and you have built a machine that returns obsolete information with a citation attached.
- Standardise the granularity. Decide whether a file corresponds to one process or one line, and apply it consistently. This translates directly into how well retrieval performs.
- Cut and tag the video. Hour-long clips do not get watched. Split by work step and attach a text description of what each segment shows.
- Produce the language versions. Fix the glossary first, then translate. Build review by local staff into the process rather than treating it as an optional check.
None of this is glamorous, but the bulk of a skill transfer AI system’s accuracy is determined here. However sophisticated the retrieval layer, if the underlying documents are stale or contradictory the output will not be trusted — and trust, once lost on a shop floor, is not recovered by a model upgrade.
Tier 2 implementation: turn decision criteria into Q&A and cases, and put them into RAG
Decision criteria as elicited in interview are not in a form suited to retrieval. They have to be shaped.
| The raw form | The shaped form | The question it will be searched with |
|---|---|---|
| “On humid days I raise the temperature a bit” | Condition table: when humidity exceeds X%, raise the set temperature by Y°C. Record the rationale and the limits alongside it | “Why do defects increase in the rainy season?” |
| “If there’s an odd noise, suspect the bearing first” | Case entry: symptom → candidate causes in order of likelihood → order of checks → decision criteria | “What is this noise?” |
| “For this material we run heavier incoming inspection” | Caution and prohibition entry: which supplier, under what conditions, which additional inspection items, and why | “Can we skip incoming inspection?” |
| “We got badly burned by this once” | Failure case: the situation → what happened → the warning signs → the countermeasure adopted afterwards | “Is it safe to keep running in this state?” |
Making knowledge organised this way searchable and citable across internal documents is what RAG (retrieval-augmented generation) is for. The user asks in natural language, the system retrieves the relevant internal documents, and the answer is generated on the basis of their content. What makes RAG a good fit for skill transfer specifically is that it can show the source document behind the answer. “The AI said so” carries no weight on a shop floor. “It is written in this entry of the case library” does. How to approach the build and how to think about the cost are set out in detail in what RAG implementation costs and how to approach it.
One caveat: RAG is not a universal solution. It cannot return what is not written, and if several contradictory documents exist it will produce contradictory answers. That is precisely why the Tier 1 stock-take and the Tier 2 structuring are prerequisites rather than nice-to-haves. Invert the order — “let’s get RAG in first and sort the content out afterwards” — and the first month of underwhelming answers costs you the shop floor’s confidence, which is far more expensive to rebuild than it was to keep.
Examples of what exists today
In the skill transfer and knowledge utilisation space, concrete services have begun to appear in Japan. Three of a deliberately different character are worth noting as reference points.
| Initiative | Provider | What it is | What is instructive about it |
|---|---|---|---|
| Tacit knowledge transfer PoC (interview agent) | NTT DATA | A proof of concept targeting design work at Kawasaki Heavy Industries, in which AI interviews experienced engineers to draw out knowledge, supporting the externalisation stage of the SECI model | The design philosophy of placing AI on the asking side. Directly relevant to the Tier 2 elicitation step |
| Manufacturing Shop-Floor Knowledge AI | Institute for Language Understanding | A service aimed at knowledge utilisation on the manufacturing floor, reported as becoming available on 1 July 2026 | Knowledge utilisation services specialised for manufacturing are now reaching the market as products rather than projects |
| Takumi AI | Mitsubishi Research Institute | A service positioned around supporting the transfer of expert insight and know-how (confirm the detailed functionality and applicable scope with the provider) | The transfer of expert skill has become a problem definition around which a commercial service can be built |
What these share is that each supports part of the process of eliciting, organising and using knowledge. Whichever you use as a reference, the work of deciding what your own company needs to keep is not something any of them replaces. Start from product selection without accepting that, and you spend the available time comparing tools and then deploy one over an empty content set — which is the most expensive way to arrive at the conclusion that AI did not work.
Tier 3 implementation: route it to training design, not to AI
Tier 3 — feel and embodied skill — is out of scope for skill transfer AI. That does not mean it can be ignored. If anything, narrowing what goes onto AI is what frees up resource to direct at training.
Realistic measures against Tier 3 look like this.
- Build a set of limit samples. Keep physical good and reject examples. In many cases it matters that they are physical objects rather than photographs.
- Design the practical training menu. Define what has to be practised, how many repetitions, and over what period, in order to reach competence.
- Make the assessment criteria explicit. Decide who judges “competent” and on what basis. Left vague, training has no end state and simply continues indefinitely.
- Consider substituting measurement. Examine whether discrimination currently done by human senses can be replaced by sensors or inspection equipment. Note that this stops being a skill transfer conversation and becomes a capital expenditure conversation, with a different approval path.
- Narrow what OJT covers. Once Tiers 1 and 2 can be retrieved through the system, OJT time can be concentrated on Tier 3.
That last item is the biggest practical dividend of splitting into three tiers. Traditional OJT depended on a senior person’s spoken explanation for everything — the procedure, the decision criteria, and the physical training alike. Move the first two onto the system and the scarce time of your most experienced people goes to the only part that genuinely requires a human being standing there.
The five layers that move the cost of skill transfer AI
“One lump-sum figure” tells you nothing
A quotation for skill transfer AI is meaningless if you look only at the AI tool licence, because most of the actual cost sits outside the AI. This article does not give specific amounts: the order of magnitude changes with the number of target processes, the state of the existing documents, and the number of languages, so there is no figure that can honestly be stated as typical. What is more useful is the set of variables that move the cost, separated into layers. Request quotations broken down along these layers and you gain the ability to compare vendors at all.
| Layer | What it covers | The main variables that move the amount |
|---|---|---|
| 1. Stock-take and interview design | Selecting target processes, taking stock of existing documents, designing and conducting interviews, structuring the output | Number of target processes; number of interviewees and how much of their time can be committed; how disordered the existing documents are; whether interpreters are needed; whether interviews are run in-house or outsourced |
| 2. Content production | Creating and updating work instructions, filming and splitting and tagging video, converting decision criteria into case libraries, producing language versions | Number and type of content items produced; whether filming is required; number of languages; whether a glossary exists; number of review rounds with the floor |
| 3. Retrieval platform (RAG) | Document ingestion, chunking design, retrieval and generation implementation, UI, permissions design | Number of documents and diversity of formats (mix of PDF, image, video); expected number of users; integration with existing systems; on-premises or cloud; the required standard of answer quality |
| 4. Operation and updating | Content updates, additional interviews, accuracy checks, usage review, licence and usage fees | Update frequency; whether updating is done in-house or outsourced; model usage volume; maintaining the language versions; scope of support for user queries |
| 5. Training design | Tier 3 training menu, limit samples, assessment criteria, redesign of OJT | Number of target processes; who delivers the training; developing assessors; degree of integration with the existing training system |
Where the cost actually concentrates
In most cases the centre of gravity is in layers 1 and 2 — not the tool, but the people who produce the content. That may be counter-intuitive, and it is also inevitable. The knowledge a skill transfer AI handles is specific to your company, which means there is nothing available to buy off the shelf. Every usable sentence in the system has to be produced by somebody who knows your process.
The corollary is that a quotation which goes light on layers 1 and 2 will produce correspondingly little. A proposal reading “simply ingest your existing documents and it is complete” looks cheap because the most labour-intensive stage has been left out of it. When you compare quotations, the informative comparison is not the total but how much effort is loaded into layers 1 and 2.
For layer 3, the structure differs depending on whether you use a general-purpose generative AI service as-is or build a retrieval platform over internal documents. Layer 4 is the one most often overlooked, and a knowledge base that is not maintained stops being used before long. Equipment changes, product mix changes, procedures change — and if only the written content stays still, the floor stops consulting it, quietly and without anybody reporting it. Comparing initial build quotations while leaving the maintenance arrangement undecided is the classic failure mode.
How to phase a small start
If you are going to stage it, the following slicing is workable in practice.
Phase 1 (one process; Tier 1 plus part of Tier 2). Pick the single process with the highest attribution risk, take stock of the existing documents, and run the interviews. The deliverables are an updated, multilingual work instruction for that process, and a case library of decision criteria running to roughly 20 to 30 entries. You do not have to introduce an AI tool at this stage at all. The objective is to find out whether the elicit-and-structure cycle can actually be run inside your organisation. If it cannot, that is the finding, and it is far cheaper to learn it here.
Phase 2 (introduce the retrieval platform). Put search and question answering on top of the Phase 1 deliverables. Because the scope is narrow, accuracy is easy to verify. This is where you establish whether it is actually used and whether the answers are believed. If it is not used, the cause is not the tool — it is the content or the operating routine.
Phase 3 (roll out and build the operating pattern). Widen the target processes while deciding who owns updating and at what frequency. This is the stage at which the organisation decides who continues the interviews and who approves content. Whether that capability can be brought in-house has a large effect on long-run cost; the approach set out in how to build AI capability in-house is a useful reference for structuring that decision.
At the end of each phase, judge whether proceeding is worth it. The advantage of this slicing is that stopping partway still leaves you the deliverables from the stages completed — a current work instruction and a case library, both of which have standalone value. A tool-first approach has no such property: stop halfway and you have a subscription and an empty index.
What to hand over when you request a quotation
A quotation prepared with insufficient information will always be padded upwards. At minimum, prepare the following.
- Candidate target processes and why they were selected (attribution risk, business impact)
- The number of veterans and expatriates in scope, and how much time they can give to interviews
- Where the existing documents are and roughly how many there are (folder structure, file counts, formats, when they were last updated)
- The current language of your work instructions, and the languages required
- The composition of the audience (how many operators, how many technicians, and the language each group reads)
- The state of existing systems (what you run for production management, maintenance, document management)
- Internal IT policy (whether cloud is permitted, restrictions on data leaving the company)
- Who is expected to own updating (if it is undecided, say “undecided” — knowing that alone changes the proposal)
That last item matters more than it looks. A system designed on the assumption that maintenance ownership is unresolved differs from one designed for in-house maintenance in its UI, its permissions and its editing workflow. Vendors who are told nothing will assume the expensive case.
Measuring the effect — what to put on the scoreboard
“Number of times AI was used” is not a metric
The standard mistake in measuring skill transfer AI is to use query counts or access numbers. Figures of that kind rise on novelty immediately after launch and fall away after a few weeks. When they fall, they tell you nothing beyond “it did not stick” — not why, not what to change.
The purpose of skill transfer is for knowledge to move from person to person. The metrics therefore have to measure knowledge transfer. Four are practical to work with.
| Metric | How to measure it | What it tells you | What to watch out for |
|---|---|---|---|
| Days to competence | Days until an individual is certified able to complete the target process alone. Recorded per new starter | The reduction in the training period itself. The most fundamental metric available | You cannot measure it unless the certification criteria are documented first. Individual variation is wide, so read it as an average across several people |
| Share of the work a person can complete alone | Divide the process into work units and track the proportion of units the person can perform unaided | Shows partial progress. Moves earlier than days-to-competence | Fix the definition of the work units. Change them midway and comparison is lost |
| Number of questions directed at veterans | The count and content of queries concentrating on specific individuals | The degree to which key-person dependency is being resolved. The content also tells you what has not yet been captured | Recording it takes effort. Substitute chat tool logs or a simple tally sheet |
| Number of knowledge updates | Count of Q&A entries and cases added or corrected, and the lag before updates land | Whether the system is alive. Zero is a sign it has become a formality | Chasing the count alone invites padding. Pair it with a review of content quality |
Take the baseline first
None of these metrics supports a claim about effect unless it was measured before deployment. “Days to competence” and “questions directed at veterans” are particularly unforgiving: start measuring after go-live and you have nothing to compare against, and any improvement claim becomes an assertion.
The realistic approach is to gather the baseline at the point you select the Phase 1 process: ask how many new starters have gone through that process historically and how long each took to become competent, and write it down. It does not have to be rigorous. If the shop floor shares the understanding that “it used to take about six months,” that is a serviceable baseline.
For question volume there is a cheap technique: record it for one or two weeks only. Give the veterans a small tally sheet and ask them to write one line — who asked, and what about. That record serves as the measurement baseline and, at the same time, doubles as the list of what to ask about in the interviews. The things people ask about most often are, by definition, the knowledge most worth capturing, which makes this the highest-yield hour of preparation available in the whole project.
The conditions for operations that stick
Metrics do not move unless the operating routine works. The conditions for a system that keeps running are these.
- The moment of use is built into the work. Not “look it up if you get stuck,” but “check the relevant entry before starting this operation,” written into the procedure itself.
- It runs in the audience’s language. The screen, the search, and the answer all complete in Thai or English. Leave one point where Japanese is required and that is where usage stops.
- Update ownership and frequency are decided. “Whoever notices fixes it” means nobody fixes it. Name an owner and set a monthly or quarterly review forum.
- There is an approval rule. Decide who approves content and who guarantees technical correctness. A system anybody can add to without review will eventually stop being believed, and rebuilding that belief costs more than the review ever did.
- Answers show their basis. Every answer carries a link to the source instruction or case. Without it, staff go and check the original themselves and the system has added a step rather than removed one.
- There is a route for reporting errors. Anybody can flag “this answer is wrong,” and the flag feeds into an update. Systems with that loop get more accurate over time; systems without it get less accurate, because reality moves and the content does not.

A practical checklist for avoiding the known failures
The material above, reduced to something you can check before you start.
Before you begin
- Have you narrowed the target to three processes or fewer, rather than starting company-wide? The three selection axes are attribution risk (does it stop if this person leaves), business impact (what the stoppage costs), and whether the people involved will cooperate. Choosing “the hardest process” — the one where feel and embodied skill dominate — as your first target means you hit the Tier 3 wall and the only conclusion the organisation retains is “AI could not do it.”
- Have you classified the knowledge at risk in that process into one of the three tiers?
- Have you excluded Tier 3 (feel and embodied skill) from scope? If not, why not?
- Have you decided who conducts the interviews and who does the structuring? Does the plan quietly assume the veterans will write?
- Have you confirmed the reading language of the audience (Thai staff) and put multilingual delivery into the requirements?
- Have you placed the stock-take and disposal of existing documents *before* the AI deployment, not after?
- Have you decided update ownership and frequency? If undecided, did you tell the vendor that so the design accounts for it?
- Do you have a plan to capture the effect metrics and their baselines before go-live?
Signs to watch for while it is running
- Interview sessions keep getting postponed → either the veteran’s workload has not been reduced, or management has not made the priority real
- The veteran does not return confirmations on the structured content → the confirmation granularity is too fine, or the format is hard to read
- No Thai staff member has said “this doesn’t explain it well enough” → the review has become a formality. Silence here is the unnatural outcome, not the good one
- People are re-checking retrieved answers against the original documents → the system is not trusted. Either the source citation is inadequate or the underlying document quality is poor
Frequently asked questions
How much does skill transfer AI cost?
There is no single figure. Most of the cost sits not in the tool but in the human effort of eliciting and organising the knowledge. Whether you target one process or ten changes the volume of work by an order of magnitude, and whether existing documents are organised or scattered changes it again. The number of languages feeds in directly. When you request quotations, ask for them broken into the five layers described above (1. stock-take and interview design, 2. content production, 3. retrieval platform, 4. operation and updating, 5. training design). Comparing totals alone tends to select the cheap proposal that has omitted layers 1 and 2 — the two that determine whether the thing works. The most accurate approach is to run one process as Phase 1 and estimate the whole from the effort that process actually consumed.
Can tacit knowledge really be captured by AI?
It depends on the type. Procedure (Tier 1) and decision criteria (Tier 2) can be captured. Feel and embodied skill (Tier 3) cannot. Think “tacit knowledge, with AI” without making that separation and there will always be a gap between expectation and outcome. What matters in practice is that in most workplaces the thing holding juniors up is Tier 2, not Tier 3. The discrimination ability learned through the body cannot be transferred, but the sequence that follows it — where to check first, on what basis to stop the line — can be. Capturing only that reduces how often a veteran gets called. It does not reduce it to zero, and zero is not the right target.
We already have work standards. Isn’t feeding those to AI enough?
No, for two reasons. First, what a work standard contains is Tier 1 (procedure) and almost no decision criteria. The situation that stops a junior is “I followed the procedure and it did not work,” and that lives outside the standard by definition. Second, ingesting existing documents as they are means superseded versions, unapproved drafts and duplicates all enter the retrieval set. AI will return wrong content while citing a source for it, which damages trust more than having no system at all. Take stock, discard and consolidate before you ingest anything. That clean-up is unglamorous and it decides most of your eventual output quality.
I’m worried the veterans won’t cooperate with interviews. What can I do?
Start by not asking them to write. Being asked to produce a document puts people on the defensive; “could we talk through how you do this?” is far easier to accept. Beyond that, three things help. First, reduce their normal workload. Book interview time formally as work time and arrange for somebody to absorb the tasks displaced. It needs a manager to state explicitly that this is work, not a favour. Second, show them the result. Building in the step where they confirm the structured version gives them the tangible experience of seeing their own knowledge take a usable shape — and it is often the point at which reluctance turns into engagement. Third, ask on the floor. Conducting it beside the work rather than summoning them to a meeting room changes the perceived burden completely. In most cases the main obstacle to cooperation is not the individual’s attitude but how the request was framed and whether the time was actually protected.
Couldn’t we create everything in Japanese and translate it into Thai afterwards?
You can do it in that order, but designing for multiple languages from the start works out cheaper. The reason is terminology control. Translate afterwards and jig names, machine nicknames and in-house abbreviations vary between translators, producing exactly the “which one does this mean?” confusion on the floor that the document was meant to prevent. The required order is to build the glossary, fix it, and then translate — and that has to be decided while the Japanese version is being written, not after. The second point is to include review by local staff from the outset. A translation can be linguistically correct and still not be phrased the way people actually speak on that floor, in which case it will not be used. Translation quality itself has improved considerably with generative AI; terminology control and floor review still require people.
Can this be used for expatriate handovers as well?
At Japanese-owned plants in Thailand, this is often the application where the effect is most visible. Where expatriate managers rotate every three to five years, the decision criteria built locally — how quality thresholds are applied per customer, how equipment behaves differently by season, incoming inspection policy by supplier — are lost when the assignment ends. Structurally this is exactly the same problem as tacit knowledge in the skilled trades. When you do it, treat it as separate from the standard handover pack. What goes into a handover pack is an organisation chart and a list of responsibilities; decision criteria are essentially absent from it. What you need is “why is this judgement made this way,” and the question patterns described above (branch point, exception, near miss, the thing you never do) work directly for drawing that out. One addition: keep it in a language local management can read. Capture it in a form only the next Japanese arrival can read and the same loss happens at the next rotation.
Summary
Whether skill transfer AI succeeds is decided not by AI performance but by how clearly you separate what is worth keeping. The key points:
- Split tacit knowledge into three tiers: 1. procedure (easy to make explicit), 2. decision criteria (capturable if elicited), 3. feel and embodied skill (not capturable). Put only 1 and 2 onto AI and route 3 to training design
- Giving up on Tier 3 is what protects Tiers 1 and 2: a project that tries to keep everything inflates and stalls. The objective is to reduce how often a veteran is called, not to eliminate it
- The failures reduce to three patterns: film the work and leave it on a server; assign the writing to veterans and watch it stall; feed every existing document to AI and get confident wrong answers. All three come from the missing separation
- “Difficulty converting veterans’ knowledge and experience into explicit knowledge” — 68.6%: per commentary on the 2026 White Paper on Manufacturing Industries, alongside “we do not know how” at 28.6% and a shortage of management resources at 28.1%. Among firms that have implemented data integration, 71.8% are not using AI
- The elicitation step is the centre of gravity: do not make people write; interview them and have the listener structure it. The four question patterns are the branch point, the exception, the near miss, and the thing you never do. NTT DATA’s PoC demonstrates the design of placing AI on the asking side
- In Thailand, expatriate decision criteria are in scope too: knowledge lost at each three-to-five-year rotation has the same structure as shop-floor tacit knowledge. The audience is Thai staff, so multilingual delivery is a requirement, not an option
- Take quotations across five layers: 1. stock-take and interview design, 2. content production, 3. retrieval platform (RAG), 4. operation and updating, 5. training design. The centre of gravity is not the tool but the effort in layers 1 and 2
- Measure by knowledge transfer: days to competence, share of the work a person can complete alone, questions directed at veterans, knowledge update counts. Capture the baselines before you deploy
- Start with one process: select on attribution risk, business impact, and whether cooperation is available. Choose the hardest process first and the project ends at “AI could not do it”
Skill transfer is not the kind of problem that is finished by buying AI. But it is also true that every step in it — eliciting, organising, making retrievable, delivering in the reader’s language — is meaningfully more reachable than it was. Follow the order and the investment stays with you as an asset. Reverse the order and what remains is a folder nobody opens and a chat window nobody types into.
TOMAS TECH is a systems integrator based in Bangkok, Thailand, working across factory IT and OT/FA for Japanese manufacturers operating in the country. From our experience handling shop-floor work and equipment data through the PEGASUS production management system, we are particularly at home designing information for sites where Japanese, Thai and English coexist and where “who is going to read this” has to be answered before anything is written. Conversations even at an early exploratory stage are welcome — “which process should we start with?” or “given the state of our work instructions, can they even go onto AI?” If we can see the current condition of your documents and the process you are at risk of losing, we can start with a view on which of the three tiers your problem actually sits in. Get in touch via our contact page.
References
- 2026 White Paper on Manufacturing Industries | Ministry of Economy, Trade and Industry
- Commentary on the 2026 White Paper on Manufacturing Industries | SmartF
- Institute for Language Understanding launches Manufacturing Shop-Floor Knowledge AI | IT Leaders
- PoC on the transfer of tacit knowledge (Kawasaki Heavy Industries, design work) | NTT DATA
- Takumi AI | Mitsubishi Research Institute
- Why Thailand’s Manufacturing Sector Is Facing a Talent Bottleneck | A Plus Career
- Skilled-labour Crisis in Thailand: Prognosis, Policies and Prospects | RSIS