Blog

2026.08.16

Work Instruction Video AI|Multilingual Video Manuals and Payback in Thai Factories 2026

Work Instruction Video AI|Multilingual Video Manuals and Payback in Thai Factories 2026

You hand out a written work instruction on a Thai shop floor, and the job still does not come out the way you intended. The instruction exists only in Thai, the line runs on operators who speak several different languages, and the explanation changes depending on who happens to be teaching that day. The approach that finally became practical in 2026 for closing this structural gap is work instruction video AI. This article uses a Japanese-affiliated automotive parts plant in Thailand with 280 employees as a model, breaks the effect down into three components — OJT hours, first-time pass rate and procedural deviation risk — and works through a three-tier investment model all the way to payback years.

What Work Instruction Video AI Is|What Changed in 2026

Turning work instructions into video is not new in itself. Film the job on a camcorder, add captions in editing software. Some plants have been doing exactly that for ten years. The reason video manuals never took hold is easy to state — editing, revision and translation cost far more than filming ever did. Half a day to produce one video, a full rebuild whenever the procedure changes, and an outside vendor for every additional language. No plant can sustain 120 processes on that structure.

What changed in 2026 is that the editing, revision and translation layer has started to be automated by generative AI.

Generative AI Turns a Written SOP into a Blueprint for Video

The AI video manual products now in real use cluster around three core capabilities.

First, drafting from existing unstructured data. Feed the AI your scattered material — text SOPs, scanned instructions in PDF, maintenance logs, past trouble reports — and it extracts the key points, then assembles a narration script and a chapter structure for the video. The step of writing a script from nothing disappears, which lowers the barrier to getting started enormously.

Second, semantic understanding of the footage itself. Given a continuous clip of a job filmed on a smartphone, the AI recognises what task is being performed and in what order, in meaningful units, and splits the footage into procedural steps automatically. The work a human used to do — playing the video back and dropping a marker at “this is where step 3 begins” — is now close to fully automatic. Since this step splitting accounted for most of the editing effort, the impact is far from trivial.

Third, multilingual localisation. Because the text is already structured step by step, translation runs against step-level text rather than against a whole video. Applying a glossary keeps proper nouns consistent, and adding speech synthesis generates narrated versions in other languages automatically.

Bundle those three together and you can drop in a text SOP, get a multilingual video manual narrated by an AI avatar in a short time, and have localisation handled automatically as well. Reports indicate this pattern is accelerating, led by large manufacturers outside Japan. Japanese vendors publish language coverage figures such as 192 languages or more than 100 languages, so the language barrier is ceasing to be a technical constraint.

Why AI Speech Synthesis Matters in a Way Subtitles Do Not

When people hear “multilingual” they tend to picture subtitles. On a factory floor, however, subtitles and audio are worth completely different things.

The reason is simple — an operator in the middle of a job is not looking at the screen. Both hands are occupied, protective equipment narrows the field of view, the posture the job demands puts the tablet outside the line of sight. Subtitles do not reach an operator in that state; audio does. On top of that, literacy and spoken fluency do not necessarily match. On any floor there are some operators who speak their mother tongue perfectly well but do not necessarily read it quickly.

Multilingual audio narration, in other words, is not a convenient upgrade over subtitles. It is a device that directly reduces the risk arising from instructions not being understood correctly. That single point is why the three-tier investment model discussed later shows such a large step between Tier A and Tier B.

Drawing the Line Against Skills Transfer AI

It is worth being explicit about how this relates to Getting Started with Skills Transfer AI 2026 on this blog. That article divided tacit knowledge into three layers — procedure, judgement criteria and sensory perception — and examined how far AI can realistically reach. This article takes the opposite side. It is confined to the range AI can carry, namely the standard work procedure itself, and to how you turn that into an asset in video form and roll it out in multiple languages.

Put differently, the domain where a veteran senses an abnormality by sound is outside the scope here. What this article covers is what can be written into an instruction, what already is, and what should be writable but has never been fully written down. Draw that line at the outset, or expectations of the video manual inflate and the project stalls in disappointment at “we made videos and they still do not replace an experienced operator.”

Work Instruction Video AI|Multilingual Video Manuals and Payback in Thai Factories 2026 - figure 1

Why Multilingual Video Manuals Work Especially Well in Thai Factories

Most Plants Do Not Know the Method, or Cannot Make It Work

In commentary on the 2026 edition of Japan’s Monodzukuri White Paper, the challenge cited most often in skills transfer was “it is difficult to codify veterans’ knowledge and experience into explicit form” at 68.6% (multiple answers allowed). What deserves attention is what came next — “we do not know the methods or tools for codification, or cannot make use of them (video recording, work logs, AI analysis and so on)” at 28.6%, followed by “insufficient time or management resources for codification” at 28.1%.

What the second and third places show is that the blockage is not a shortage of motivation or of awareness that this matters. It is on the side of method. People understand intellectually that recording on video would solve it, yet the project sits still because nobody has settled who films, who edits and who handles revisions. Turn that around, and if AI cuts the effort of editing, revision and translation, those 28.6% and 28.1% segments convert directly into addressable demand.

The same white paper states that keeping work procedures, processing conditions, inspection criteria and defect responses as manuals and videos converts skill from “an individual’s experience” into “the organisation’s information.” The accurate way to position work instruction video AI is as a means of lowering the cost of that conversion.

Turnover and Layered Languages, Specific to Thai Plants

Set against a plant in Japan, manufacturing sites in Thailand differ structurally in two ways.

The first is how often people change. Thai manufacturing carries a large share of labour-intensive processes, so new assignments arising from headcount increases, reassignment and turnover occur more frequently than in a Japanese plant. That means the number of times you teach is simply higher, which in turn means a large multiplier sits on top of the cost of each round of training. It also means that systematising training produces a proportionally larger effect.

The second is the layered language structure. Alongside operators whose first language is Thai, a certain share of the workforce comes from neighbouring countries and does not speak Thai as a first language. When the instruction exists only in Thai, those operators end up relying on a colleague’s verbal explanation and gestures, and understanding varies. The real problem here is not the lower level of comprehension itself but the fact that the gap is invisible. The operator nods and starts working, so the failure of communication only surfaces once a deviation occurs.

The share of non-Thai-speaking operators on a given floor varies widely by industry, region and company size, and there is no reliable primary statistic that applies uniformly. This article therefore avoids stating a headcount or a percentage and goes no further than the qualitative premise that a Thai-only instruction creates a structure in which comprehension varies. You can establish your own situation in a single day by taking the operator roster for each process and inventorying first languages.

Three Limits of Paper and Excel Instructions

Compared with video, your current paper or Excel-plus-photo instructions run into the following limits.

  • They cannot express the time axis. Information tied to movement and timing — “turn it slowly,” “stop when the sound changes” — is lost in a still image
  • Version control breaks down on the floor. You distribute a revised PDF, the old version stays in the folder at the line, and nobody can tell which is current
  • Maintenance targets multiply with every language you add. A workflow where fixing the Thai version means also fixing the English version will break without fail once you pass a few dozen processes

The essential value of work instruction video AI lies in the third of these. Build a structure in which one master procedure is the single source and every language version is generated from it, and revisions happen in one place only. Whether you run two language versions or five, the revision effort barely changes.

Estimating the Effect of OJT Video AI at a Model Plant

From here the discussion turns to concrete numbers. What follows is a TOMAS TECH estimation model, and actual figures will differ from plant to plant. The premises are stated explicitly, so please substitute your own numbers as you read.

Model Plant Premises

ItemSetting
Industry and locationJapanese-affiliated automotive parts plant in Thailand
Employees280 (60 indirect, 220 direct)
Standard work instructions120 processes
Current mediumPaper or Excel plus photos, Thai only
Indirect hourly rate150 THB/hour (line leader, skills trainer)
Direct hourly rate50 THB/hour (400 THB per day ÷ 8 hours)

The hourly rates are aligned with the other AI-related articles on this blog. The 400 THB daily figure is an approximation from the minimum wage level, and since actual employer cost also carries social insurance and bonuses, this is a conservative (understated) estimate.

Baseline Annual OJT Cost

Start by counting the current training load.

Annual new-assignment and process-transfer events — replacement of leavers, headcount increases and rotation combined — are set at 55. Each event involves learning an average of 2.5 procedures from scratch, so annual learning events come to 55 × 2.5 = 137.5, rounded to 138.

OJT time per procedure is 2.0 hours for the line leader (skills trainer) and 2.0 hours for the new operator. The leader stays alongside to demonstrate and verify, so the same duration falls on both.

  • Line leader annual OJT hours = 138 events × 2.0 hours = 276 hours
  • New operator annual OJT hours = 138 events × 2.0 hours = 276 hours
  • Annual OJT cost = 276 × 150 + 276 × 50 = 41,400 + 13,800 = 55,200 THB

We also add an in-house indicator we use in shop floor assessments, the first-time pass rate. This is the share of operators who clear the hands-on check immediately after OJT on the first attempt. The Baseline is set at 62%. The main causes are the language gap and variation in how each trainer explains the job.

  • Failures (Baseline) = 138 events × 38% ≈ 52
  • Rework time per failure = 0.75 hours of the line leader plus 0.75 hours of the new operator
  • Rework cost per failure = 0.75 × 150 + 0.75 × 50 = 150 THB
  • Rework cost (Baseline) = 52 × 150 = 7,800 THB
Work Instruction Video AI|Multilingual Video Manuals and Payback in Thai Factories 2026 - figure 2

Breaking Down Post-AI OJT Hours as a Formula

Now to the effect calculation. In a tebiki deployment case (Nihon Closures Co., Ltd.), assembly, disassembly and die work were filmed on smartphones and turned into video manuals with AI captions, roughly 70% of OJT was replaced by video, and comprehension reportedly improved. On the basis of that result, the video substitution rate is set at 70%.

  • Video-substituted events = 138 × 70% ≈ 97
  • Events kept face to face = 138 − 97 = 41 (complex processes and hazardous work stay in person)

Hours for the substituted portion are set as follows. Operator viewing time is 40% of the in-person demonstration, because editing removes dead waiting time and setup, leaving playback and rewinding of the parts that matter. The line leader does not attend, but 0.5 hours of key-point verification remain in order to protect quality. Set that to zero and operators would go straight from watching a video to the hands-on check, which would undermine the first-time pass rate assumption.

  • Line leader annual OJT hours = 97 × 0.5 hours + 41 × 2.0 hours = 48.5 + 82.0 = 130.5 hours
  • New operator annual OJT hours = 97 × (2.0 hours × 40%) + 41 × 2.0 hours = 77.6 + 82.0 = 159.6 hours

The saving is taken as the difference, not as the hours remaining after the reduction.

  • Line leader hours saved = 276 − 130.5 = 145.5 hours → 145.5 × 150 = 21,825 THB
  • New operator hours saved = 276 − 159.6 = 116.4 hours → 116.4 × 50 = 5,820 THB
  • Total OJT time benefit = 21,825 + 5,820 = 27,645 THB per year

Gains from the First-Time Pass Rate

With procedure videos carrying multilingual narration, the content of the explanation becomes identical for everyone, and every operator can hear it in their own language. The first-time pass rate is set to improve from 62% to 84%.

  • Failures (After) = 138 × 16% ≈ 22
  • Rework cost (After) = 22 × 150 = 3,300 THB
  • Quality benefit = 7,800 − 3,300 = 4,500 THB per year (the same as 30 fewer failures × 150 THB)

Note that this improvement is independent of the video substitution rate. Even for the 41 events still taught face to face, the line leader has the same video master to hand, so the sequence and terminology of the explanation are standardised. The improvement in first-time pass rate therefore depends on all 120 processes existing as multilingual video, not on what share was replaced by video. That distinction does real work in the sensitivity analysis below.

Labour Savings Alone Do Not Pay for the Investment

Adding up the benefits so far.

  • Total labour-saving benefit = 27,645 (OJT hours) + 4,500 (first-time pass rate) = 32,145 THB per year

To put it plainly, that amount does not pay back an investment in work instruction video AI. At 32,145 THB per year, it does not even reach the 150,000 THB annual running cost of Tier B described below. Converted to Japanese yen it comes to roughly 140,000 to 150,000 yen, which the tool licence alone consumes.

This conclusion matches what this blog has consistently shown in its other AI-related articles. Try to justify factory AI investment on reduced working hours alone and the numbers almost always fall short. In Thailand, where hourly labour rates are lower than in Japan, that tendency is stronger still. Time savings are a secondary effect and cannot be the lead argument in an investment decision.

So what does take the lead role? The losses already being generated by instructions that fail to get through.

Risk Avoidance Layer|Multilingual Audio Pays Off Most in Curbing Procedural Deviation

How to Estimate the Expected Loss

Language gaps and communication failures sometimes lead to work being performed outside the procedure. Skipping the tightening torque check, setting the jig in the wrong order, misreading the stop condition for an abnormality. Among such deviations, serious events that ended in a line stoppage, a defect escape or a corrective action report to the customer are set at 3.5 per year at the model plant. Average loss per event is 70,000 THB.

  • Expected loss (Baseline) = 3.5 events × 70,000 THB = 245,000 THB per year

These are estimates, not measurements. The frequency of 3.5 is an assumed value derived from the number of corrective action reports over the past several years that can be judged to stem from language or communication, and the 70,000 THB per event is an approximation combining the opportunity cost of downtime, sorting labour and freight. If you are estimating for your own plant, start by opening the list of corrective action reports held by quality assurance and counting the cases whose cause column contains wording such as “operator did not understand” or “misread the instruction.” Once that figure is replaced with real data, every calculation downstream becomes your own.

Reduction Through Multilingual Audio

Assume that adding narration in English and in the other languages actually spoken on the floor, on top of Thai, reduces the incidence of these deviations by 55%.

  • Expected loss (After) = 245,000 × (1 − 55%) = 110,250 THB per year
  • Risk avoidance benefit = 245,000 − 110,250 = 134,750 THB per year

The 55% reduction rate is also an assumption. However, the direction of the conclusion holds even if that assumption is somewhat off. Even at a 30% reduction the benefit would be 73,500 THB, far above the 32,145 THB of labour savings. Which layer plays the lead role in the return-on-investment argument does not shift.

Total Benefit

  • Total benefit = 32,145 (labour savings) + 134,750 (risk avoidance) = 166,895 THB per year

About 81% of the total benefit comes from the risk avoidance layer. Understanding this structure matters, because a capital request written around “video reduces training time and therefore costs” will be rejected on insufficient value. The framing that fits the numbers is “this is an investment that reduces losses from procedural deviation, and the reduction in training time is a by-product.”

Three-Tier Investment Model|Tiers A, B and C and Their Payback

Even within the same category of AI video manual deployment, how far you go changes both the investment and the benefit completely. Here is a comparison across three tiers.

Tier A|Smartphone Filming Plus AI Captions (Thai Only)

Film the job on the smartphones you already have and auto-generate captions with AI. Step splitting and translation are either unused or limited to displaying translated captions.

The investment breaks down as follows.

  • Initial investment 10,000 THB (6,000 for filming accessories such as tripods and external microphones, 4,000 for tool setup and template preparation. Existing smartphones are repurposed for filming)
  • Annual running cost 90,000 THB (54,000 for a lower-tier licence covering captions and translation only, 27,000 for refilming on revision = 40 processes × 4.5 hours × 150 THB, 9,000 for the administrator’s monthly effort = 5 hours per month × 12 months × 150 THB)

It is worth being explicit about why revision effort is set differently for Tier A and Tier B. The number of processes revised per year is 40 in both cases. What differs is the effort per process.

Tier A has no step splitting function, so the video is treated as one continuous piece of footage. Even when only part of the procedure changes, filming itself takes 1.0 hour, but reapplying every caption, shifting the timecodes and re-editing the overall running time takes 3.5 hours, for a total of 4.5 hours. In Tier B, steps are independent objects, so refilming the affected step takes 0.5 hours and verifying the regenerated multilingual narration takes 1.0 hour, for a total of 1.5 hours. The same 40 revised processes therefore split into 27,000 THB and 9,000 THB. That threefold difference is the practical value of the step splitting function.

The premises behind the initial investment are not aligned either. Tier A assumes the first round of filming across 120 processes is absorbed within existing OJT hours, so no additional cost is booked. Tier B books 27,000 THB up front as dedicated effort for the initial filming and verification of the AI step splitting, in order to protect the quality of the multilingual narration. This difference in premises favours Tier A, and yet Tier A still runs at a loss, as the next figures show.

Because Tier A cannot produce multilingual audio narration, no benefit accrues from the risk avoidance layer.

  • Annual benefit = 32,145 THB (labour savings only)
  • Annual balance = 32,145 − 90,000 = −57,855 THB

That is a loss of 57,855 THB every year, so operations do not stand up even before you consider recovering the initial investment. What is more, the 32,145 THB benefit used here is itself an optimistic figure. The improvement from 62% to 84% in first-time pass rate assumes the language gap has been closed, so a Thai-only Tier A would deliver the same improvement only among Thai speakers. Include non-Thai speakers and the actual benefit falls below this.

Tier A has value as a pilot to test the water, but it is not a tier you should choose as a permanent operating model.

Tier B|Automated Multilingual Narration Through AI Speech Synthesis

Tier A plus automatic step splitting and multilingual narration generated by AI speech synthesis, covering all 120 processes.

  • Initial investment 150,000 THB (40,000 for tool setup and template preparation, 45,000 for a full filming kit = 2 smartphones, 2 tablets, tripods and external microphones, 27,000 for filming 120 processes and verifying the AI step splitting = 120 processes × 1.5 hours (1.0 hour filming plus 0.5 hours verifying the split) × 150 THB, 38,000 for terminology proofreading of the multilingual narration and shop floor review)
  • Annual running cost 150,000 THB (108,000 licence = 9,000 THB per month, 9,000 internal effort for revision filming and regeneration = 40 processes × 1.5 hours × 150 THB, 24,000 outsourced multilingual proofreading, 9,000 administrator’s monthly effort)
  • Annual benefit = 166,895 THB
  • Annual balance = 166,895 − 150,000 = 16,895 THB
  • Simple payback = 150,000 ÷ 16,895 ≈ 8.9 years

It does run a surplus. But with an annual surplus just under 17,000 THB, the arithmetic works out at 8.9 years. Hold the initial investment down to 130,000 THB and it is still 7.7 years; let it swell to 180,000 THB and it becomes 10.7 years. That is not a payback period a capital request for equipment gets approved on.

The point to note is not that Tier B is a bad investment but that a single site is not large enough. Break down the cost side and you find a mix of items that scale cleanly with the number of sites and items that do not. The licence fee (108,000 THB) rises roughly in proportion as sites or user accounts increase, with a volume contract shaving perhaps 10% off the unit price. Terminology proofreading of the multilingual narration (24,000 THB) and the internal effort of content production, on the other hand, can be shared across sites once the common processes have been produced, so they do not rise anything like in proportion to site count. Going to three sites raises them not by a simple factor of three but by somewhere between 1.2 and 2.3 times.

The benefit, by contrast, scales almost exactly with site count. Every site takes on new assignments, and every site carries its own procedural deviation risk. This asymmetry — benefit scales, cost only partly scales — is what makes Tier C work.

Work Instruction Video AI|Multilingual Video Manuals and Payback in Thai Factories 2026 - figure 3

Tier C|Sharing a Content Platform Across Multiple Sites

Three sites in the group share the video manual platform, and each films only the processes unique to it. The master for common processes and its multilingual narration are produced once and rolled out horizontally.

  • Combined annual benefit across 3 sites = 166,895 × 3 = 500,685 THB

Stack Tier B up three times as is, and the investment would be 450,000 THB initial and 450,000 THB annual. Here is how far sharing brings each line item down.

Initial investment is 330,000 THB, which is 2.2 times a single site (150,000 THB) rather than three times. The breakdown is 80,000 for tool setup and template preparation (120,000 if taken three times over; one common template suffices, but site-specific settings remain), 135,000 for the filming kits (no compression from the three-site figure of 135,000, since each site needs its own equipment), 61,000 for filming the common processes and verifying the AI step splitting (81,000 if taken three times over; the saving is exactly what comes from handling the duplicated processes once), and 54,000 for terminology proofreading of the multilingual narration and shop floor review (114,000 if taken three times over; the glossary and audio for common processes need only one round of proofreading, which is where compression is greatest).

Annual running cost is 360,000 THB, 2.4 times a single site. The breakdown is 288,000 for a three-site volume licence (324,000 otherwise), 18,000 of internal effort for revision filming and regeneration (27,000 otherwise, since one representative site handles revisions to common processes once), 30,000 for outsourced multilingual proofreading (72,000 otherwise), and 24,000 for the administrator’s effort (27,000 otherwise). Because the licence barely compresses, the compression ratio on the annual side is smaller than on the initial side.

  • Annual balance = 500,685 − 360,000 = 140,685 THB
  • Simple payback = 330,000 ÷ 140,685 ≈ 2.3 years

The same mechanism, with nothing changed except going to three sites, sees payback shrink from 8.9 years to 2.3 years. Video manual content has the property that once produced it costs almost nothing to duplicate, so economies of scale apply cleanly in this area.

Read the other way, if you are considering an investment in work instruction video AI for a single site, the decision becomes distinctly marginal. In that case you should settle the question of how many sites the platform is meant to serve — including other group sites and any planned expansion — before you debate the investment amount.

Comparing the Three Tiers

ItemTier ATier BTier C (3 sites)
ConfigurationSmartphone filming plus AI captionsA plus multilingual narration via AI speech synthesisB plus content sharing across sites
LanguagesThai onlyThai plus the other languages on the floor plus EnglishSame as B
Initial investment10,000 THB150,000 THB330,000 THB
Annual running cost90,000 THB150,000 THB360,000 THB
Annual benefit32,145 THB (optimistic, as noted above)166,895 THB500,685 THB
Annual balance−57,855 THB16,895 THB140,685 THB
Simple paybackNever pays backAbout 8.9 yearsAbout 2.3 years

When comparing the initial investment row, mind the difference in premises. The 10,000 THB for Tier A excludes the effort of the first round of filming across 120 processes (assumed absorbed within existing OJT hours), whereas Tier B and above include it as dedicated effort. In other words, Tier A is presented in a way that makes its initial investment look smaller than it really is, and it still runs an annual loss.

Sensitivity Analysis|If the Video Substitution Rate Were 60%

The 70% substitution rate is an assumption grounded in the tebiki deployment case. Depending on your own process mix, a high proportion of complex or hazardous work may mean you cannot substitute that far. Here is the calculation at 60%.

The most important thing here is which effects the sensitivity factor is applied to. Only the reduction in OJT hours depends on the video substitution rate. The first-time pass rate improvement and the risk avoidance layer do not. The former is tied to all 120 processes existing as multilingual video, the latter to the existence of multilingual audio. Apply 0.6/0.7 uniformly to every benefit and the conclusion comes out worse than reality.

Recalculating at a 60% substitution rate.

  • Video-substituted events = 138 × 60% ≈ 83, events kept face to face = 55
  • Line leader annual OJT hours = 83 × 0.5 + 55 × 2.0 = 41.5 + 110.0 = 151.5 hours
  • New operator annual OJT hours = 83 × 0.8 + 55 × 2.0 = 66.4 + 110.0 = 176.4 hours
  • OJT time benefit = (276 − 151.5) × 150 + (276 − 176.4) × 50 = 18,675 + 4,980 = 23,655 THB
Item70% substitution60% substitution
OJT time benefit27,645 THB23,655 THB
First-time pass rate benefit4,500 THB4,500 THB (unchanged)
Risk avoidance benefit134,750 THB134,750 THB (unchanged)
Total benefit166,895 THB162,905 THB
Tier B annual balance16,895 THB12,905 THB
Tier B paybackAbout 8.9 yearsAbout 11.6 years
Tier C annual balance140,685 THB128,715 THB
Tier C paybackAbout 2.3 yearsAbout 2.6 years

There are two things to take from this.

First, total benefit falls only 2.4%, from 166,895 to 162,905. The video substitution rate, a variable that looks important at first glance, in fact barely moves the total benefit, because 80% of the benefit sits in the risk avoidance layer and that layer does not depend on the substitution rate.

Second, Tier B payback nonetheless stretches by 2.7 years, from 8.9 to 11.6. When an investment runs on a thin annual balance, a small movement in benefit is heavily amplified in the payback figure. Tier C, by contrast, moves only 0.3 years, from 2.3 to 2.6. Judged on robustness against uncertainty in the premises as well, the multi-site rollout is the safer design.

How to Choose an AI Video Manual Tool

Step Splitting Is the Biggest Fork in the Road

The first thing to check in tool selection is whether the product performs automatic step splitting with AI. That is what determines editing effort.

Tools without step splitting head in the direction of captioning and translating the footage as filmed. Because the file is handled as one long video, “we only want to revise step 5” means either refilming the whole thing or swapping the relevant section in a video editor. On processes that are revised often, that difference decides whether the operating model survives.

With step splitting, the video is decomposed into step-level objects. A revision means refilming only the affected step, and translation runs against step-level text. That is exactly why revision is set at 1.5 hours per process in the Tier B annual running cost and 4.5 hours in Tier A.

Language Coverage and Whether Speech Synthesis Is Included

Every vendor publishes a language count, but rather than the size of the number, look at whether the languages actually spoken in your plant are covered and whether speech synthesis is supported. As noted above, 100 languages with subtitles only will not generate any benefit in the risk avoidance layer.

Comparison of the Main Tools

The following is compiled from information published as of 2026. Prices and language counts are subject to change, so always confirm the latest information when you evaluate.

ToolIndicative pricingLanguagesPublished characteristics
DiveFree tier available, paid from 10,000 JPY per month for 10 accounts (Japan domestic pricing)192 languagesAutomatic AI step splitting is patented technology
Teachme BizContact vendor20 or more languagesStep splitting could not be confirmed from published information
tebikiContact vendorMore than 100 languagesComparison articles state there is no step splitting function; the focus is captions and translation on the footage as filmed
Manual.toCould not be confirmed from published informationMore than 200 languages, speech synthesis (TTS) availableGenerates an end-to-end multilingual manual from one video in about 60 seconds
SynthesiaCould not be confirmed from published information160 languagesCould not be confirmed from published information

Vendors do not all publish step splitting information at the same level of detail. For the products marked as unconfirmed above, the reliable approach is to ask directly in the sales meeting — “when we want to replace only step 5 of one process, what has to be refilmed?” If the answer is that the whole process must be refilmed, the operating model will be equivalent to Tier A.

For use at a Thai site, note that tools built for the Japanese domestic market have UIs and support centred on Japanese and English. Your selection criteria change depending on whether the screens operators touch directly need to display in Thai, or whether only administrators use the Japanese UI while operators simply watch the generated videos.

Three Conditions for Making It Stick on the Floor

The failure mode where a tool is introduced and then never used is common with video manuals too. Three things decide whether it takes hold.

  • Name who films. “Everyone films” means nobody films. Assign a filming owner per process and protect that effort as official working hours
  • Embed the moment of viewing into the workflow. Not “watch it if you need to” — put “watch the relevant video” on the checklist for new assignments as a line item
  • Define the revision trigger. Document a rule that the relevant video is updated whenever any of three things occurs — a process change, a 4M change, or a corrective action

Our thinking on designing for adoption is set out in more detail in AI Adoption Enablement 2026. This part deserves as much effort as tool selection, if not more.

Implementation Steps|Putting Your First 30 Processes on Video in 90 Days

Target all 120 processes at once and you will run out of steam before filming ends. We recommend finishing 30 processes in 90 days, locking down the operating pattern there, and then expanding to the rest.

Days 1 to 30|Selecting Target Processes and Piloting One

Choose target processes from the overlap of three sets — processes newly assigned operators handle first, processes with a low first-time pass rate, and processes that have produced a serious corrective action in the past. The key is to select on training frequency and cost of failure, not on the importance of the process.

The one thing you must do in these 30 days is take a single process all the way through. Filming, AI step splitting, proofreading the text, generating the multilingual narration, and a review with non-Thai-speaking operators actually watching it. The need for a glossary and problems with narration speed will surface here without fail, so settle them before expanding to 30 processes.

Days 31 to 60|Filming 30 Processes and Building the Glossary

Allow 30 minutes to one hour of filming per process. Filming processes on the same line together makes the setup more efficient.

In parallel, build a glossary of proper nouns. Equipment names, jig names, the abbreviations used internally. Without registering these as a dictionary, AI translation will assign different renderings from one process to the next and operators will be confused. Glossary work is unglamorous, but it is the single biggest determinant of quality in multilingual rollout.

Days 61 to 90|Documenting the Operating Rules and Setting a Measurement Baseline

Once the 30 processes are complete, put the “who films, when it is watched, what triggers a revision” points above into a written document. At the same time, capture the numbers that will serve as the baseline for measuring effect — specifically, the actual first-time pass rate and OJT hours for the processes you have put on video. Start without taking a Baseline and you will have nothing to say about whether it worked six months later.

Common Stumbling Blocks

Here are the failures that occur most readily in practice.

  • Filming that is too polished. Video shot under added lighting and retaken many times takes a long time to produce and looks different from the real working environment. Handheld footage from the actual floor is entirely sufficient
  • Narration that is too fast. The default speed of AI speech synthesis is often too fast for a non-native speaker listening while working, so slow it down and check it on the floor
  • Videos that run too long. Nobody watches a ten-minute video to the end. Aim for one video per procedure, each under two minutes
  • Treating production as the finish line. A video manual that is never revised will be in the same state as a paper instruction within six months

For the wider picture of how other plants are using AI, see Shop Floor AI Use Cases 2026. For the higher-level question of how to position generative AI across the plant as a whole, see Generative AI Implementation Guide 2026.

Frequently Asked Questions

What is work instruction video AI

It is a mechanism in which generative AI processes existing text instructions and job footage shot on a smartphone, then automates splitting into procedural steps, generating the narration script, translating into multiple languages and synthesising speech, to produce a video manual. The difference from a conventional video manual is that the effort of editing, revision and translation drops sharply. Filming itself is done on a smartphone, so no special equipment is required.

How much does an AI video manual tool cost

In the model plant used in this article (280 employees, 120 processes), Tier A with captions only comes to 10,000 THB initial and 90,000 THB annual running cost, and Tier B including multilingual narration comes to 150,000 THB initial and 150,000 THB annual. However, single-site Tier B takes about 8.9 years to pay back, so making it work as an investment realistically requires a design in which multiple group sites share the platform. Tier C shared across three sites pays back in about 2.3 years.

How do you produce a multilingual video manual

The standard approach is to build a master version in Japanese or Thai and generate each language version from it. What matters is running translation against step-level text rather than against a whole video. Choose a tool with step splitting and that structure follows naturally. Register a glossary of equipment and jig names in advance as well, or renderings will drift. Supporting speech synthesis and not just subtitles has a large effect on how well it works in practice on the floor.

How does this differ from skills transfer AI

Skills transfer AI addresses the question of how far tacit knowledge — a veteran’s judgement criteria and sensory perception — can be carried by AI. The work instruction video AI covered in this article sits at the opposite pole, taking the range that can already be made explicit, namely the standard work procedure, turning it into an asset in video form and rolling it out in multiple languages. A video manual will not replace an experienced operator, but by raising the reproducibility of standard work it frees experienced people to spend time on the non-routine judgement they should be handling. The two do not compete, and in terms of sequence, taking work instruction video AI first gives a clearer return on investment.

Summary

Work instruction video AI became a method capable of sustaining operation at a scale of 120 processes only once generative AI sharply reduced the effort of editing, revision and translation in 2026. As the Monodzukuri White Paper shows, many plants recognise the need and are blocked on the side of method, and that blockage is now being resolved technically.

At the same time, the investment numbers deserve an honest look. In the model plant estimate, labour savings — reduced OJT hours plus improved first-time pass rate combined — come to only 32,145 THB per year, which does not even cover the annual running cost of Tier B. What makes the investment work is the risk avoidance layer worth 134,750 THB per year, achieved by curbing procedural deviation through multilingual audio. Even then, single-site Tier B needs about 8.9 years to pay back, and only Tier C, sharing the content platform across three sites, brings it to a practical level of about 2.3 years.

The first question to ask when weighing this investment, in other words, is not which tool to choose but how many sites the platform is for. Settle that and both the functions you need and the amount you invest narrow down automatically.

How this estimate shifts for your own process count and language mix can be approximated from two pieces of data — the number of corrective action reports and the first-language composition of your operators. TOMAS TECH supports Japanese-affiliated manufacturers in Thailand with designing video manual platforms that fit the reality of the shop floor and with building investment plans that anticipate multi-site rollout. There is no need to have decided on deployment; if you would like to work out how far your own organisation should go, please get in touch through our contact page.

References and Sources