Enterprise translation automation almost always stalls at the same point. You swap the engine, you refine the prompts, and the shop floor still says the translations are inconsistent. The problem is not accuracy. It is design. As long as every internal document flows through one pipeline, a controlled work instruction and a daily production report get identical treatment, and nobody can say how much of the job the machine is allowed to own. This article sets out a different structure: sort your documents into four classes, then rebuild the workflow in three layers — terminology, generation, and approval. The context throughout is a Japanese-owned manufacturing plant in Thailand, where Japanese, Thai, and English circulate every day.
Why Enterprise Translation Automation Stalls
At Japanese-owned manufacturers with sites in Thailand and Vietnam, three languages are in daily circulation — Japanese, Thai, and English — and often four or five once Vietnamese or Chinese is added. Technical bulletins from head office. Local work standards. Responses to customer audits. Daily production reports. Every one of these generates translation work, and most of it is absorbed by somebody’s overtime.
The Myth That Switching Engines Fixes It
When translation becomes a bottleneck, the first thing most organizations examine is the engine. Move from legacy machine translation to an LLM-based system, or from a general-purpose LLM to a translation-specialist service. Yet after the swap, the complaints from the plant barely change.
The reason is simple. The complaint was never “the translation is bad.” It was “the translation is not consistent.” Last month’s work standard rendered “changeover” as การเปลี่ยนรุ่น; this month’s revision uses a different word. The same piece of equipment is called something different in every document it appears in. This is terminology drift, and with a general-purpose LLM used without context it happens structurally, not occasionally. The model produces a plausible rendering every time. It has no memory of how it rendered the same term last time.
Assessments published in 2026 make the same point from the other direction: the range machine translation can handle has clearly widened, but it has not reached the level where human judgment can be removed from enterprise localization. One comparison found that LLM translation without supplied context was rated “good” in 55.7% to 80% of cases depending on the language pair. Note what those figures cover — English into German, Polish, and Russian, all European pairs. For Japanese into Thai or Vietnamese, which is what this article is about, you should assume worse numbers, simply because the training data is thinner. Even taking the European figures at face value, 20% to 40% of cases are not rated good. And you cannot tell in advance which 20% will fail.
How a Single Pipeline Breaks Down
The first architecture most companies build routes all internal documents into one translation workflow. Drop in a file, get back a translation. It looks efficient, and it contains a fatal conflation of assumptions.
Push a work standard and an internal email through the same process and one of two things happens. Set the quality bar at the work standard’s level, and a single email now requires an approval workflow — so nobody uses the system. Set the bar at the email’s level, and work standards reach the shop floor unverified, and the finding shows up in your next audit. Split the difference and you get a compromise that satisfies neither requirement.
Enterprise translation automation fails not because the engine is not accurate enough. It fails because the boundary of what the machine may own — which should differ document by document — is being fixed with a single global setting.
Move the Unit of Decision to the Document Class
The fix is straightforward in principle: move the unit of decision from the tool to the document class. The question is not which engine to use. It is which class this document belongs to, how far the machine goes within that class, and where a human signs. Engine selection comes afterward. In fact, using different engines for different classes is perfectly reasonable.

Sorting Internal Documents Into Four Classes
The four classes below come from looking at the documents actually circulating in Japanese-owned plants in Thailand and dividing them by one question: when a mistranslation occurs, what breaks? Details vary by industry, but these four cover most of the ground.
| Class | Typical documents | What breaks on mistranslation | Machine’s scope | Human’s role |
|---|---|---|---|---|
| A Controlled documents | Work standards, quality manuals, inspection criteria, drawing notes, SOPs | Basis of quality assurance, audit conformity | First draft only | Term base checking and approval signature |
| B Operational documents | Daily reports, internal notices, meeting minutes, email, routine reports | Temporary misunderstanding (recoverable) | Fully automated is fine | Spot checks only |
| C External documents | Customer submissions, audit responses, contracts, complaint responses | Commercial relationships, legal liability | Consistency checking only | Translation and final judgment |
| D Shop-floor UI text | HMI screens, andon displays, form labels, poka-yoke indicators | Operator error, broken layouts | Out of scope for translation | Built in as design |
Class A: Controlled Documents — AI Goes as Far as the First Draft
These are the documents under ISO 9001 control. Clause 7.5, “Documented information,” requires that it be available and suitable for use where and when needed, that it be identifiable, and that it be adequately protected from unintended alteration. Controlled documents are issued after formal review and approval, and their distribution and retention are managed.
When translation drifts here, the consequence is not merely awkward reading. If the Japanese work standard and the Thai version express a tightening torque differently, an auditor will ask which one governs. In a compliance context, the term base (approved terminology) and the translation memory (approved sentence-level translations) form a controlled vocabulary layer, and auditors expect that layer to be documented. In other words, a term base is simultaneously a quality tool and an audit artifact you can produce on request.
In this class, AI’s mandate ends at the first draft. Draft generation is fast, and correcting a draft is measurably faster than translating from a blank page. But the approval signature belongs to a person, and the record must show who approved which revision, and when.
Class B: Operational Documents — Full Automation Is Fine
Daily reports, internal notices, meeting minutes, email. The cost of a mistranslation in this class is low, and speed itself is the value. Readers already hold the context, so an unnatural rendering still carries the intent. If it does not, they ask.
This is where automated meeting minutes AI and AI transcription pay off most directly. Picture a meeting with Japanese managers and Thai staff in the room, where Japanese and Thai alternate sentence by sentence. Traditionally somebody took notes by hand and produced both language versions afterward. With generative AI, one continuous process handles it: transcribe the recording, extract decisions with owners and due dates, and output in both languages. This is also the class where return on investment is easiest to read.
Document drafting AI and automated report generation should start here as well. Routine production reports and monthly summaries have fixed templates and pull their figures from systems. Automating both generation and translation at once is safe here because the blast radius of a failure is small.
“Fully automated is fine” does not mean “nobody looks.” Pull a handful of documents a month, read them, and check only whether any mistranslation has changed the meaning. The purpose is not quality assurance. It is a sensor that tells you when an engine change or a configuration change has caused degradation. A few minutes per document is enough.
Class C: External Documents — Use AI for Consistency Checking, Not Drafting
Customer submissions, audit responses, contracts, complaint responses. Because a mistranslation here reaches commercial relationships and legal liability directly, avoid handing the translation itself to the machine.
That does not mean not using AI. It means using it at a different point. Take a human-produced translation and have AI back-translate it against the Japanese original. Have it detect dropped or transposed figures, model numbers, and dates. Have it check whether any term conflicts with the internal term base. Machines are good at this kind of consistency checking, and it is exactly where human reviewers tend to lose concentration. As applications of AI translation for business go, it is unglamorous. Its effect on incident rates is not.
Class D: Shop-Floor UI Text — This Is Design, Not Translation
This is the class most often overlooked. Button labels on HMI screens, andon messages, column headers on printed forms, warning text on poka-yoke devices. All of them carry hard constraints on character count and display width.
A Japanese status label such as 異常停止 (“abnormal stop”) fits in four characters. The Thai equivalent, หยุดเนื่องจากความผิดปกติ, expands dramatically. Push it through a standard document translation workflow and you get back a translation that is grammatically correct and overflows the screen. The shop floor trims it by hand, the trimmed version never makes it back into the term base, and the same thing happens at the next revision.
Separate shop-floor UI text from the translation process and treat it as part of screen design. Fix the maximum displayable character count as a specification first, then write the text for each language inside that constraint. The person who decides is not a translator; it is someone who knows the shop floor. If you automate without making this split, UI text rework never ends, and the verdict becomes “we automated and nothing got easier.”

A Three-Layer Design That Stops Terminology Drift
If the class split answers “what do we delegate, and how far,” the three-layer design answers “how do we build it.” Keep the terminology layer, the generation layer, and the approval layer conceptually separate.
Layer 1: Terminology — Term Base and Translation Memory
At the bottom sit the term base and the translation memory. The term base is a dictionary that fixes, at the word level, how a given term is translated. The translation memory is a record that accumulates, at the sentence level, how a given sentence was previously translated and approved.
Without this layer, even the best LLM will render the same term differently every time. With it, you can swap out the engine in the generation layer and terminology consistency survives. Engines turn over every few years. A term base persists as an asset. The first investment is the terminology layer, not the engine.
What to register: company-specific terms, product names, process names, equipment names, department names, form names, quality terminology. Around 200 terms is enough to start. Starting with the 200 most frequent terms and adding entries whenever drift is found survives far better than trying to be exhaustive on day one and stalling.
Layer 2: Generation — LLMs Plus Internal Context
The middle layer is where translations are actually produced. You use an LLM here, but you do not throw text at a bare model. You pass the term base, previously approved translations, and the process context the document belongs to alongside the source. The baseline pattern is retrieval-augmented generation (RAG), with approved past translations as the retrieval corpus.
How to inject internal context is where most of the implementation judgment lives. Which documents go into the index. Which revision governs when several exist. How to isolate confidential material. We cover that thinking in our guide to RAG implementation cost and rollout for factory knowledge, which is worth reading at the point where you are designing this layer.
Varying the model by use case is also worth doing. Comparisons show that the model best suited to technical documentation and software localization is not the model best suited to brand-heavy marketing translation. Factory documents sit firmly in the first category, but externally facing company profiles behave more like the second. Do not try to cover everything with one model.
Layer 3: Approval — Version Control and Sign-Off
The top layer is approval. Record who approved which revision and when, and ensure that only approved revisions reach the shop floor. Putting translations under ISO 9001 clause 7.5 control means, concretely, implementing this layer.
Four things need to be decided here.
- Who approves (define per class — the quality assurance manager for Class A, a department head for Class C, and so on)
- What constitutes approval (electronic signature, an approval record in a workflow tool, or a physical stamp — pick one)
- How superseded revisions are handled (physically recalled from the floor, or invalidated in the system)
- How approved translations flow back into the terminology layer (skip this and the three layers never close the loop)
The fourth point matters most. If an approved translation does not return to the translation memory, a human will make the same correction again next time. The three-layer design is not a one-way path from terminology to approval. It only works when there is a cycle, with approved translations flowing back down.
Why the Three-Layer Diagram Matters Internally
The structure is also a useful instrument for internal communication. Say “we’re going to automate translation with AI” and the executive response is “who guarantees quality?” while the shop-floor response is “another new tool.” Show the three layers, say that humans own the approval layer and that the terminology layer registers the names people actually use on the floor, and the discussion becomes concrete. Buy-in for this kind of project moves on an explanation of responsibility boundaries, not on an explanation of technology.
Replace “99% Accuracy” With TTE as Your KPI
The classic way a translation automation project loses its bearings is setting the KPI as a translation accuracy percentage.
Accuracy Percentages Do Not Describe Workload
Ninety-five percent accuracy sounds ample and means nothing operationally. If a document has 100 sentences and five are wrong, a human reads all 100 to find the five. The reading time barely falls as accuracy rises. At 99% accuracy you still have to read the whole document, because you do not know which single sentence is the bad one.
Worse, LLM translation errors do not announce themselves. The output is fluent and natural, and only a term has been replaced with a different word. Without the tell-tale awkwardness of legacy machine translation, the error is easier to miss. The higher the accuracy figure climbs, the less alert the reviewer is, and the higher the miss rate goes.
What TTE (Time to Edit) Is
The metric to use instead is TTE (Time to Edit): the measured time for a human to bring machine output up to usable quality. In 2026 machine translation evaluation, TTE is defined as the time a professional translator needs to raise machine translation output to human quality, and it is treated as the metric suited to observing real deployment impact in production. Whether the work actually decreased or merely moved downstream is something you cannot know without measuring this.
Measurement is simple. For the same document, compare the time taken to translate from scratch against the time taken to correct a machine-generated draft. If the latter is clearly shorter, the automation is working. If it is about the same or longer, automation is not functioning for that document class.
| Metric | What it measures | Usable for operational decisions? |
|---|---|---|
| Accuracy (%) | Proportion of correct translation | No. Does not correlate with review effort |
| Automatic scores such as BLEU | Similarity to a reference translation | Reference only. Blind to terminology drift |
| TTE (Time to Edit) | Actual time to reach usable quality | Yes. Maps directly to labor cost saved |
| Term base violations | Deviations from approved terminology | Yes. Directly tied to Class A quality control |
Cautions When Measuring TTE
TTE has to be measured separately per document class. If Class B shows an 80% reduction in TTE while Class A shows zero, then Class A either comes out of scope for automation or its terminology layer is not built out yet. An overall average hides that decision.
Second, TTE is sensitive to how practiced the editor is. Immediately after rollout, nobody has a rhythm for correcting machine output, and TTE reads longer than the true figure. Do not judge on the first month’s numbers; watch the trend over roughly three months. Conversely, an area where TTE has not fallen after three months is not a familiarity problem. It is a design problem.
Where Thai and Vietnamese Break the Workflow
Ignore language-specific realities and a correct design still fails in practice. Here are the points that deserve particular attention in ASEAN.
Thai — Segmentation Errors Cascade
Thai does not place spaces between words. Machine processing therefore has to determine where the character string divides into words — word segmentation — before anything else. Get that wrong and everything downstream collapses with it.
Proper nouns, internally coined terms, and Japanese loanwords transliterated into Thai script are all absent from dictionaries, which makes their boundaries easy to misjudge. Manufacturing terms such as poka-yoke, andon, and kanban written in Thai script are the textbook case. Registering these in the term base therefore carries more weight than simple terminology unification.
Token Counts and the API Cost Asymmetry
Multilingual LLM tokenizers are not equally efficient across languages. For the same content, Thai script has been noted to consume more tokens than English. Since API billing is per token, that is a direct cost difference. Japanese also fares worse than English, so the gap against Thai is not as wide as it looks from an English baseline — but starting from English-based unit costs will still mislead you.
Use English-based rates at the estimating stage and you will exceed your assumptions the moment you move to Thai operations. At the modeling stage, run actual samples in the target language, measure the token counts, and then set your unit cost. The larger your translation volume, the less this gap can be ignored.
It is also worth noting that LLM evaluation for Thai and other Southeast Asian languages has matured considerably in recent years. Alongside region-wide efforts such as SEACrowd and SEA-HELM, Thai now has instruction-following datasets such as WangchanThaiInstruct and benchmarks that measure cultural and foundational capability. When selecting a model, these regional evaluations are worth consulting alongside the English-centric benchmarks.
Vietnamese — Dropped Tone Marks
In Vietnamese, the presence or absence of tone marks changes what a word means. When text with stripped marks enters the pipeline, the machine treats it as a different word entirely. The cause is usually not the translation engine but something upstream of it: CSV exports from legacy core systems, character encoding mismatches, form templates using fonts that do not support the marks.
The countermeasure sits outside the translation process. Put encoding and diacritic-integrity checks on the input side and reject bad text before it reaches the engine. Treat it as a translation quality problem and you will never find the cause.
Handling Multilingual Internal Questions
As languages multiply, “I don’t know who to ask” becomes a bigger problem than translation itself. When a Thai staff member reads a Japanese bulletin and has a question, and there is no way to search internal regulations written in Japanese, the only path left is asking a Japanese manager directly. That is what consumes management time.
We have set out approaches to making first-line internal support multilingual in our guide to chatbot implementation cost and rollout. Translation automation and support automation tend to get chartered as separate projects, but the term base they reference can and should be shared.
The Work Automation Does Not Remove
The single biggest reason translation automation business cases miss is that nobody counted the work that will not be automated.
Text Inside Drawings — Out of Range for Text Translation
Notes on CAD drawings, nameplates visible in equipment photographs, tables inside scanned PDFs. None of this text is addressable as text, so none of it rides the normal translation workflow. And in factory documentation, the share of content in this state is far from trivial.
Most of the residual shop-floor effort lives here. Auto-translate the body of a work standard and leave the instructions inside the figures in Japanese, and the operator still cannot read it. Handling this requires pairing translation with AI-OCR: detect text regions in the image, extract the text, translate it, and place it back in position. That last step — placing it back — is the hard one, and full automation is often not realistic.
Before committing, count the proportion of in-scope documents that contain text inside images. Decide to “automate work standard translation” without knowing that number and you will get half the expected benefit after go-live.
The “Readable but Ignored” Problem
The other ambush sits on a different axis from translation quality altogether. A translation can be grammatically correct and still get skipped, because the shop-floor term in it is head office’s literal rendering rather than the name people actually use.
Suppose head office calls an activity 予防保全 — preventive maintenance. On the Thai shop floor the same activity is called something else, often an English-derived term. Put the literal translation in the standard and the staff reading it do not recognize it as describing their own work. The document is correct, it will pass an audit, and behavior does not change.
What belongs in the term base is not the correct translation but the term actually in use on the floor. Determining that means asking shop-floor leaders, not the translation function. Outsource term base construction wholesale to a translation vendor and you will receive a list of terms that are correct and unused.
Designing for Adoption
Installing a tool does not change operations. With translation automation, the job that changes is not the job of the person who used to request translations; it is the job of the person who checks them. Staff who previously sent work out to a vendor now sit in the position of approving machine output. When the role changes, training is required.
We have covered the usual reasons generative AI training fails to stick in four reasons generative AI training does not take hold in manufacturing. The same wall shows up across AI-driven process automation generally, not just in translation, so it is better to budget for training inside the rollout plan from the start.
In Thai manufacturing, the language barrier in on-the-job training is a documented source of learning load, and developing materials for Thai-speaking staff is a standard recommendation. Translation automation is, among other things, a way to make that material development cheaper. Count only the reduction in translation hours as your benefit and this dimension disappears from view.

Cost Model for a Mid-Sized Japanese-Owned Plant in Thailand
The following is an annual model for a Japanese-owned plant in Thailand with roughly 300 employees. These are assumed values, not measurements — standard levels observed across multiple engagements. Treat them as a skeleton to recalculate with your own numbers.
Assumptions
| Item | Value used | Note |
|---|---|---|
| Headcount | 300 | Including 3–5 Japanese expatriates |
| Class A controlled documents | 1,200 pages/year | Including revisions. A4, 400 Japanese characters/page equivalent |
| Class B operational documents | 6,000 pages/year | Daily reports, minutes, internal notices |
| Class C external documents | 400 pages/year | Customer audit responses, submissions |
| Class D shop-floor UI text | 150 page-equivalents | Screens, forms, display labels |
| Outsourced translation rate | 600 THB/page | Japanese to Thai technical documents |
| Internal staff cost | 300 THB/hour | Administrative staff |
| Japanese manager cost | 900 THB/hour | Engaged in approval and audit response |
| LLM API cost | 1.5 THB/page | Assumes the Thai token overhead |
All figures are set in THB. Exchange rates move, so avoid comparing in yen; deciding in THB is the practical approach. Vendor lead time is also a design variable. For translating quality manuals and inspection checklists from Japanese to Thai for automotive parts makers, translation companies have quoted turnaround on the order of three weeks. For controlled documents that get revised often, those three weeks land directly as a delay in getting the revision to the floor.
Three Scenarios Compared
| Cost line (THB/year) | 100% outsourced | 100% fully automated AI | Class-split |
|---|---|---|---|
| Direct translation / API cost | 4,650,000 | 11,625 | 251,400 |
| Initial setup and design | — | 150,000 | 570,000 |
| Approver review effort | 252,000 | — | 252,000 |
| Terminology checking and rework | 180,000 | 950,000 | 154,600 |
| Ordering and acceptance admin | 72,000 | — | 20,000 |
| Fixing broken UI layouts | — | 250,000 | — |
| Explaining at audit | 45,000 | 90,000 | 20,000 |
| Annual total | 5,199,000 | 1,451,625 | 1,268,000 |
The bold row is the total. The top two lines are costs that appear on a quotation or an invoice; the bottom five are costs that never appear on a document and instead disappear as internal time. That is where the gap opens, and that gap is the argument of this article.
Approver review effort is loaded identically into the 100% outsourced and class-split columns. Approval signatures for Class A and Class C occur regardless of who produced the translation, and sending work to a vendor does not make them disappear. The 100% fully automated column is blank on that line only because that scenario never seated an approver at all — it is not cheaper, it has skipped a step. Note also that the 100% outsourced column assumes, for comparability, that Class D shop-floor UI text is also sent to the vendor with display-width constraints written into the order specification, which is why no layout rework cost is booked against it.
How to Read the Table
The 100% outsourced scenario is dominated by direct cost. Roughly 89% of the total is payment to translation vendors and the rest is internal effort. It is easy to manage, and it is structured so that cost rises in lockstep with volume.
The 100% fully automated scenario collapses direct cost and inflates terminology checking and rework to 950,000 THB. That line absorbs the effort of retroactively correcting Class A documents where terminology drifted, reissuing them, and redistributing them. The 250,000 THB of UI layout rework appears in this column and nowhere else for the same underlying reason: there is nobody to hand the display-width constraint to as a specification.
Most importantly, there is a risk the fully automated total does not show. A controlled document with no record of approval sign-off does not cost money on a line item, but it does become a finding at audit. The cost when that turns into a corrective action depends too heavily on circumstances to be booked here. Choosing full automation because the total is lowest is a decision that overlooks that omission.
The class-split scenario keeps the 252,000 THB of approval effort at exactly the level of the outsourced scenario, while pulling down both direct cost and rework. Approval is not a reduction target; it is a cost retained deliberately. Because the term base is doing its job, terminology rework falls to 154,600 THB, and because an approval history exists, the effort of explaining at audit stays at 20,000 THB.
What This Model Deliberately Leaves Out
The 570,000 THB of initial setup in the class-split scenario covers term base construction, translation memory preparation, and carving Class D text out into design. A substantial share of that is one-time cost weighted to year one and falls thereafter. But the payback period in years depends too heavily on translation volume growth and document revision frequency for us to justify a figure, so we are not putting one here.
Likewise, the cost of a quality incident caused by mistranslation is not booked, because there is no defensible basis for assigning a probability. If you encounter a cost model that includes numbers of that kind, check where the assumptions came from.
A 90-Day Rollout
Rather than building an elaborate plan, it is faster to run one full 90-day cycle and come out with measured numbers. Here is the minimum viable sequence.
Days 1–30: Inventory and Class Assignment
Spend the first month writing down every piece of translation work that currently happens. Who translates what, how often, and how many hours it takes. Separate what goes to vendors from what is handled internally.
Then assign each documented item to one of the four classes. Do this with quality assurance, manufacturing, and administration all in the room. Done by one person, the Class A / Class C boundary will always be drawn inconsistently. At the same time, count the proportion of documents containing text inside images.
Days 31–60: Build the Term Base and Pilot Class B
Month two runs terminology layer construction and a Class B pilot in parallel. Start the term base at 200 frequent terms and fill it in by confirming actual shop-floor names with line leaders. Do not let the translation function decide alone.
For Class B, pick one or two use cases — automated meeting minutes, automatic translation of daily reports, automatic generation of routine reports — and actually run them. The purpose here is less about measuring benefit than about establishing an operating pattern. Build the habit first in the class where failure costs little.
Days 61–90: Measure TTE and Trial Class A
In month three, trial the draft-plus-human-approval pattern in Class A. Limit the scope to a handful of frequently revised work standards. Measure TTE here and compare it against translating from scratch.
At the same time, compile the actuals from the Class B operation. What you should be holding at the end of 90 days is: measured TTE by class, a 200-term term base, an operating track record for Class B, and your own numbers to support the next investment decision. With those four in hand, you can decide on a full rollout.
| Period | Main activity | Deliverable at period end |
|---|---|---|
| Days 1–30 | Inventory of translation work, four-class assignment | Document class register, in-image text ratio |
| Days 31–60 | Term base construction, Class B pilot | 200-term term base, Class B operating procedure |
| Days 61–90 | Class A trial, TTE measurement | Measured TTE by class, investment decision pack |
Frequently Asked Questions
How much of enterprise translation can actually be delegated to AI?
It depends on the document class. Operational documents — daily reports, minutes, internal notices — are fine fully automated, with a few spot checks a month retained. Controlled documents such as work standards and quality manuals go as far as a machine draft, with human approval. Customer submissions and contracts are translated by people, with AI used for consistency checking. Display text such as HMI screens comes out of the translation process entirely and is handled as design. Any attempt to answer “how far can we delegate” with one company-wide rule will break somewhere.
What does AI translation cost?
At the document volumes modeled in this article, the API cost itself lands in the ten-thousand-odd THB per year band. The costs that matter are the surrounding ones: term base construction, approval workflow setup, and cleaning up existing documents run to several hundred thousand THB in initial cost. When you request quotations, compare on the scope of that initial setup work, not on the API unit rate. Note also that Thai consumes more tokens for the same content, so applying an English-based rate card directly will put you over budget.
Does automated meeting minutes AI work in meetings that mix Japanese and Thai?
Yes, with conditions. It assumes either separate microphones per speaker or a quiet recording environment. In a noisy factory meeting room, or when several people talk at once, AI transcription accuracy drops. Fix the recording environment first. Company-specific terms and equipment names also need to be registered in a dictionary in advance or they will be misheard. Sharing the term base with the transcription side as well is an effective configuration.
Why is Thai the only language where translation quality won’t stabilize?
Several factors compound. Thai has no spaces between words, so the system must first determine where the string divides, and errors there break everything downstream. On top of that, training data volume is smaller than for English or Chinese. And company-specific and manufacturing terms transliterated into Thai script do not exist in any dictionary, so both the segmentation and the rendering go wrong. Registering these in the term base works because it acts directly on that third factor.
How should we build the term base (glossary)?
Three principles. First, start at around 200 frequent terms; aiming for coverage means never finishing. Second, decide renderings by asking shop-floor leaders, not by consulting a dictionary (the reasoning is in the “readable but not followed” section above). Third, build a mechanism that automatically harvests terms from approved translations and adds them. Maintained by hand alone, updates stop within six months.
What should we do with our existing translation vendor contracts?
You do not need to bring everything in-house. For Class C external documents, continuing to outsource is often the rational call. Commissioning a vendor to build the initial term base is also a legitimate option, though confirming shop-floor terminology remains your own job in that case. The goal is not to reduce outsourcing to zero; it is to pick the optimal processing path for each class.
What will an ISO 9001 audit ask for?
The scope is control of documented information under clause 7.5. Documents must be available and legible, identifiable, and protected from unintended alteration. For translated versions, you need to be able to explain which revision is approved and how it was distributed. If your management of approved terminology is documented, the explanation to the auditor becomes straightforward. Conversely, a state where raw machine translation output is distributed directly to the floor is difficult to explain.
Summary
Enterprise translation automation stalls because of design, not engine performance. As long as documents are handled without distinction, the question of who is accountable for which translation never gets settled.
There are three moves. First, split documents into four classes — controlled, operational, external, and shop-floor UI text — and vary the machine/human split by class. Second, build in three layers — terminology, generation, approval — with a loop that returns approved translations to the terminology layer. Third, replace the accuracy-percentage KPI with TTE (Time to Edit) and measure whether the work actually decreased.
In Thailand and across ASEAN, word segmentation, token efficiency, tone marks, text inside drawings, and the “readable but ignored” problem are the specific barriers. None of them is solved by engine selection. They are solved by how you design your document classes and how you build your term base. Starting with 90 days of class assignment and TTE measurement is the realistic first step.
Working out how many of your documents sit in which class, and where automation should start, is a perfectly good place to begin — no commitment required. TOMAS TECH implements IT, OT, and AI systems for Japanese-owned manufacturers in Thailand, and translation automation is something we are happy to discuss starting from an inventory of your current state. If you would like to compare notes while you are still evaluating, get in touch through our contact page.
References
- Machine Translation 2026 Assessment (Translated)
- Best LLM for Translation 2026 (Alconost)
- What is the Best LLM for Translation (Lokalise)
- Compliance Best Practices for Translation: 2026 Guide (Adverbum)
- Assuring Quality Through Document Control (QualityWeb360)
- WangchanThaiInstruct: Thai instruction-following dataset and evaluation (arXiv:2508.15239)
- Cultural and foundational capability benchmarks for Thai LLM development (arXiv:2410.04795)
- An empirical study of multilingual LLM tokenizers (arXiv:2606.15044)
- Thai translation services (Kawamura International)