When a Thailand operation evaluates Thai-language generative AI, a fluent Japanese or English demo is not proof of production fitness. The buyer should procure an operating capability: the system retrieves the right Thai document, preserves names, numbers, dates and negation, stops when evidence is missing, protects restricted data, and can be governed by local staff. This guide, current as of 31 August 2026, turns that principle into an RFP, a Thai test set, RAG evaluation, a sample 90-day PoC, acceptance testing and contract requirements.
Executive decision: procure operational feasibility, not a model name
Product proposals often lead with model names, parameter counts, context windows and polished responses. Those facts may help shortlist candidates, but they do not answer the business questions. Can the system distinguish “replaced” from “not yet replaced” in a Thai maintenance note? Can it retain a decimal point in a pressure reading? Can it cite the current purchasing policy rather than an obsolete copy? Can a supervisor understand why it answered?
Before issuing an RFP, the buyer should define five assets:
- The work process, permitted actions and prohibited actions.
- A Thai production-work corpus and business-approved answers.
- Risk-tiered mandatory gates and scored criteria.
- Automated checks plus human review by Thai-speaking practitioners.
- Contractual controls for monitoring, change, incidents and exit.
This makes the evaluation reusable when a provider changes its model. Starting with a preferred model and inventing a use case later reverses the logic: the scorecard will be shaped around the product’s strengths. The procurement goal is not “the best AI in general.” It is a system that remains controllable in the buyer’s real workflow.
Why a Japanese or English demo misses Thai production risk
Head-office demonstrations usually use clean documents, short questions and carefully chosen examples. Thai operations handle mixed Thai and English part names, department abbreviations, OCR noise, Buddhist Era and Gregorian dates, omitted units, speech-like daily reports and policies with multiple versions. Data state matters as much as language: approved versus draft, current versus obsolete, and accessible versus restricted.
A response can sound natural and still be unacceptable if it:
- summarizes “not replaced yet” as “replacement completed”;
- confuses part AB-120 with AB-102;
- changes 3.5 bar to 35 bar;
- converts a Buddhist Era date incorrectly;
- cites an obsolete purchasing rule as current;
- invents an approver, price or due date; or
- answers in English when Thai was explicitly required.
Do not hide these errors inside one average score. Safety, quality, payment and employment-related items need mandatory gates. One prohibited output may be sufficient to fail a candidate even if its writing style scores highly. The Qwen-SEA-LION-v4.5-27B-IT model card, for example, treats an answer in the wrong required language as a failed evaluation response. That is not evidence that one model is superior; it is a useful example of specifying evaluation behavior before testing.
Define the Adopter, Customizer or Maker scope before the RFP
The ETDA Generative AI Governance Guideline for Organizations describes three organizational application patterns in increasing complexity: Adopter, Customizer and Maker. An Adopter uses an off-the-shelf service. A Customizer builds a tailored solution, for example through RAG or further adjustment. A Maker develops a foundation model. The categories do not rank quality; they clarify what data, talent, governance and contractual responsibility the organization must carry.
| Decision area | Adopter | Customizer | Maker |
|---|---|---|---|
| Typical scope | Ready-made GenAI service | RAG, workflow and adaptation | New foundation-model development |
| Buyer evaluates | Tenant controls, data terms, standard functions | Retrieval, integration, prompts and change control | Training data, model development and full lifecycle |
| Thai evaluation | Fit of real inputs and outputs | Retrieval and generation tested separately | Layered evaluation from training through operations |
| Core RFP question | How are inputs, logs and access controlled? | Which evidence is retrieved and how is it cited? | Who continually assures quality, safety and rights? |
A Japanese manufacturer in Thailand may use both an Adopter tool for general drafting and a Customizer system for internal documents. Classify each use case. Otherwise, procurement may demand bespoke behavior from a standard license or evaluate a RAG project as a simple subscription comparison.
ETDA’s AI 2026 direction is framed as “Driving Trust AI Governance.” Trust is not a feature label. It is produced by accountable owners, test evidence, rules of use, monitoring and corrective action. The technical specification and the operating model therefore belong in the same procurement decision.

Build a Thai production-work corpus and approved-answer set
Public benchmarks inform the design; company data decides acceptance
Public datasets can teach evaluation patterns. The SEA-NLI dataset card describes 2,160 examples: 1,443 “normal” and 717 “hard.” It also describes the involvement of native speakers and linguistic experts and discloses limitations. That is useful design evidence, but it contains neither your equipment names nor your approval rules. A public benchmark score is not a factory acceptance result.
Build company cases from realistic material rather than rewriting everything into demonstration-friendly prose. Before using personal, customer or confidential data in an evaluation environment, define purpose, minimization, masking, access, retention and deletion. Obtain organization-specific advice from the DPO or legal team for PDPA and any other applicable obligations.
Give every case an input, answer, prohibition and source
| Process | Realistic input | Approved answer data | Typical fail condition |
|---|---|---|---|
| Factory daily report | Colloquial Thai, English asset name, measurements | Facts, asset, time and downtime reason | Negation reversal, changed number, invented cause |
| Maintenance | Failure history, standard, bill of materials | Machine, procedure, warning and revision | Unapproved procedure or wrong part |
| Quality | Defect record, inspection rule, corrective report | Lot, specification, disposition and evidence | Changed disposition or missing evidence |
| Purchasing | Quotation, contract and approval rule | Price, currency, lead time and approval condition | Mixed supplier, price or currency |
| Human resources | Work rule and request FAQ | Eligibility, workflow and revision | Personal-data exposure or invented employment decision |
A test record should contain more than a question. Store expected facts, acceptable variation, required citations, prohibited claims, risk tier, reviewer and source version. If a consultation has no single correct answer, the expected behavior may be a clarifying question or escalation to the responsible department. That is more defensible than rewarding a confident guess.
Deliberately include difficult Thai business features
The set should include:
- entity names: people, Thai industrial estates, suppliers, machines and products;
- numbers: decimals, separators, zero, negative values, ranges, units and currencies;
- dates: Gregorian and Buddhist Era, deadlines, relative dates and shifts crossing midnight;
- negation: not done, no issue found, replacement not required and scoped negative clauses;
- politeness and register: supervisor, operator, customer and abbreviated chat language;
- department shorthand whose meaning changes between production, quality, maintenance and purchasing;
- mixed language: Thai sentences with English part codes, Japanese-origin shop-floor terms and symbols; and
- noisy sources: OCR errors, damaged tables, extra line breaks and old templates.
Create paired cases where a small wording difference reverses the decision: “replaced” versus “not replaced,” or “approval required” versus “approval not required under this condition.” Separate the person who drafts the answer from the business owner who approves the gold record. Ambiguity between human reviewers is itself a risk signal; do not ask AI to decide a process that the organization has not defined.
Treat Thai LLM model cards as shortlist evidence, not a warranty
The Typhoon2.1-Gemma3-12B model card describes a text-only 12B model with Thai and English as primary languages and a 128K context length. The Qwen-SEA-LION-v4.5-27B-IT model card describes a 27B model, a 262K context length and post-training that includes Thai and other Southeast Asian languages.
These are publisher-stated specifications. They do not guarantee accuracy on your documents, safety, latency, cost or operational fit. More parameters, a longer context window or Thai in a language list is not a procurement winner by itself. A hosted service and a self-hosted model also create different responsibilities for audit, patching, upgrades and incident response.
Use model cards to screen licenses, deployment options, regional availability, technical interfaces and declared limitations. Run finalists against the same document snapshot, retrieval configuration, output constraints and blind test set. Record the model version, prompt version, index configuration and execution time so a later result can be reproduced.
Create an RFP matrix with mandatory gates and weighted criteria

An average score can allow a severe failure to be offset by attractive usability. Separate mandatory gates from scored comparison criteria. All weights and thresholds below are labelled samples; your organization must set them from its own risk assessment.
| Type | Criterion | Required evidence | Sample decision rule set by buyer |
|---|---|---|---|
| Gate | No unauthorized transfer of restricted data | Architecture, contract and traffic record | Pass only with zero critical violation |
| Gate | Respect safety and quality prohibitions | Adversarial-test log | Zero prohibited response in defined set |
| Gate | Abstain or clarify when evidence is missing | Unknown-question set | Uses specified safe behavior |
| Gate | Answer in Thai when Thai is required | Language test log | Required language in every gated case |
| Score | Facts, numbers and dates | Gold-answer comparison | Example: 30 points, buyer sets weight |
| Score | RAG retrieval quality | Recall@k and error analysis | Example: 20 points, buyer sets weight |
| Score | Thai naturalness and work fit | Blind practitioner review | Example: 15 points, buyer sets weight |
| Score | Operations, audit and change | Procedures, SLA and evidence | Example: 20 points, buyer sets weight |
| Score | Cost and response performance | Load test and commercial terms | Example: 15 points, buyer sets weight |
Specify the evidence format in the RFP rather than accepting a vendor declaration. Ask for test logs, failed-case lists, a data-flow diagram, subprocessors, escalation routes, version-change notifications, recovery steps and Thai user material. For the demonstration, use blind cases supplied by the buyer and do not share approved answers in advance.
Combine automated and human evaluation
Automated tests are effective for exact entities, JSON shape, part numbers, measurements, dates, citation identifiers, latency, refusal behavior and repeatability. Human review is necessary for Thai naturalness, appropriate politeness, ambiguity, the quality of clarifying questions and practical usefulness.
A defensible evaluation flow is:
- Version the prompts, cases and approved results.
- Run each candidate repeatedly under fixed conditions.
- Automatically check format, facts, numbers, citations and prohibited behavior.
- Have Thai-speaking practitioners assess anonymized outputs.
- Send disputed cases to the accountable business owner.
- Classify failures and add them to the regression suite after correction.
Give reviewers a rubric rather than asking whether they “like” the answer. Possible dimensions are semantic accuracy, completeness, ambiguity, register and clarity of next action, with a separate critical-error flag. Blind the candidate name where possible. Reviewer disagreement should be captured, not averaged away.
Evaluate RAG retrieval separately from generation
A wrong RAG answer can have several causes: the right document was not retrieved; it was retrieved but ranked outside the supplied context; or the generator received correct evidence and still changed the conclusion. A final-answer score alone does not show what to repair.
Retrieval tests
- Did the approved document appear within the top k results?
- Was the correct revision, site, department and language selected?
- Could the system retrieve content from tables, attachments and scanned sources?
- Did Thai spelling variation and shorthand reach the same authoritative record?
- Were inaccessible documents excluded from results and context?
Recall@k or Mean Reciprocal Rank may be used, but the buyer sets k and pass levels. High-risk use cases should also test the behavior when no approved source exists.
Generation tests
- Does every material claim have evidence?
- Are revision number and effective date correct?
- Are quantities, units, dates and negation preserved?
- Does the answer expose conflicts between sources rather than hide them?
- Does it avoid adding unsupported facts?
First hold the retriever fixed while comparing generators, then hold the generator fixed while comparing retrieval settings. Also test index update time, deletion, permission changes and rollback. Operational retrieval is a lifecycle, not a one-off benchmark.
Operate hallucination, confidentiality and PDPA controls by risk tier
The NIST Generative AI Profile, published on 26 July 2024 and shown by the source page as updated on 8 April 2026, provides a companion resource for incorporating GenAI-specific risks into AI risk management. It is neither Thailand-specific legal advice nor certification, but it is a practical frame for identifying, measuring, managing and monitoring risk.
| Risk tier | Example use | AI authority | Human control |
|---|---|---|---|
| Low | Rewrite general text or generate ideas | Draft only | Ordinary user review |
| Medium | Internal FAQ, report summary or purchasing search | Evidence-backed proposal | Responsible staff approves use |
| High | Safety, quality, HR or payment decision | No autonomous final decision | Authorized person makes decision |
This table is an example. Set your tiers from impact, reversibility, sensitivity and legal effect. In high-risk processes, limit AI to retrieval and drafting; do not let it finalize a disposition, operate equipment or send externally.
Training users not to enter confidential data is not sufficient. Combine data classification, access controls, tenant boundaries, DLP, masking, logging, retention and output controls. For personal data, have the DPO and legal team assess purpose, necessity, access, processors, retention and any cross-border handling under PDPA and other applicable rules. This article is not legal advice.
Hallucination controls should test citations, trusted-source boundaries, abstention, clarification and handoff. Incident evidence should allow the organization to reconstruct the input, output, source documents, versions, configuration, user and timestamp.
Sample 90-day PoC: create acceptance evidence, not a demo

The following is an example plan, not a universal duration or benchmark. Adapt it to process complexity, data readiness and internal review.
Days 1–15: fix the process and risk scope (sample)
Limit the first scope to one or two processes. Name the business owner, IT, security, DPO/legal contact and Thai reviewers. Capture a baseline such as current handling time, rework or inquiry volume. Agree on prohibited data, prohibited autonomous actions and stop conditions.
Days 16–35: build the corpus and approved answers (sample)
Select representative and difficult daily-report, maintenance, purchasing or HR cases, then mask data as required. Add expected facts, citations, prohibitions and risk tiers. Prepare metadata for revision, site, language and access. The business owner approves answers even if preparation is outsourced.
Days 36–55: compare candidates under fixed conditions (sample)
Use the same retrieval settings, prompts and output format. Conduct blind review. A candidate that fails a mandatory gate does not advance merely because its average score is high. Classify each failure as document, retrieval, prompt, model, permission or operating-model related.
Days 56–75: conduct a controlled user pilot (sample)
Let a limited group of local staff work in a near-real environment. Apply full human confirmation or an organization-defined review sample. Record new wording, missing sources, misuse, confusion and education needs. Measure verification effort and stop decisions, not only convenience.
Days 76–90: run acceptance and make the transition decision (sample)
Freeze the evaluation set and run the final test. Present gates, scores, residual risks and operating cost for approval. Distinguish accepted, conditionally accepted, another PoC and stopped. A conditional acceptance should state users, sites, data, duration and release conditions.
Acceptance-test checklist
Acceptance should be a contractual completion test, not a statement that the demonstration worked.
- Pass the mandatory gates on the frozen Thai test set.
- Retrieve and cite the right revision and source.
- Keep critical entity, number, date, unit, negation and abbreviation errors within buyer-set limits.
- Exclude unauthorized documents from retrieval, context, output and logs.
- Follow defined behavior for unknown questions, conflicting sources and prompt injection.
- Re-run regression tests after model, prompt, index or policy change.
- Demonstrate audit logs, alerts, stop, recovery and support routes.
- Supply Thai operating instructions and training material.
- Validate response, concurrency and cost limits under realistic conditions.
Do not offset a prohibited safety or quality answer with a high general FAQ score. Preserve every failed case and make it part of future regression testing.
Contract clauses for evaluation, change and accountability
The agreement or SOW should define the capability that must remain operable. Qualified counsel should draft or validate legal wording.
- Scope and exclusions: process, sites, languages, users, data and prohibited use.
- Acceptance: test-set version, gates, scoring, retest and evidence retention.
- Change notice: model, version, prompts, retrieval and subprocessors.
- Regression: who runs it, when, with which set and at whose cost.
- Data: input, output and log retention, training use, deletion and return.
- Security: access, encryption, vulnerability, incident notice and audit support.
- Service levels: response, support and recovery, not only uptime.
- Intellectual property: source content, configuration, evaluation data and deliverables.
- Exit: export, deletion evidence, logs and fallback process.
- Accountability: AI proposal, human approval, external transmission and business decision.
“Supports Thai” is not an acceptance clause. Attach the Thai processes, outputs, prohibited errors and test procedure. If the service changes models automatically, consider a right to pause high-risk use until regression tests pass.
Make local-staff AI education part of acceptance
Local-staff AI education must cover more than prompt tips. Staff need Thai examples of what data may be entered, what the output cannot decide, how evidence is checked and where an error is reported.
Train by role. Users practice classification, questioning, source checks and escalation. Business administrators review failures, permissions and stopping rules. IT administrators manage configuration, logs, updates and regression. Executives approve residual risk and scope.
Use scenario exercises rather than attendance as proof of understanding. Show a request containing restricted information, an unsupported answer and conflicting dates, then assess the correct response. Maintain a Thai support channel and a no-blame reporting route.
For the organizational phase after evaluation, see AI adoption support for Thailand operations. For deploying an off-the-shelf productivity environment, see Microsoft Copilot implementation in Thailand.
Final checklist for Japanese companies using AI in Thailand
In the decision meeting, examine how failures remain—not only the average score
Do not present only the rank by total score. Put mandatory-gate results, critical errors, unresolved cases, excluded processes, risks controlled operationally and risks allocated contractually on one decision sheet. Two candidates with the same score are not equivalent if one has many minor style issues and the other has fewer but material errors in negation or price.
For each failure, record the reproducible input, expected result, actual result, source, cause hypothesis, interim control, permanent correction and owner. Do not close every issue with “improve the prompt.” Compare document revision control, permission metadata, retrieval chunking, interface warnings and human approval. After a correction, run a surrounding regression set rather than only the original failed case.
Put expected benefits and residual risks in the same management view. Benefits should include search, verification, rework, education and audit effort, not just generation time. If management uses a financial benefit or reduction rate, label it as company-measured PoC data; do not borrow an unrelated public percentage. If evidence is insufficient, choose conditional approval or additional measurement instead of inventing a number.
Go-live is not limited to yes or no. The organization may restrict departments, data classifications, output to draft status, require full review for a defined period or freeze model updates. Every temporary control needs an expiry date and release criterion. If prohibited behavior persists, evidence cannot be reconstructed or change notice is unavailable, stopping remains a valid decision even when the demo is convenient.
Preserve results for candidates not selected. Future price, model or regional-availability changes may require reevaluation, and the same evidence supports a fair comparison. If evaluation data is sensitive, govern what may be shared with vendors, where it is retained and when it is deleted. This turns testing into an organizational procurement capability rather than a one-time demonstration.
- A Thai production-work set exists separately from the Japanese/English demo.
- Thai business owners approve the gold answers.
- Entities, numbers, dates, negation, politeness and abbreviations are tested.
- RAG retrieval and generation are evaluated separately.
- Mandatory gates are separated from weighted scores.
- Model-card specifications are not treated as performance warranties.
- Abstention, clarification and human handoff are tested.
- Confidentiality, personal data, permissions, logs and retention are designed.
- Regression can run after model or retrieval changes.
- Thai training, support and incident procedures are ready.
- Acceptance and change control appear in the contract or SOW.
- Quality, cost and permitted scope are reviewed after go-live.
Summary
The right way to procure Thai-language generative AI is not to declare a winner from a public specification or a fluent demo. Build a Thai production corpus, business-approved answers, risk-tiered gates, automated and human evaluation, separate RAG tests and repeatable regression. Model names will change; your evaluation assets and acceptance evidence remain reusable.
TOMAS TECH can support an early-stage Thai test-set design, RFP matrix or controlled PoC before a product is selected. We connect Thailand operations with Japanese-head-office governance and turn requirements into evidence. Contact TOMAS TECH.
FAQ
Which Thailand generative AI use case should be evaluated first?
Start with a process where approved answers can be checked and the effect of an error can be contained. Internal retrieval or drafting may be candidates, but select by reversibility, data sensitivity and human-review capacity. Keep quality, safety and payment decisions as human decisions unless a separate high-risk case has been approved.
What must a Japanese company include in a Thailand AI RFP?
Specify the Thai test set, prohibited use, data handling, permissions, evidence, behavior when uncertain, audit logs, change notice, regression and acceptance. “Thai support” and one overall accuracy number are not enough. Create mandatory gates for high-impact errors.
Is a public benchmark enough for Thai LLM evaluation?
No. It helps shortlist or design tests, but it does not contain your equipment, forms, abbreviations, policy versions or permissions. Acceptance requires protected real-work cases and approved answers owned by the business.
What belongs in local-staff AI education?
Teach permissible data, source verification, checking of numbers/dates/negation, prohibited decisions, stopping and reporting. Test understanding through Thai scenarios and include correct operating behavior in acceptance.
Which Thai generative AI model performs best?
Public information does not justify a universal ranking. Suitability changes with the workflow, documents, risk, latency, cost and deployment model. Compare candidates with the same company set and include operational controls and change management. Model-card values are not a capability guarantee.