Selecting a Bangkok AI company on the strength of a polished demo or a low development rate creates expensive uncertainty. A buyer must assess problem definition, Thai-language data, legacy integration, PDPA controls, AI governance, acceptance testing, source-code rights and operational handover as one procurement package. This guide gives Japanese and international companies in Thailand a practical route from vendor screening and RFP to PoC, FAT/UAT, SLA and production ownership.
The answer: select an AI development company by evidence you can accept and operate
An AI project does not fail only when the model refuses to run. It also fails when a prototype works but cannot process the spelling variants used by Thai operators, an ERP master change degrades the result, nobody can trace the source behind a generated answer, or only the original vendor can update prompts and evaluation tests.
The evaluation should therefore answer five questions:
- Can the vendor explain the business problem and a non-AI alternative?
- Are responsibilities for data, rights, PDPA, safety and security explicit?
- Can it run reproducible tests on representative Thai, English and Japanese data?
- Can it integrate with ERP, MES, documents, equipment and approval workflows safely?
- Can the customer monitor, change, restore and transfer the service after acceptance?
This treats AI as an operating business system, not a one-off experiment. For an earlier sourcing decision, see our guide to AI development outsourcing in Thailand. For test planning, use the detailed guide to AI PoC costs and success criteria.
Why Thailand AI development procurement needs stronger discipline now
Thailand BOI reported THB 1.47 trillion across 1,299 investment applications in the first half of 2026, with application value up 37% year on year. These are application figures across the announced investment portfolio, not approved or realized investment and not an AI-only total. BOI also published an August 2026 update on AI and technology inflows as Thailand prepared its national semiconductor and advanced-electronics strategy. Together, the releases show a fast-moving environment for compute, cloud and digital services; they do not certify any individual supplier.
Market growth can make selection harder. Labels such as “AI-ready,” “generative AI developer” and “agent builder” reveal little about cross-border data, Thai evaluation, audit logs, model updates or incident recovery. Without a common buyer specification, vendors quote different scopes. Prices, schedules and claimed accuracy then cannot be compared fairly.
ETDA promotes AI governance as a foundation for trustworthy use and provides organizational support through its AI Governance Practice Center. ISO/IEC 42001 specifies requirements for establishing, operating and continually improving an AI management system. Neither guarantees the accuracy of a product. They are useful because they move procurement toward accountable roles, risk records, monitoring and improvement.
Four useful types of Bangkok AI company
Classify candidates by the capabilities the project needs rather than reputation alone.
| Vendor type | Typical strengths | Questions to test |
|---|---|---|
| AI specialist or startup | Model experiments, GenAI, specialist talent, speed | ERP integration, long-term support, staffing continuity, handover |
| SI or enterprise software firm | Requirements, ERP/MES/API, operations, contracts | Depth of AI evaluation, current methods, experimentation speed |
| Cloud or product partner | Platform, security controls, standard services, scalability | Lock-in, business fit, cost at higher usage |
| Factory and OT integrator | Site work, equipment, networks, maintenance, Thai operations | Data science, GenAI assurance, MLOps experience |
One supplier may not cover every layer. If the delivery team combines a prime contractor, an AI specialist, a cloud provider and an OT integrator, name one incident owner and one party accountable for the overall acceptance result. Contract the conditions for changing subcontractors. A design that makes the customer arbitrate every supplier boundary will slow change and recovery.

AI development company scorecard: fix weights and disqualifiers before proposals
Scoring criteria created after reading proposals invite brand and presentation bias. Agree on weights, minimum scores, disqualifying conditions and evidence formats before issuing the RFP. The following 100-point model is an editorial example, not an industry benchmark.
| Area | Sample weight | Evidence to request |
|---|---|---|
| Business understanding and value | 15 | Current process, KPI definition, non-AI alternative |
| Data and Thai-language capability | 15 | Data diagnosis, Thai test set, missing-data treatment |
| AI quality and evaluation | 15 | Baseline, reproducible tests, error taxonomy |
| Integration and architecture | 15 | API, identity, audit, ERP/MES integration, failure design |
| Security and PDPA | 15 | Data flow, processors, storage, deletion, incident procedure |
| Operations, SLA and handover | 15 | Monitoring, restore, source, documents, training, staffing |
| Commercial continuity | 10 | Assumptions, change rates, licenses, subcontractors, viability |
| Total | 100 | Compare through one evidence template |
Possible disqualifiers include using confidential inputs for external training without authorization, inability to name storage locations, no material-incident notification process, undefined rights to deliverables, or testing on the same examples used to tune the system. The buyer must adapt these conditions to its information classification and legal advice.
Score evidence, not “yes” answers. Ask for anonymized samples, screens, test records, diagrams, procedures and interviews with the proposed delivery people. A vendor that can describe a relevant failure, root cause and correction often provides more useful evidence than a large but unexamined project count.
Build an RFP that makes AI adoption in Thailand comparable
A useful RFP does not pretend the buyer already knows the complete solution. It defines a common problem, reveals constraints and requires every bidder to label assumptions.
Business scope and exclusions
Describe users, sites, languages, inputs, outputs, approvers, operating hours, peak volume and current systems. Replace “automate customer enquiries” with a boundary such as: “classify Thai and Japanese maintenance enquiries, draft an answer with approved-source citations, and require a person to approve delivery.” State what the PoC must not do, such as automatic pricing, personnel decisions or unsupervised customer sending.
Baseline and target
Measure current handling time, volume, rework, misclassification, missing records and queue time. If the baseline does not exist, make baseline creation an initial deliverable. Avoid a single “95% accuracy” target. Define severe and tolerable errors. Missing a safety-related maintenance instruction is more material than a stylistic paraphrase.
Data inventory and conditions
List data owner, format, time period, language, personal-data status, confidentiality, location, refresh frequency, missingness and labels. Before raw data can be shared, provide a redacted schema, volume bands, synthetic examples and classification. Require bidders to identify every assumption.
Integration and non-functional requirements
Identify ERP, MES, CRM, SharePoint, file, email and machine-data connections. Define identity, network zone, concurrency, audit, maintenance windows, backup and recovery needs. Derive response and recovery objectives from actual business tolerance and criticality; do not insert universal numbers without a basis.
Deliverables and acceptance
List more than application source. Include a data dictionary, prompts and system prompts, evaluation data and scripts, model/service configuration, IaC, settings, APIs, runbooks, incident response, administrator training, license inventory and known limitations. Assign an owner, format and acceptor to each item.
Use a PoC to reduce a Go/No-Go uncertainty, not to stage a success demo
The AI adoption roadmap for Thai SMEs helps keep a PoC connected to deployment and operations. “Can AI do this?” is too broad. Test decision questions such as:
- Can the system classify high-impact cases despite Thai abbreviations, spelling variants and mixed alphanumeric input?
- Can it ground answers only in approved documents and abstain when evidence is missing?
- Can it integrate through actual API permissions without creating duplicate ERP transactions?
- Can the team measure latency and cost as usage grows?
- Can customer administrators update prompts, knowledge and thresholds and then roll back?
Evidence gates for a PoC
Duration varies by project. The sequence below is a governance example, not a fixed market schedule.
| Gate | Decision | Minimum evidence |
|---|---|---|
| G0 Problem | User, decision, baseline, alternative | Process map, baseline, risk register |
| G1 Data | Quality, rights, language, representation | Data profile, exclusions, test-set specification |
| G2 Technology | Improvement over baseline, severe errors | Reproduction steps, results, error taxonomy |
| G3 Workflow | Human approval, exception, training, UX | Scenario records, user observations |
| G4 Production | Security, performance, monitoring, cost | Load results, data flow, operating estimate |
| G5 Investment | Benefit, residual risk, next stage | Go, conditional Go or No-Go record |
Keep final evaluation examples separate from development feedback. Include critical but rare cases, native Thai, colloquial text, OCR noise, obsolete documents, conflicting sources and adversarial input. For generative outputs, exact-match accuracy alone is weak. Score grounding, prohibited behavior, coverage, abstention and human rework.
The NIST AI RMF Generative AI Profile is a companion resource for applying AI RMF to generative-AI risks. It is not a certification or vendor ranking, but it helps structure how risks are governed, mapped, measured and managed. OWASP Top 10 for LLM Applications 2025 highlights prompt injection, sensitive information disclosure, supply-chain risk, data and model poisoning, improper output handling and excessive agency, among other risks. Convert these headings into tests against the actual permissions and data flow.

Test Thai-language operations, not a translated demo
A Thai user interface is not proof of Thai business quality. Operating records commonly mix Thai and English, product codes, Roman characters, local abbreviations, levels of formality, speech transcripts and OCR errors. Japanese managers and Thai operators may use different terms for the same defect or maintenance event.
Stratify the evaluation set across relevant sites, departments and shifts. Include content originally written in Thai, not only Japanese test cases translated after the fact. Reviewers need both language fluency and domain knowledge; grammatical Thai and a correct production decision are different outcomes.
Version the terminology list, prohibited expressions, approved answers, language-specific prompts and evaluation sets. Assign who updates them after a new product, process, regulation or organization change and which regression tests must run. “The foundation model is multilingual” is not an acceptance criterion.
Manufacturing AI requires OT and safety boundaries
When AI influences equipment, quality disposition, work instructions or production planning, its impact is larger than an ordinary web feature. The architecture must show whether AI writes to a PLC, whether a person or deterministic logic approves the output, and how the process enters a safe state on failure.
For some early deployments, AI can remain an advisory layer separated from deterministic safety control and machine interlocks. The appropriate structure depends on risk assessment; adding AI does not itself improve safety. After model or prompt changes, regression testing should cover output ranges, timeouts, abnormal values and manual fallback as well as model quality.
depa’s AI & Digital Transformation Catalog can be an entry point for exploring solutions available in Thailand. Inclusion must not be interpreted as a performance guarantee, legal compliance finding or TOMAS TECH endorsement. Each option still needs project-specific RFP evidence and acceptance tests.
Put data rights and PDPA controls into the contract
“We comply with PDPA” is not a complete control. Legal conclusions require advice for the specific processing, but procurement should document at least:
- Assumed controller, processor and subprocessors.
- Processing purpose, lawful handling process and data minimization.
- Countries and services holding source data, embeddings, logs, backups and test data.
- Whether provider settings permit inputs to be used for model training.
- Retention, deletion, backup treatment and return at termination.
- Access, download and administration permissions with audit records.
- Escalation for leakage, misdelivery, cross-border risk or re-identification.
Separate ownership from license rights. Define customer source data, customer-funded labels, vendor background libraries, customer-specific code, generic improvements, prompts, fine-tuned weights, evaluation sets and outputs asset by asset. A blanket statement that “all deliverables belong to the customer” may conflict with cloud APIs, open-source licenses and vendor background components.
Accept source code, models and operational handover as a working system
A Git repository is not a handover if it cannot be built, production privileges remain in a vendor employee’s account, or the test data is absent. Acceptance should include deployment into customer-controlled accounts from a clean environment, a restore from backup, and rerunning representative tests.
| Layer | Handover package |
|---|---|
| Code | Application, integration, evaluation, migration, IaC, dependencies, build steps |
| AI configuration | Model/version, parameters, prompts, guardrails, routing |
| Data | Dictionary, schema, labeling rules, test sets, provenance, rights, retention |
| Environment | Customer accounts, access matrix, secret handling, network diagram |
| Operations | Monitoring, alerts, incidents, change, re-evaluation, rollback, cost monitoring |
| Knowledge | Admin/user training, recordings, FAQ, known limits, open issues |
Consider source-code escrow where appropriate, included transition hours at termination and staffing continuity if a key person leaves. Maintain an inventory of open-source and third-party model licenses, API terms and end-of-support dates.
Connect FAT, UAT and final acceptance through one evidence model
FAT may cover functions and integration in a supplier-controlled environment, while UAT confirms business fitness with real users. Names matter less than documenting which risk is tested, in which environment, by whom and with which data.
FAT coverage
Freeze code, model and prompt versions. Test APIs, permissions, input validation, audit logs, error handling, timeouts, retries, idempotency and backup. Simulate an external AI service that is unavailable, slow or rate-limited. If generated output feeds another system, test that it is treated as untrusted input rather than executed as a command, SQL or markup without validation.
UAT coverage
Japanese and Thai users should operate normal, ambiguous, incomplete, prohibited, low-confidence and unauthorized scenarios. Record grounding visibility, correction effort, abstention, escalation, accessibility and training burden, not only an accuracy percentage. A nominal human-in-the-loop design is weak if the screen prevents the reviewer from detecting a plausible AI error.
Acceptance records
Every case should contain preconditions, input, steps, expected and actual results, evidence ID, configuration version, reviewer and open issue. Agree thresholds in the RFP or at the end of PoC. If a threshold is relaxed after seeing results, approve the reason and residual risk. Do not allow a high average to offset failure of a mandatory critical case.

An SLA should cover AI quality and change, not uptime alone
A cloud endpoint may be available while sources are stale, Thai classification has drifted or API cost is uncontrolled. Separate three layers of service indicators:
- System: availability, latency, errors, queue, integration, backup and recovery.
- AI quality: severe errors, unsupported answers, abstention, drift, performance by language and rework.
- Business: handling time, backlog, adoption, approval time and operational effect.
Not every indicator needs a financial service credit. Separate contractual guarantees, operating targets and observed metrics. For external cloud or model failure, define notification, degraded mode, alternate model or manual process and state reconciliation, not only liability exclusions.
Treat model versions, prompts, knowledge documents, retrieval settings, thresholds, code and schemas as configuration items. Even a small wording change can affect critical behavior. Keep evaluation, approval, release, monitoring and rollback in one change record.
Ask operational AI governance questions
Do not reduce ISO/IEC 42001 or ETDA guidance to a certificate checkbox. Check the scope and ask how the company works:
- Are use case, model, data, owner, risk and users recorded in an AI inventory?
- Who approves or stops a high-impact use case?
- When is risk assessment updated during design, release and operation?
- How are model information, data records, evaluations and changes retained?
- Who reviews fairness, explainability, privacy, safety and security?
- How are users told that AI is involved, what its limits are and where to report problems?
- How are harmful answers, leakage, prompt injection and excessive action detected and contained?
- How are provider changes and model retirement monitored?
If an AI agent can send email, place orders, issue equipment instructions or delete files, minimize its tools and permissions. Put approvals and limits around high-impact actions and audit inputs and results. OWASP’s excessive-agency category is directly relevant: do not give an agent more functionality, permission or autonomy than the task needs.
Compare estimate assumptions before comparing totals
Quotes with the same title can include very different work. Break out discovery, data preparation, labeling, translation, APIs, cloud, model use, security testing, user training, hypercare and maintenance. Do not rely on fabricated market rates. Provide a price sheet with the same quantities and assumptions.
| Cost layer | Quantity and assumption to normalize |
|---|---|
| Discovery | Departments, interviews, data sources, site visits, languages |
| Build | Screens, APIs, workflows, model candidates, environments, migration |
| Data | Records, pages, audio hours, labels, OCR, translation, anonymization |
| Platform | Dev/test/prod, log retention, storage, network, compute/GPU |
| Usage | Requests, tokens, retrieval, concurrency, peak and reprocessing |
| Operations | Coverage, support, incidents, reassessment, model changes, reporting |
| Exit | Data return, deletion evidence, documents, training, transition support |
Clarify volume tiers, currencies, minimums, cancellation, prepayment, third-party increases and change-request rates. A PoC quote is not a production price. Update low, base and high production scenarios with measured usage.
Governance for Japanese companies adopting AI in Thailand
Thai subsidiaries often involve Japan headquarters, the local company, factory users, IT, legal or HR and regional management. Sending every decision to headquarters slows delivery; deciding everything locally may conflict with group policy and contracts. Build a RACI that separates use-case owner, data owner, technical owner, risk acceptor and budget owner.
Language is part of accountability. Do not make critical decisions only in Japanese minutes. Thai users need to confirm acceptance conditions, prohibited use and incident steps. Even where the contract is in English, local procedures, screens and training should be usable in the operating language. Define the authoritative version and translation synchronization.
Steering meetings should review assumptions, unresolved risks, evidence, data rights, cost forecasts and Go/No-Go conditions—not only percentage complete. A procurement culture that punishes early red flags encourages concealment until acceptance. Reward prompt disclosure and reproducible evidence.
Final contract checklist
- Scope, exclusions, automation level and final decision owner are clear.
- Users, sites and Thai/Japanese/English evaluation are defined.
- Purpose, storage countries, subprocessors, training use, deletion and return are contracted.
- Rights to data, labels, code, prompts, evaluation sets and outputs are asset-specific.
- A non-AI baseline and critical-error definition exist.
- PoC Go/No-Go gates and an exit path exist.
- FAT/UAT environment, data, reviewers, evidence and version are named.
- ERP/MES/API duplicate, timeout, retry, cancellation and reconciliation are tested.
- Prompt injection, leakage, excessive permission and output validation are tested.
- SLA, degraded mode, notice, restore, change and model retirement are addressed.
- Customer accounts can deploy the source, configuration, documents and tests.
- Termination supports data return/deletion and transfer to another provider.
FAQ
How many Bangkok AI companies should we compare?
There is no universal number. Use an initial information request to test capability and disqualifiers, then invite a number whose questions and evidence your team can evaluate properly. Include more than one vendor type and compare all on the same assumptions.
What does an AI PoC cost in Thailand?
Cost varies with data preparation, integration, languages, security and evaluation. Compare decomposed deliverables, role-based effort, third-party charges, volume assumptions and change rates. A cheap demo and a PoC capable of supporting a production decision are different purchases.
Is it safe to give confidential data to an AI development company?
Safety cannot be inferred from a company name. Minimize or anonymize data, prefer controlled environments, and verify permissions, location, training-use settings, logs, deletion, incident notice and subprocessors. Obtain Thai legal advice for the actual processing.
How should Thai-language AI quality be evaluated?
Use an independent set containing native Thai, mixed English, abbreviations, OCR noise and rare critical cases. Domain-capable Thai reviewers should judge correctness, evidence, abstention and correction effort.
Should headquarters or the Thai subsidiary own the AI project?
Split accountability by use case, data, technology, risk and budget. Align group controls with Thai operations and legal requirements, and place local users in UAT and change approval.
Should an AI agent automate actions from day one?
Begin with low-impact, reversible actions, restricted permissions, approvals, limits, logs and a stop mechanism. Payments, external messages, equipment control and deletion should not receive broad autonomy without a risk assessment and acceptance evidence.
Conclusion: select for sustainable customer operation
A Bangkok AI company should be compared through business value, Thai data, integration, PDPA, AI governance, FAT/UAT, SLA, rights and handover—not demo polish, model names or a day rate. A PoC is an evidence gate that reduces production uncertainty. Final acceptance should prove that the customer can deploy, evaluate, restore and continue the service even if the original team changes.
TOMAS TECH can help structure a comparison scorecard or RFP while the project is still being defined. For AI adoption involving Thai operations, manufacturing sites or ERP/MES integration, contact us through the TOMAS TECH English enquiry page.
References
- Thailand BOI, Thailand Secures 1.47 Trillion Baht in Investment Applications in 1H 2026: https://osos.boi.go.th/EN/news/2430/Thailand-Secures-43-6bn-1H-2026-Investment-Surge-as-Big-Tec/
- Thailand BOI, Thailand AI and Tech Inflows Surge: https://osos.boi.go.th/EN/news/2462/Thailand-AI-and-Tech-Inflows-Surge-as-Country-Prepares-Natio/
- ETDA, Driving Trust with AI Governance: https://www.etda.or.th/th/pr-news/aigc_Driving-Trust_AI_Governance.aspx?feed=cb66f430-5546-4dd8-b279-3827e88d154b
- ETDA, AI Governance Practice Center: https://www.etda.or.th/th/Our-Service/AIGC/index.aspx
- ISO, ISO/IEC 42001:2023: https://www.iso.org/standard/42001?browse=ics
- NIST, AI RMF Generative AI Profile: https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence
- depa, AI & Digital Transformation Catalog: https://www.depa.or.th/storage/app/media/file/ai-digital-transformation-catalog.pdf
- OWASP, Top 10 for LLM Applications 2025: https://genai.owasp.org/llm-top-10/