Search “best AI development company” and you get listicles — 25 vendors compared, 50 vendors compared. Yet shortlisted projects still stall at PoC, break down at acceptance, or leave the model owned by the vendor. Success is decided by contract type, staging, acceptance criteria and IP allocation — not by the comparison table. Here is the order to decide them in.
Why AI work is hard to promise as “finished”
Comparison articles rest on an unstated assumption: pick a good company and what you ordered will get built. For conventional system development that assumption mostly holds. Build the screens and functions written in the specification and the job is done. Once an AI model is in scope, the assumption breaks down structurally.
Japan’s Ministry of Economy, Trade and Industry (METI) published a checklist on contracts for the use and development of AI on 18 February 2025. On performance and accuracy it makes a specific point: because AI models are developed by an inductive method grounded in training data, the limits of accuracy may already be baked into the training data itself. The consequence it draws is that setting an obligation to complete, or a performance guarantee, is difficult.
This is not a statement about vendor competence. It is a statement about information: no engineer, however good, can extract more than the data contains. The shop floor makes it concrete. If three years of inspection records say only “defect” — with no defect type, no process step, no operator, no ambient conditions — then no model built on that data will predict which conditions produce defects. And you usually only discover that after someone has spent a few weeks with the data.
So the first job for the buyer is not “find a good company”. It is to decide, on the assumption that the work may not complete, how much you are willing to pay and where you have the right to stop.
A note for readers working outside Japan. METI’s documents are Japanese public guidance, and they are worth reading wherever you sit, because they are the most detailed publicly available treatment of exactly this problem, and because for many Japanese manufacturers in ASEAN the counterparty, the group standard, or the parent-company approval process is Japanese. Use them as a checklist of issues rather than as applicable law — which law applies follows from the contracting entity, covered later in this article.
First, classify your own project
The same checklist splits AI-related contracts into three use-case types. Which one you are in changes the issues you need to close out, and even whether a bespoke contract is needed at all.
| Type | What it is | Buyer’s main concerns |
|---|---|---|
| General-purpose AI service use | Using an off-the-shelf AI service as-is | How input data is handled, retention period, deletion, security posture |
| Customized type | Using a service tuned for the user | The above, plus scope of tuning, ownership of tuning results, treatment on termination |
| New development type | Co-developing a bespoke AI system | The above, plus whether an obligation to complete exists, acceptance criteria, ownership of the trained model |
The checklist positions itself as a way of “showing where the issues lie” when contracting for AI use, and is organised as 14 input-related items and 14 output-related items. The input side covers definition of the data the user supplies to the vendor, purpose of use, conditions of use, management and security arrangements, retention and deletion. The output side covers obligation to complete, conditions of provision, purpose of use, conditions of use and security management.
The buyer’s first piece of work here is to run those 28 items against your own project and identify the items you do not yet have an answer for. Doing this before you talk to any AI development company changes the quality of the conversation entirely, because questions like “do you anticipate using our data on other clients’ projects?” and “when and by what procedure is it deleted?” then come from a systematic list rather than from whatever occurs to you in the room.
Classifying into one of the three types also feeds directly into the build-versus-buy line — how much you keep in house and where outsourcing starts. We set out a layer-by-layer view of who owns what in our article on AI insourcing support.
Exploratory staged development — splitting the contract into four phases
If completion is hard to promise, how should you contract? Japan’s official answer sits in METI’s Contract Guidelines on the Use of AI and Data, published in June 2018 and partially revised in December 2019. The AI volume of those guidelines proposes a development method known as exploratory staged development.
The idea is simple. Do not contract for the whole thing up front. Break the development process into phases, contract phase by phase, and decide at the end of each one whether to continue or stop. The guidelines set out four phases: assessment, PoC, development, and additional learning. In the table below the phase names come from the guidelines; the “what happens” and “what the buyer holds at the end” columns are our own summary of common practice.
| Phase | What happens | What the buyer holds at the end |
|---|---|---|
| 1. Assessment | Use a limited sample of data to judge whether a model is feasible at all | A technical feasibility report. “Not feasible” is a legitimate result |
| 2. PoC | Build a prototype on real data and test whether accuracy is good enough for the job | A prototype and test results. An outlook expressed in business metrics |
| 3. Development | Build the production model and surrounding system on the validated approach | Trained model, surrounding systems, documentation |
| 4. Additional learning | Add data during operation to maintain and retrain the model | Sustained accuracy, ability to follow changes in the business |
The guidelines set out contracting method, factors to consider, and sample clauses for each phase. They also note that, given that development of an AI model on its own proceeds by an inductive method dependent on data, a quasi-mandate contract (junin’nin) is a natural fit for that work.
The real point of phasing is creating a place to stop
Phasing sounds like a delivery-method topic. For the buyer, the practical value lies elsewhere. The point is creating a point at which you can legitimately terminate.
Run a project under a single lump-sum contract and, once it becomes apparent that the thing is harder than expected, you cannot stop — the contract is still running. Stop and you are into a breach discussion; continue and you keep funding an investment with no prospect. That is the anatomy of “stuck at PoC, and somehow still spending”. Under phased contracts, if the assessment phase concludes that the project is not feasible with the data as it stands today, you pay for that phase and end it. The value is in being able to end it.
One point deserves emphasis. When the assessment phase concludes “not feasible”, that is not a failure. In the sense that you avoided a full investment decision for the cost of the investigation alone, it is a success. If your organisation creates an internal climate in which PoCs must succeed, the teams on the ground will keep hopeless projects alive. If you introduce phased contracting, design the internal evaluation criteria at the same time as the contract.
Write entry and exit conditions for each phase
Splitting the work into phases is not enough on its own. For each phase, write entry conditions (when you may start this phase) and exit conditions (when this phase can be said to be finished) into the contract or into a phase-specific agreement.
For a PoC phase, an entry condition might read: “the assessment report concluded that the target data contains information sufficient for discrimination”. An exit condition might read: “the business metric measured against the validation dataset meets the level set out in Schedule X, or, where it does not, a root-cause analysis report has been delivered”. The trick is to include a deliverable for the case where the target is missed. Without it, a phase that goes badly leaves you with nothing at all.

Contract for work or quasi-mandate — and the “deliverable-completion quasi-mandate” middle ground
Once the phases are set, the next question is which contract type each phase uses. Two Japanese civil-law categories dominate the discussion, and IP BASE, run by the Japan Patent Office, frames them as follows.
| Type | What the vendor owes | Buyer’s main remedy |
|---|---|---|
| Contract for work (ukeoi) | Promises completion of the work and bears legal responsibility for performing it | Where the deliverable is non-conforming, can demand cure or damages |
| Quasi-mandate (junin’nin) | Performs to a given standard with the care of a good manager. Owes no obligation to complete | Pursuing breach of the duty of care. Cannot demand completion itself |
The difference is large, and leaving it vague is a reliable source of trouble. The buyer thinks “we ordered it, so it will be finished”. The vendor thinks “we explained that accuracy depends on the data”. The gap surfaces for the first time at acceptance.
Two variants of the quasi-mandate
The quasi-mandate splits further into two.
- Performance-ratio type: the obligation to pay arises when the delegated work has been performed
- Deliverable-completion type: the obligation to pay arises in exchange for delivery of the agreed deliverable
Performance-ratio type is, in plain terms, “pay for the work done”. It suits exploratory work like the assessment and PoC phases, where nobody knows in advance what will come out. From the buyer’s side, though, it means paying even if nothing comes out, which gets harder to justify internally as the amount grows.
Deliverable-completion type means the payment obligation arises only once a pre-defined deliverable has been handed over. That is easier to explain internally — “nothing delivered, nothing paid” — but it still does not mean the vendor owes an obligation to complete. So the deciding factor for whether you can use this type is whether you can define in writing what counts as the “deliverable”. If you can, it favours the buyer. If you pick deliverable-completion type while the definition stays vague, the dispute over whether delivery occurred is worse than anything performance-ratio type would have produced.
Assume the vendor will try to avoid a contract for work
Because AI accuracy depends on training data, vendors tend to avoid contracts for work. That is not bad faith; it reflects the structural reason for the absence of an obligation to complete set out in the first section. Go into a negotiation without knowing this and you end up in an unproductive argument of the form “if you won’t take it on an ukeoi basis, we can’t trust you”.
The middle ground on offer is the deliverable-completion quasi-mandate. No obligation to complete, but payment arises in exchange for delivery of the agreed deliverable. Note the condition attached: where this form is used, consultation with a specialist on the definition of “completion” is treated as essential. That connects straight back to the core of this article. If you cannot write down what constitutes delivery of the deliverable — in other words, the acceptance criteria — the deliverable-completion type does not function.
Pairing phases with contract types (a way of thinking, not advice)
The following reflects the thinking in the guidelines. It is not a guarantee of the legality or suitability of any individual project.
| Phase | Natural fit | Why |
|---|---|---|
| 1. Assessment | Quasi-mandate (performance-ratio) | “Not feasible” is a legitimate result. Completion is hard to conceive of |
| 2. PoC | Quasi-mandate (performance-ratio or deliverable-completion) | If the validation report can be defined as the deliverable, deliverable-completion is an option |
| 3. Development | Quasi-mandate (deliverable-completion), or partly contract for work | Where specifications can be fixed, e.g. surrounding systems, part of the scope can be carved out as a contract for work |
| 4. Additional learning | Quasi-mandate (performance-ratio) plus an SLA | Continuing services. Service levels set separately in the SLA |
For phase 3, it helps to treat the AI model and the surrounding system as separate design problems. Screens, permissions management, integration with existing systems and report output all have specifiable requirements, so there is room to carve them out as a conventional contract for work. The model stays under a quasi-mandate. Telling prospective vendors that you intend to split the scope this way, at the point where you request the estimate, improves the quality of the proposals you get back.
This section sets out general contract categories and is not legal advice. When you draft and sign an actual contract, have it checked by a lawyer or equivalent specialist. This matters particularly where a Thai entity is the contracting party: governing law, jurisdiction and the treatment under Thai law do not necessarily line up with the Japanese analysis.
Read the estimate as five layers
With the design above in place, you read estimates differently. A common question is what AI development costs. This article gives no price ranges. Most published “market rate” information originates from AI development companies’ own marketing sites, with no disclosed basis for the figures, and using unsourced numbers as decision inputs makes for worse decisions, not better ones.
What we propose instead is a procedure: break the estimate into five layers and match each layer against the line items yourself. You do not need to know a market rate to spot a missing layer.
The five layers
| Layer | Content | Line-item names commonly used |
|---|---|---|
| 1. Problem definition and data survey | Interviewing the business, inventorying the target data, judging feasibility | Requirements definition, assessment, preliminary survey |
| 2. Data preparation | Collection, entity matching, correcting gaps and inconsistent notation, annotation, building training data | Data pre-processing, data cleansing, training data creation |
| 3. Model / application build | Model selection, training and tuning, inference API, screens | Model development, AI implementation, application development |
| 4. Integration with existing systems and permissions design | Connections to production management, accounting, PLC/SCADA and similar; authentication, permissions, audit logs, test environments | System integration, interface development, infrastructure build |
| 5. Operation and retraining | Monitoring, detecting accuracy drift, additional learning, model replacement, support desk | Maintenance and operation, SLA, MLOps |
The “cheap” estimate is usually missing layers 4 and 5
Collect estimates from several vendors and the totals will differ widely. Concluding that the cheapest is the most honest is premature. Structurally, the main driver of the gap is not engineering skill or day rates but scope: whether layer 4 (integration with existing systems and permissions design) and layer 5 (operation and retraining) are inside the quoted scope.
Why do those two drop out so easily? Layer 4 cannot be sized without looking inside your existing systems, so at proposal stage it tends to become “quoted separately”. Layer 5 sits after go-live, so it falls naturally outside a capital-investment discussion. But in a factory, without layer 4 the model never reaches the work, and without layer 5 nobody can fix it when accuracy degrades six months in.
Thought experiment: how much cheaper does a missing layer look?
The following is a thought experiment using assumed values. It is neither real data nor a market rate. It is here so you can substitute your own project’s numbers and check the arithmetic yourself.
Assumptions (all hypothetical):
| Layer | Assumed effort |
|---|---|
| 1. Problem definition and data survey | 10 person-days |
| 2. Data preparation | 25 person-days |
| 3. Model / application build | 30 person-days |
| 4. Integration with existing systems and permissions design | 20 person-days |
| 5. Operation and retraining (first year) | 12 person-days/year (assumed 1 person-day per month) |
The arithmetic:
- Initial build including layers 1 to 4: 10 + 25 + 30 + 20 = 85 person-days
- Quoted effort if layer 4 is missing: 10 + 25 + 30 = 65 person-days
- Apparent discount: (85 − 65) ÷ 85 = 20 ÷ 85 ≈ 0.235 → looks about 23.5% cheaper
- Three-year total effort (layers 1 to 4 plus three years of layer 5): 85 + 12 × 3 = 85 + 36 = 121 person-days
- The 65 person-day estimate as a share of three-year total effort: 65 ÷ 121 ≈ 0.537 → about 53.7%
So under these assumptions, an estimate missing layers 4 and 5 reflects roughly half the effort actually required over three years. When vendor A quotes 65 person-days and vendor B quotes 85, the possibility you have to eliminate first is that B is not expensive at all — that A has simply left layer 4 out.
To repeat: the person-day figures above are assumptions used for illustration. Real values move a great deal depending on the target process, the state of the data, and how many systems you are integrating with. What matters is not the numbers but the exercise: working out, by hand, which of the five layers has no corresponding line in your estimate.
The matching procedure
- Tag every line in the estimate to one of the five layers (if a line fits none of them, ask what it covers)
- Identify the layers to which no line was assigned
- For each empty layer, ask in writing whether it is out of scope, provided free of charge, or simply not anticipated
- For any layer answered with “quoted separately”, agree the assumptions behind a ballpark up front — number of systems to integrate with, expected number of interfaces, expected support hours
- Restate the whole thing as a three-year total for internal approval
If you want to improve the accuracy of layer 4 in particular, mapping out the configuration of your existing systems in advance is effective. General verification steps for choosing a development partner in Thailand are collected in our article on choosing a system development company in Thailand. And since a three-year approval will not pass without a story about benefits, it is worth reading our article on measuring the effect and ROI of AI adoption first.

Designing acceptance criteria — never accept “99% accuracy”
With contract type decided and the estimate layers aligned, acceptance criteria come next. We consider this the section with the greatest practical impact.
When a proposal says “99% accuracy”, the buyer’s first question should not be “really?” It should be “99% of what?”
Change the denominator and you change the checking workload
Take a project that reads business forms with AI. “99% accuracy” has at least three readings.
| Denominator | What “99%” means | How it feels to the buyer |
|---|---|---|
| Per character | 99% of all characters recognised correctly | With hundreds of characters per form, almost every form contains an error |
| Per field | 99% of all extracted fields extracted correctly | The more fields per form, the lower the form-level pass rate |
| Per form | 99% of all forms contain not a single error | The strictest definition. Few vendors will guarantee it |
The same “99%” produces order-of-magnitude differences in how many items a person has to check. And proposals normally do not state the denominator at all.
Thought experiment: how the denominator changes the manual checking workload
The following is a thought experiment using assumed values, not real data. It also simplifies by assuming errors occur independently across fields. Real forms cluster errors in particular fields and particular form types, so this simplification does not match reality.
Assumptions (all hypothetical):
| Assumption | Value |
|---|---|
| Forms processed per month | 1,000 |
| Fields extracted per form | 20 |
| Case A: per-field accuracy | 99% (= 1% error rate per field) |
| Case B: per-field accuracy | 99.9% (= 0.1% error rate per field) |
| Visual check time for one form containing at least one error | 3 minutes |
| Errors assumed independent across fields | — |
Case A (99% per field)
- Total fields per month: 1,000 forms × 20 fields = 20,000 fields
- Expected erroneous fields: 20,000 × 1% = 200 fields
- Probability all 20 fields in a form are correct: 0.99 to the power of 20 ≈ 0.8179 (about 81.8%)
- Share of forms with at least one erroneous field: 1 − 0.8179 = 0.1821 (about 18.2%)
- Forms needing a manual check: 1,000 × 18.2% ≈ 182 forms/month
- Checking effort: 182 forms × 3 minutes = 546 minutes ≈ 9.1 hours/month
Case B (99.9% per field)
- Expected erroneous fields: 20,000 × 0.1% = 20 fields
- Probability all 20 fields in a form are correct: 0.999 to the power of 20 ≈ 0.9802 (about 98.0%)
- Share of forms with at least one erroneous field: 1 − 0.9802 = 0.0198 (about 2.0%)
- Forms needing a manual check: 1,000 × 2.0% ≈ 20 forms/month
- Checking effort: 20 forms × 3 minutes = 60 minutes = 1.0 hour/month
Measured per field, “99%” and “99.9%” are 0.9 percentage points apart. Under these assumptions, though, the forms needing a manual check are 182 versus 20 and the checking effort is 9.1 hours versus 1.0 hour — a gap of roughly nine times. Raise the fields per form from 20 to 40 and the same calculation gives about 331 forms and about 39 forms, so the gap in absolute form counts widens from 162 to 292 (the ratio itself narrows slightly, from about 9.1 times to about 8.4 times). The more fields a form has, the more directly a misread denominator lands on the shop floor as workload.
The implication is clear. Do not let the contract say only “99% accuracy”. An accuracy metric with no stated denominator will always be read two different ways at acceptance.
Write acceptance criteria as business metrics, not “did it read”
The formulation we recommend defines acceptance in business metrics rather than AI performance metrics.
| Poor wording | Better wording |
|---|---|
| Accuracy of 99% or above | Against the acceptance dataset (1,000 records listed in Schedule A, with the breakdown by type stated), no more than X forms require manual correction |
| Few misrecognitions | For designated fields (amount, quantity, part number), forms containing one or more errors held to X% or below |
| Fit for practical use | In monthly closing, operator time on the process reduced by X hours or more compared with before implementation |
Three points. First, fix the acceptance dataset in a schedule. Move the population and the numbers move with it. Second, weight the fields. An error in a part number or an amount has an entirely different business impact from an error in a remarks column. Third, write it in terms of the work that remains. Not “the read rate” but “how many items a human still has to touch” is the buyer’s real metric.
We go deeper into pinning down accuracy denominators, including how individual products approach it, in our AI-OCR comparison. Reading it before you draft acceptance criteria sharpens the questions you ask.
Write down what happens when acceptance fails
The other commonly forgotten piece is what happens if the acceptance criteria are not met. Put the following in the contract or the phase agreement.
- Number of retries and the time allowed (how many rounds of remediation at no charge, over how many weeks)
- What happens if the criteria are still not met after retries (price reduction, termination at that phase, not proceeding to the next phase)
- What is handed over to the buyer at that point (interim deliverables, data, documentation, treatment of the trained model)
- The deletion procedure for data the buyer supplied, and the evidence of deletion
The third is the important one. Even when acceptance is never reached, having the prepared data and the annotation results in your hands gives the next attempt a starting point. Leave it out and you reach the worst outcome: money paid, nothing retained.
IP allocation — who ends up owning the trained model
Alongside acceptance criteria, IP allocation is where disputes cluster. METI’s contract checklist calls for distinguishing foreground IP (results newly created) from background IP (pre-existing assets) and clarifying the position in advance.
Foreground IP and background IP
- Background IP: assets each party held before the contract. On the vendor side, existing libraries, general-purpose models, in-house frameworks. On the buyer side, your operational data, existing systems, know-how.
- Foreground IP: results newly created under this contract.
The tangle happens where the two mix. A trained model produced by additionally training the vendor’s general-purpose model on your shop-floor data — whose asset is it? If the contract has no answer, the project stops dead the moment you want to switch vendors after go-live, or roll the system out to another plant in the group.
Four things to treat separately
When you discuss IP allocation, do not lump everything together as “the AI model”. Split it four ways.
| Object | Issues |
|---|---|
| Raw data (your shop-floor data) | Who owns it, how far the vendor may use it, retention period and deletion |
| Training dataset (after preparation and processing) | The preparation effort may have been borne by the vendor. Set ownership and scope of use separately |
| Trained model (parameters) | Ownership or licence. Exclusive or non-exclusive. Handed over on a vendor switch or not |
| Derived models (results of additional learning, distillation, transfer learning) | The most commonly overlooked. This is where reuse on other clients’ projects is decided |
The checklist states that development-type contracts need to define clearly the permitted scope and specific conditions of the vendor’s use of inputs. Read the other way round: write nothing and the scope stays undefined while the project proceeds.
The one sentence to confirm in the contract
There is one issue every manufacturing buyer should confirm explicitly.
Will the vendor be permitted to reuse the model trained on your shop-floor data — and its derived models — on other clients’ projects? If so, on what conditions?
Many AI development companies build accumulated general know-how into their business model, and there is nothing wrong with that in itself. The problem is the state in which the buyer believes a model was built exclusively for them while, contractually, the vendor may supply the same model to a direct competitor. In ASEAN, where competitors often sit in the same industrial estate, that difference is not something management can ignore.
Drafting options include:
- Prohibit outright: prohibit provision to third parties of the project data and of trained and derived models originating from it
- Permit outside your industry: prohibit provision only to specified industries or to a named list of competitors
- Exclusive for a period: exclusive for X years after go-live, converting to non-exclusive thereafter
- Permit abstracted know-how only: permit reuse of generalised methods and know-how, but not individual data or parameters
Option 1 is the safest but may raise the vendor’s price. Options 3 and 4 tend to be realistic landing points. Whichever you choose, we recommend putting the issue into the RFP. Raise it just before signature and you reopen price and schedule.
Specify the form of handover too
Writing “the model belongs to the buyer” means little if what actually gets handed over is left vague.
- File format and version of the trained model
- Training scripts, pre-processing code, hyperparameter settings
- Documentation sufficient for you or a third party to carry out retraining
- Dependent libraries and their licences (including open-source licence constraints)
- Specification of the runtime environment needed for inference
If future insourcing or transfer to a third party is in view, this level of detail needs to be written down. On how to design the layer that connects your own data, our article on building RAG on factory knowledge is also useful.
Extra decisions when you contract from a Thai base
From here on the issues are specific to companies operating in Thailand and ASEAN. For the same AI project, where you place the contracting entity changes the tax and legal treatment.
Contract from the Japanese parent or from the Thai entity?
This is the entry decision. Contract through the Japanese parent and it is easier to keep everything within Japanese law and Japanese tax, but the plant that actually uses the system is in Thailand. Contract through the Thai entity and you run it on local budget and local approval, but Thai tax rules apply. According to JETRO, Thai corporate income tax has been set at 20% on a permanent basis for accounting periods beginning on or after 1 January 2016, and VAT is currently 7% (the statutory headline rate is 10%, reduced to 7% by royal decree).
Contracting with a vendor inside Thailand
Service fees paid to a service provider inside Thailand attract withholding tax (WHT). For payments by ordinary companies the rate is 3% (5% where the recipient is a non-permanent foreign company branch). Payments to companies are filed on form PND53, filed and paid by the 7th of the month following the month of payment.
There is also a measure worth knowing about: the reduced 1% e-Withholding Tax rate. The Thai cabinet approved an extension of the reduced electronic-withholding rate, so what had been due to end at the close of 2025 continues, with retroactive effect, to the end of 2027. The rate stays at 1%, down from the usual 5%, 3% and 2%, and the period runs 1 January 2026 to 31 December 2027. It applies to payments of assessable income to companies (excluding foundations and associations) and to individuals, made through the e-Withholding Tax system — including service fees, professional fees, rent, advertising fees, royalties, promotional expenses and prize money.
Alongside this, additional corporate income tax deductions were approved for investment in e-Tax Invoice / e-Receipt systems, investment in e-Withholding Tax systems, service provider fees relating to their use, and information system evaluation services through ETDA (same period, 1 January 2026 to 31 December 2027). These are an extension of measures that had lapsed on 31 December 2025.
Thought experiment: how withholding changes the amount remitted
The following is a thought experiment using an assumed amount. It does not represent actual transaction terms or a tax determination. Only the tax rates are published figures drawn from the sources.
Assumption (the amount is hypothetical): a Thai entity pays a service fee of 100,000 THB (excluding tax) to an AI development company inside Thailand.
| Item | Standard WHT at 3% | e-Withholding Tax at 1% |
|---|---|---|
| Service fee (excluding tax) | 100,000 THB | 100,000 THB |
| VAT at 7% | +7,000 THB | +7,000 THB |
| Amount withheld | −3,000 THB (3%) | −1,000 THB (1%) |
| Actual remittance to the vendor | 104,000 THB | 106,000 THB |
The arithmetic: 100,000 + 7,000 − 3,000 = 104,000. And 100,000 + 7,000 − 1,000 = 106,000.
The total payable (service fee plus VAT) is the same; what changes is the split between what is withheld and remitted to the state and what goes directly to the vendor. From the vendor’s side more cash lands in hand, so in some cases e-Withholding Tax capability can be used as a negotiating point. Note that the treatment of VAT (whether and when input tax can be credited) and the refund or offset of withholding tax require case-by-case judgement — confirm with a tax specialist.
Contracting with an overseas vendor — reverse-charge VAT and PND54
If you contract with a Japanese AI development company, or with a vendor in Singapore, Europe or the US, there are additional issues.
Where a company outside Thailand supplies services from abroad and those services are used in Thailand, the supply is deemed to take place in Thailand and VAT at 7% arises. In that case the Thai entity receiving the service must file on behalf of the overseas supplier, using form P.P.36. This is the reverse charge.
Further, payments out of Thailand are described as highly likely to be subject to withholding tax (PND54) at the same time, and taking specialist advice in advance is recommended. The applicable rate and whether a tax treaty applies vary by case, so this article does not assert a figure.
The practical implication: an overseas vendor that looks cheaper may score differently once the filing work and cost arising on the Thai side are included. And these filing obligations are not the kind you can excuse by saying you did not know about them. At the negotiation stage, settle these three points with your own finance and tax team.
- Which entity contracts (Japanese parent or Thai entity)
- Where the “place of use” of the service is determined to be
- Whether VAT and withholding tax are inside the contract amount or borne separately (whether there is a gross-up clause)
The third is a drafting question. A contract that says only “amounts exclude taxes” leaves who bears them to be argued later.
Thailand’s AI Act is not yet enacted — which is exactly why it belongs in the contract
One further issue is time-sensitive. Thailand’s Electronic Transactions Development Agency (ETDA) published a new draft of the Draft Act on Artificial Intelligence on 2 July 2026 and ran a public consultation of about 30 days. That consultation window is understood to have closed in early August 2026. As of writing, the Act has not been enacted.
The Act is not yet enacted, and the timing of enactment is not fixed. So this article does not tell you to “comply with the Thai AI Act”. The draft does, however, contain elements a buyer can usefully act on now.
The draft borrows a risk-based structure from the EU AI Act while carrying Thailand-specific elements: strict liability for AI-related harm, an obligation for foreign providers to appoint a local representative, and a labelling obligation for AI-generated content. The risk tiers comprise prohibited-risk AI (manipulation of cognition and behaviour through subliminal techniques; AI systems producing unfair and widespread discrimination) and high-risk AI (to be designated by royal decree, covering systems affecting national security, health, the environment, energy, telecommunications and transport). Subsequent royal decrees may additionally require notification to, registration with, or licensing by the regulator before certain AI systems are deployed.
One practical conclusion follows. Write into the contract now who handles compliance if and when the Act is enacted. Clauses worth considering:
| Possible requirement | What to put in the contract |
|---|---|
| Local representative obligation for foreign providers | If you contract with an overseas vendor, who arranges the local representative and who bears the cost, should the obligation arise |
| Labelling obligation for AI-generated content | If a labelling function has to be implemented, whether it is treated as additional development or as within the maintenance scope |
| Notification and registration of high-risk AI | Division of work in preparing filings if the system falls in scope, and the obligation to supply the necessary technical documentation |
| Change of risk classification | The duty to consult if amendments to legislation or royal decrees change the classification, and the principle for allocating cost |
Even a single sentence to the effect that “the parties will consult separately if and when the legislation is enacted” is far better than writing nothing. Ideally you also fix the framework for that consultation: who, by when, and on what basis costs are shared. Country-level differences in practice around AI adoption in Thailand are also covered in our article on AI implementation in Thailand.
On both tax and legislation, this article provides general information only. Confirm any actual decision with tax and legal specialists.

Ten questions to always ask an AI development company
Here is the whole discussion reduced to a question list you can use in a meeting. Each comes with why you ask it and what answer should worry you.
1. Which of the three checklist types (general-purpose AI service use / customized / new development) do you think this project falls into?
Why ask: different types raise different issues. The answer tells you how much contracting experience the vendor actually has.
Red flag: an instant “we’ll build everything bespoke”. Parts that an off-the-shelf service would cover may be getting pushed into new development.
2. Can we run this as separate contracts per phase, along the lines of exploratory staged development?
Why ask: this is the test of whether you can create a place to stop.
Red flag: pushing back with “a single contract works out cheaper”. You would be trading the freedom to withdraw for the discount.
3. If the assessment phase concludes the project is not feasible, what gets delivered?
Why ask: it tells you what you hold if things go badly.
Red flag: no clear answer, or “that doesn’t happen to us”.
4. For each phase, is the contract a contract for work or a quasi-mandate? If a quasi-mandate, performance-ratio or deliverable-completion?
Why ask: it forces the presence or absence of an obligation to complete into the open.
Red flag: vague framing like “it’s a quasi-mandate but effectively the same as a contract for work”. Legally they are not remotely the same.
5. What is the denominator of the accuracy figure in your proposal — per character, per field, or per form?
Why ask: as the thought experiment above showed, the denominator changes the required effort by an order of magnitude.
Red flag: no immediate answer, or “generally we get around 99%” with no denominator attached.
6. How will the acceptance dataset be decided? How many records, with what breakdown?
Why ask: acceptance does not work unless the population is fixed.
Red flag: trying to settle it with “we’ll run it and see if there are problems”.
7. If the acceptance criteria are not met, how many rounds of remediation over how many weeks are free of charge? And what happens after that?
Why ask: it checks the exit route when things go wrong.
Red flag: “we’ll keep going until it’s met”, with no time limit. Unlimited liability is never actually performed.
8. Will the trained model and derived models built on our data be reused on other clients’ projects?
Why ask: this is the core IP issue.
Red flag: stopping at “we may use it as generalised know-how” without going on to define the scope.
9. Is integration with our existing systems and permissions design included in this estimate? If not, what has to be fixed before you can quote it?
Why ask: it closes the gap where layer 4 goes missing.
Red flag: leaving integration out of scope with “your side handles that” and giving no conditions.
10. After go-live, who does the retraining, at what frequency, and with what team? And which entity is the contracting party — the Thai company or the Japanese one?
Why ask: it checks layer 5 and the entry point to the tax and legal issues at the same time.
Red flag: a sudden change of personnel when operations come up, or brand-new information such as “maintenance is a different company”.
Ask these ten in writing and get the answers in writing wherever you can. Verbal explanations do not survive into the contract. If the vendor attaches the answer document to the proposal, it becomes a shared reference point at acceptance.
Six common failure patterns
Here are the failures we see most often, in cause-and-countermeasure form.
Pattern 1: the project stalls at PoC and quietly goes dormant
Cause: the PoC exit condition was defined at the level of “try it and get a feel for it”. With no criterion for progressing to the next phase, nobody can make the call.
Countermeasure: agree in writing, before the PoC starts, which numbers have to reach what level to move into the development phase. Write the termination procedure for the other case at the same time.
Pattern 2: “accuracy” is read two different ways at acceptance
Cause: the contract contains only an accuracy metric with no denominator.
Countermeasure: fix the acceptance dataset in a schedule and write the criteria in business metrics (the count and hours of manual work that remain).
Pattern 3: the model works but never connects to the business systems
Cause: layer 4 (integration with existing systems and permissions design) was out of scope. After go-live, someone is still shuttling CSV files by hand.
Countermeasure: run the five-layer matching on the estimate and close the layer 4 gap before you place the order. Build the list and count of systems to integrate with yourself, in advance.
Pattern 4: accuracy drops six months in and nobody can fix it
Cause: layer 5 (operation and retraining) was not contracted for. The engineers who built it have moved to another project and no retraining procedure was left behind.
Countermeasure: include the additional learning phase in the contract from the start. Make handover of the scripts and documentation needed for retraining part of the acceptance criteria.
Pattern 5: you cannot switch vendors
Cause: ownership of the trained model is ambiguous, or the runtime environment is vendor-specific and the form of handover was never defined.
Countermeasure: put the handover list from the IP section into the contract. State background IP and foreground IP separately.
Pattern 6: a tax problem surfaces just before go-live
Cause: reverse-charge VAT and PND54 on payments to an overseas vendor come to light after signature. With no gross-up clause, the parties argue over who bears them.
Countermeasure: put the choice of contracting entity at the very front of the negotiation. Involve finance and tax from the estimate-comparison stage onwards.
What the six have in common is that none of them is a technical problem; all are problems of what was agreed in advance. Which is precisely why comparing lists of vendors does not prevent them.
A 90-day plan
Finally, a roadmap for execution. The following is an example of a standard sequence; it moves either way depending on project size and the state of your data.
| Week | What to do | Owner | Output |
|---|---|---|---|
| Weeks 1–2 | Narrow the target process. List three processes that hurt, then pick one | Business function + IT | Target process decided |
| Week 3 | Run the 28 METI checklist items against your project and identify the ones you cannot answer | IT + admin functions | List of open items |
| Week 4 | Determine which of the three types applies. Test first whether an off-the-shelf service is insufficient | IT | Type classification memo |
| Weeks 5–6 | Inventory the data. Establish in-house where the target data sits, how many records, over what period, with what gaps | Shop floor + IT | Data status sheet |
| Week 7 | Decide the contracting entity (Japanese parent / Thai entity). Bring finance and tax in | Admin functions | Contracting entity decided |
| Week 8 | Write the RFP. State the staged-contract premise, your IP preferences and your approach to acceptance | IT + admin functions | RFP issued |
| Weeks 9–10 | Receive proposals from several vendors. Put the ten questions in writing and get written answers | IT | Proposals + answer documents |
| Week 11 | Match estimates across the five layers. Query empty layers in writing. Restate as a three-year total | IT + admin functions | Layer-by-layer comparison |
| Week 12 | Review the draft contract. Have lawyers and tax specialists check it. Draft the acceptance criteria schedule | Admin functions + external specialists | Contract draft finalised |
| Week 13 | Sign the assessment-phase contract. Final check of entry and exit conditions | Management | Signature and kick-off |
Notice that within these 90 days, you are actually talking to AI development companies for two to three weeks. The rest is time spent deciding, internally, what you have to decide. Reverse the order and start with vendor visits, and you will have the vendor decide your open items for you — and those decisions will always land in the vendor’s favour.
If the data inventory in weeks 5–6 slips, everything slips. It is the hardest piece to size, so start it with time in hand. The specialist review in week 12 also depends on external availability, so make the approach early.
FAQ
What is the difference between an AI development company and an AI consulting firm?
There is no agreed industry definition, but in practice the emphasis differs. Firms that describe themselves as AI consulting tend to be strong at layer 1 — problem definition and data survey — judging which processes AI should be applied to and organising the investment case. AI development companies centre on layer 3, the model and application build. The problem is that layers 2 (data preparation), 4 (integration with existing systems) and 5 (operation) can fall into neither party’s remit and float. If you engage both, put on paper at the outset who holds which of the five layers. And if you contract consulting separately, check whether its deliverables come in a form the development company can actually pick up.
Why do AI development outsourcing quotes vary so much between companies?
The main driver is not day rates but difference in scope. As the estimate section showed, whether layer 4 (integration and permissions design) and layer 5 (operation and retraining) are included moves the total a great deal. A further large factor is different assumptions about layer 2. A vendor quoting on the premise that “the data is already prepared” and one quoting on the premise of “starting from correcting inconsistent notation” will produce completely different effort figures for the same project. When you compare, do not line up the totals — tag the lines to the five layers and level the scope first. Once the scope is levelled, the price gap is usually smaller than you expected.
Is it better to outsource custom AI development or hire AI developers in house?
Treating it as a binary is the wrong framing. A realistic split across the five layers keeps layer 1 (problem definition) and layer 5 (operation) close to your own organisation and sends the technically heavy parts of layers 2 to 4 outside. Two reasons. First, layer 1 does not get more accurate without your own understanding of the business. Second, if you outsource layer 5 entirely, detection of accuracy drift is delayed and you can do nothing the moment the contract lapses. Conversely, trying to build layer 3 in house from the start takes too long in recruiting and retaining people. Our article on AI insourcing support sets out the layer-by-layer division of responsibility in detail.
Can an AI system development contract be a contract for work?
It is not that it cannot be done, but if you ask for a contract for work covering the AI model itself, the general pattern is that the vendor declines or prices in a risk premium. Because AI accuracy depends on training data, vendors are said to tend to avoid contracts for work. Two designs are realistic. One is to separate the AI model from the surrounding system and carve only the surrounding system, where specifications can be fixed, into a contract for work. The other is the middle-ground deliverable-completion quasi-mandate. If you take the latter, note that consultation with a specialist on the definition of “completion” is treated as essential; if you cannot write that definition down, the deliverable-completion type is a formality only. Always confirm the final choice of contract type with a lawyer.
Which part of AI development cost varies the most?
Layers 2 and 4 — data preparation and integration with existing systems. Layer 2 varies because the true state of the data is unknown until someone starts work. Being told “we have three years of data” and then finding that the format changed midway, or that each operator followed different input rules, is a common story. Layer 4 grows more than linearly with the number of systems to integrate with and the specification of the interfaces. The best way to contain the variance in both is to make the assessment phase a standalone contract, establish the actual state of the data and the existing systems there, and then re-quote layers 3 onwards. Staged contracting does not only create the freedom to withdraw; it is also a device for improving estimate accuracy.
At what stage can we start an AI consulting conversation?
The concept stage is fine. In fact, waiting until the RFP is written can lock in assumptions about layers 2 and 4 that do not match reality. That said, three things prepared in advance make the conversation move much faster. First, candidate target processes (three listed, narrowed to one). Second, the current state of the data — where it sits, how many records, over what period, with what gaps. Third, a hypothesis about how you will measure the benefit. On the third, our article on measuring the effect and ROI of AI adoption is a useful starting point. With those three in hand, a first meeting can get as far as an initial feasibility view.
Summary
To restate the argument.
AI development proceeds by an inductive method grounded in training data, so the limits of accuracy may already be contained in that data, and setting an obligation to complete or a performance guarantee is difficult. The assumption that picking a good company means you get a finished system is structurally hard to sustain. Which is why, before comparing lists of vendors, you need to design four things.
- Contract type: contract for work or quasi-mandate. If a quasi-mandate, performance-ratio or deliverable-completion. Whether you can write down the definition of “completion”
- How to split the phases: four phases — assessment, PoC, development, additional learning — with entry and exit conditions on each. Create a place where you can stop
- Acceptance criteria: state the denominator of any accuracy figure, fix the acceptance dataset in a schedule, and write the criteria in business metrics (the count and hours of manual work that remain)
- IP allocation: distinguish foreground IP from background IP, and set terms separately for raw data, the training dataset, the trained model and derived models. Always confirm whether reuse on other clients’ projects is permitted
If you contract from Thailand, add to that the choice of contracting entity, the treatment of WHT, VAT and the reverse charge, and the allocation of work should the not-yet-enacted Thai AI Act come into force. Always confirm tax and legal points with specialists.
And as the 90-day roadmap shows, most of this design work can be finished inside your own organisation before you meet a single AI development company. Simply respecting that order prevents a large share of selection failures.
TOMAS TECH is based in Bangkok and builds IT and OT systems for the factories of Japanese manufacturers across Thailand and ASEAN. We are glad to talk at an early stage — “we are still at the concept stage and have not decided which process to target”, or “we would like help reading the estimates other vendors have sent us”. Use us as a sounding board before you choose a supplier. Get in touch through our contact form.
References
- METI, Contract Guidelines on the Use of AI and Data (published June 2018, partially revised December 2019), AI volume — https://www.meti.go.jp/policy/mono_info_service/connected_industries/sharing_and_utilization/20180615001-3.pdf
- METI, Checklist on contracts for the use and development of AI (published 18 February 2025), source PDF — https://www.meti.go.jp/policy/mono_info_service/connected_industries/sharing_and_utilization/20250218003-ar.pdf
- The same, press release — https://www.meti.go.jp/press/2024/02/20250218003/20250218003.html
- Overview material on the contract guidelines (January 2021, METI Information Economy Division) — https://www.jftc.go.jp/cprc/conference/index_files/21011902.pdf
- Logit Partners Law and Accounting Office, commentary on METI’s checklist on contracts for the use and development of AI — https://partners.logit.jp/ai-checkliist-meti/
- Japan Patent Office, IP BASE, “Contract types in AI development” — https://ipbase.go.jp/learn/point/ai/page05.php
- JETRO, “Thailand: Tax System” (corporate income tax 20%, VAT 7%, withholding tax) — https://www.jetro.go.jp/world/asia/th/invest_04.html
- HLB Thailand, “Cabinet Approves Extension of Tax Measures to End of 2027” (e-Withholding Tax reduced rate of 1%, 1 January 2026 to 31 December 2027) — https://www.hlbthai.com/cabinet-approves-extension-of-tax-measures-to-end-of-2027-to-promote-adoption-of-electronic-tax-systems/
- Mahanakorn Partners, “Thailand Approves Two-Year Extension of Electronic Tax System Incentives” — https://mahanakornpartners.com/thailand-approves-two-year-extension-of-electronic-tax-system-incentives/
- JGA, “Payments to overseas suppliers and reverse-charge VAT (P.P.36)” — https://jga.asia/?p=1241
- Mondaq, “Thailand Releases New Draft Artificial Intelligence Act” (ETDA draft published 2 July 2026; not yet enacted) — https://www.mondaq.com/it-and-internet/1811290/thailand-releases-new-draft-artificial-intelligence-act
- Tilleke & Gibbins, “Thailand’s AI Governance Framework” — https://www.tilleke.com/insights/comprehensive-policy-thailands-ai-governance-framework/