Blog

2026.08.05

How to Choose an AI Development Company in 2026 — Contracts, Scope and Cost

How to Choose an AI Development Company in 2026 — Contracts, Scope and Cost

Search “best AI development company” and you get listicles — 25 vendors compared, 50 vendors compared. Yet shortlisted projects still stall at PoC, break down at acceptance, or leave the model owned by the vendor. Success is decided by contract type, staging, acceptance criteria and IP allocation — not by the comparison table. Here is the order to decide them in.

Why AI work is hard to promise as “finished”

Comparison articles rest on an unstated assumption: pick a good company and what you ordered will get built. For conventional system development that assumption mostly holds. Build the screens and functions written in the specification and the job is done. Once an AI model is in scope, the assumption breaks down structurally.

Japan’s Ministry of Economy, Trade and Industry (METI) published a checklist on contracts for the use and development of AI on 18 February 2025. On performance and accuracy it makes a specific point: because AI models are developed by an inductive method grounded in training data, the limits of accuracy may already be baked into the training data itself. The consequence it draws is that setting an obligation to complete, or a performance guarantee, is difficult.

This is not a statement about vendor competence. It is a statement about information: no engineer, however good, can extract more than the data contains. The shop floor makes it concrete. If three years of inspection records say only “defect” — with no defect type, no process step, no operator, no ambient conditions — then no model built on that data will predict which conditions produce defects. And you usually only discover that after someone has spent a few weeks with the data.

So the first job for the buyer is not “find a good company”. It is to decide, on the assumption that the work may not complete, how much you are willing to pay and where you have the right to stop.

A note for readers working outside Japan. METI’s documents are Japanese public guidance, and they are worth reading wherever you sit, because they are the most detailed publicly available treatment of exactly this problem, and because for many Japanese manufacturers in ASEAN the counterparty, the group standard, or the parent-company approval process is Japanese. Use them as a checklist of issues rather than as applicable law — which law applies follows from the contracting entity, covered later in this article.

First, classify your own project

The same checklist splits AI-related contracts into three use-case types. Which one you are in changes the issues you need to close out, and even whether a bespoke contract is needed at all.

TypeWhat it isBuyer’s main concerns
General-purpose AI service useUsing an off-the-shelf AI service as-isHow input data is handled, retention period, deletion, security posture
Customized typeUsing a service tuned for the userThe above, plus scope of tuning, ownership of tuning results, treatment on termination
New development typeCo-developing a bespoke AI systemThe above, plus whether an obligation to complete exists, acceptance criteria, ownership of the trained model

The checklist positions itself as a way of “showing where the issues lie” when contracting for AI use, and is organised as 14 input-related items and 14 output-related items. The input side covers definition of the data the user supplies to the vendor, purpose of use, conditions of use, management and security arrangements, retention and deletion. The output side covers obligation to complete, conditions of provision, purpose of use, conditions of use and security management.

The buyer’s first piece of work here is to run those 28 items against your own project and identify the items you do not yet have an answer for. Doing this before you talk to any AI development company changes the quality of the conversation entirely, because questions like “do you anticipate using our data on other clients’ projects?” and “when and by what procedure is it deleted?” then come from a systematic list rather than from whatever occurs to you in the room.

Classifying into one of the three types also feeds directly into the build-versus-buy line — how much you keep in house and where outsourcing starts. We set out a layer-by-layer view of who owns what in our article on AI insourcing support.

Exploratory staged development — splitting the contract into four phases

If completion is hard to promise, how should you contract? Japan’s official answer sits in METI’s Contract Guidelines on the Use of AI and Data, published in June 2018 and partially revised in December 2019. The AI volume of those guidelines proposes a development method known as exploratory staged development.

The idea is simple. Do not contract for the whole thing up front. Break the development process into phases, contract phase by phase, and decide at the end of each one whether to continue or stop. The guidelines set out four phases: assessment, PoC, development, and additional learning. In the table below the phase names come from the guidelines; the “what happens” and “what the buyer holds at the end” columns are our own summary of common practice.

PhaseWhat happensWhat the buyer holds at the end
1. AssessmentUse a limited sample of data to judge whether a model is feasible at allA technical feasibility report. “Not feasible” is a legitimate result
2. PoCBuild a prototype on real data and test whether accuracy is good enough for the jobA prototype and test results. An outlook expressed in business metrics
3. DevelopmentBuild the production model and surrounding system on the validated approachTrained model, surrounding systems, documentation
4. Additional learningAdd data during operation to maintain and retrain the modelSustained accuracy, ability to follow changes in the business

The guidelines set out contracting method, factors to consider, and sample clauses for each phase. They also note that, given that development of an AI model on its own proceeds by an inductive method dependent on data, a quasi-mandate contract (junin’nin) is a natural fit for that work.

The real point of phasing is creating a place to stop

Phasing sounds like a delivery-method topic. For the buyer, the practical value lies elsewhere. The point is creating a point at which you can legitimately terminate.

Run a project under a single lump-sum contract and, once it becomes apparent that the thing is harder than expected, you cannot stop — the contract is still running. Stop and you are into a breach discussion; continue and you keep funding an investment with no prospect. That is the anatomy of “stuck at PoC, and somehow still spending”. Under phased contracts, if the assessment phase concludes that the project is not feasible with the data as it stands today, you pay for that phase and end it. The value is in being able to end it.

One point deserves emphasis. When the assessment phase concludes “not feasible”, that is not a failure. In the sense that you avoided a full investment decision for the cost of the investigation alone, it is a success. If your organisation creates an internal climate in which PoCs must succeed, the teams on the ground will keep hopeless projects alive. If you introduce phased contracting, design the internal evaluation criteria at the same time as the contract.

Write entry and exit conditions for each phase

Splitting the work into phases is not enough on its own. For each phase, write entry conditions (when you may start this phase) and exit conditions (when this phase can be said to be finished) into the contract or into a phase-specific agreement.

For a PoC phase, an entry condition might read: “the assessment report concluded that the target data contains information sufficient for discrimination”. An exit condition might read: “the business metric measured against the validation dataset meets the level set out in Schedule X, or, where it does not, a root-cause analysis report has been delivered”. The trick is to include a deliverable for the case where the target is missed. Without it, a phase that goes badly leaves you with nothing at all.

How to Choose an AI Development Company in 2026 — Contracts, Scope and Cost - figure 1

Contract for work or quasi-mandate — and the “deliverable-completion quasi-mandate” middle ground

Once the phases are set, the next question is which contract type each phase uses. Two Japanese civil-law categories dominate the discussion, and IP BASE, run by the Japan Patent Office, frames them as follows.

TypeWhat the vendor owesBuyer’s main remedy
Contract for work (ukeoi)Promises completion of the work and bears legal responsibility for performing itWhere the deliverable is non-conforming, can demand cure or damages
Quasi-mandate (junin’nin)Performs to a given standard with the care of a good manager. Owes no obligation to completePursuing breach of the duty of care. Cannot demand completion itself

The difference is large, and leaving it vague is a reliable source of trouble. The buyer thinks “we ordered it, so it will be finished”. The vendor thinks “we explained that accuracy depends on the data”. The gap surfaces for the first time at acceptance.

Two variants of the quasi-mandate

The quasi-mandate splits further into two.

  • Performance-ratio type: the obligation to pay arises when the delegated work has been performed
  • Deliverable-completion type: the obligation to pay arises in exchange for delivery of the agreed deliverable

Performance-ratio type is, in plain terms, “pay for the work done”. It suits exploratory work like the assessment and PoC phases, where nobody knows in advance what will come out. From the buyer’s side, though, it means paying even if nothing comes out, which gets harder to justify internally as the amount grows.

Deliverable-completion type means the payment obligation arises only once a pre-defined deliverable has been handed over. That is easier to explain internally — “nothing delivered, nothing paid” — but it still does not mean the vendor owes an obligation to complete. So the deciding factor for whether you can use this type is whether you can define in writing what counts as the “deliverable”. If you can, it favours the buyer. If you pick deliverable-completion type while the definition stays vague, the dispute over whether delivery occurred is worse than anything performance-ratio type would have produced.

Assume the vendor will try to avoid a contract for work

Because AI accuracy depends on training data, vendors tend to avoid contracts for work. That is not bad faith; it reflects the structural reason for the absence of an obligation to complete set out in the first section. Go into a negotiation without knowing this and you end up in an unproductive argument of the form “if you won’t take it on an ukeoi basis, we can’t trust you”.

The middle ground on offer is the deliverable-completion quasi-mandate. No obligation to complete, but payment arises in exchange for delivery of the agreed deliverable. Note the condition attached: where this form is used, consultation with a specialist on the definition of “completion” is treated as essential. That connects straight back to the core of this article. If you cannot write down what constitutes delivery of the deliverable — in other words, the acceptance criteria — the deliverable-completion type does not function.

Pairing phases with contract types (a way of thinking, not advice)

The following reflects the thinking in the guidelines. It is not a guarantee of the legality or suitability of any individual project.

PhaseNatural fitWhy
1. AssessmentQuasi-mandate (performance-ratio)“Not feasible” is a legitimate result. Completion is hard to conceive of
2. PoCQuasi-mandate (performance-ratio or deliverable-completion)If the validation report can be defined as the deliverable, deliverable-completion is an option
3. DevelopmentQuasi-mandate (deliverable-completion), or partly contract for workWhere specifications can be fixed, e.g. surrounding systems, part of the scope can be carved out as a contract for work
4. Additional learningQuasi-mandate (performance-ratio) plus an SLAContinuing services. Service levels set separately in the SLA

For phase 3, it helps to treat the AI model and the surrounding system as separate design problems. Screens, permissions management, integration with existing systems and report output all have specifiable requirements, so there is room to carve them out as a conventional contract for work. The model stays under a quasi-mandate. Telling prospective vendors that you intend to split the scope this way, at the point where you request the estimate, improves the quality of the proposals you get back.

This section sets out general contract categories and is not legal advice. When you draft and sign an actual contract, have it checked by a lawyer or equivalent specialist. This matters particularly where a Thai entity is the contracting party: governing law, jurisdiction and the treatment under Thai law do not necessarily line up with the Japanese analysis.

Read the estimate as five layers

With the design above in place, you read estimates differently. A common question is what AI development costs. This article gives no price ranges. Most published “market rate” information originates from AI development companies’ own marketing sites, with no disclosed basis for the figures, and using unsourced numbers as decision inputs makes for worse decisions, not better ones.

What we propose instead is a procedure: break the estimate into five layers and match each layer against the line items yourself. You do not need to know a market rate to spot a missing layer.

The five layers

LayerContentLine-item names commonly used
1. Problem definition and data surveyInterviewing the business, inventorying the target data, judging feasibilityRequirements definition, assessment, preliminary survey
2. Data preparationCollection, entity matching, correcting gaps and inconsistent notation, annotation, building training dataData pre-processing, data cleansing, training data creation
3. Model / application buildModel selection, training and tuning, inference API, screensModel development, AI implementation, application development
4. Integration with existing systems and permissions designConnections to production management, accounting, PLC/SCADA and similar; authentication, permissions, audit logs, test environmentsSystem integration, interface development, infrastructure build
5. Operation and retrainingMonitoring, detecting accuracy drift, additional learning, model replacement, support deskMaintenance and operation, SLA, MLOps

The “cheap” estimate is usually missing layers 4 and 5

Collect estimates from several vendors and the totals will differ widely. Concluding that the cheapest is the most honest is premature. Structurally, the main driver of the gap is not engineering skill or day rates but scope: whether layer 4 (integration with existing systems and permissions design) and layer 5 (operation and retraining) are inside the quoted scope.

Why do those two drop out so easily? Layer 4 cannot be sized without looking inside your existing systems, so at proposal stage it tends to become “quoted separately”. Layer 5 sits after go-live, so it falls naturally outside a capital-investment discussion. But in a factory, without layer 4 the model never reaches the work, and without layer 5 nobody can fix it when accuracy degrades six months in.

Thought experiment: how much cheaper does a missing layer look?

The following is a thought experiment using assumed values. It is neither real data nor a market rate. It is here so you can substitute your own project’s numbers and check the arithmetic yourself.

Assumptions (all hypothetical):

LayerAssumed effort
1. Problem definition and data survey10 person-days
2. Data preparation25 person-days
3. Model / application build30 person-days
4. Integration with existing systems and permissions design20 person-days
5. Operation and retraining (first year)12 person-days/year (assumed 1 person-day per month)

The arithmetic:

  • Initial build including layers 1 to 4: 10 + 25 + 30 + 20 = 85 person-days
  • Quoted effort if layer 4 is missing: 10 + 25 + 30 = 65 person-days
  • Apparent discount: (85 − 65) ÷ 85 = 20 ÷ 85 ≈ 0.235 → looks about 23.5% cheaper
  • Three-year total effort (layers 1 to 4 plus three years of layer 5): 85 + 12 × 3 = 85 + 36 = 121 person-days
  • The 65 person-day estimate as a share of three-year total effort: 65 ÷ 121 ≈ 0.537 → about 53.7%

So under these assumptions, an estimate missing layers 4 and 5 reflects roughly half the effort actually required over three years. When vendor A quotes 65 person-days and vendor B quotes 85, the possibility you have to eliminate first is that B is not expensive at all — that A has simply left layer 4 out.

To repeat: the person-day figures above are assumptions used for illustration. Real values move a great deal depending on the target process, the state of the data, and how many systems you are integrating with. What matters is not the numbers but the exercise: working out, by hand, which of the five layers has no corresponding line in your estimate.

The matching procedure

  1. Tag every line in the estimate to one of the five layers (if a line fits none of them, ask what it covers)
  2. Identify the layers to which no line was assigned
  3. For each empty layer, ask in writing whether it is out of scope, provided free of charge, or simply not anticipated
  4. For any layer answered with “quoted separately”, agree the assumptions behind a ballpark up front — number of systems to integrate with, expected number of interfaces, expected support hours
  5. Restate the whole thing as a three-year total for internal approval

If you want to improve the accuracy of layer 4 in particular, mapping out the configuration of your existing systems in advance is effective. General verification steps for choosing a development partner in Thailand are collected in our article on choosing a system development company in Thailand. And since a three-year approval will not pass without a story about benefits, it is worth reading our article on measuring the effect and ROI of AI adoption first.

How to Choose an AI Development Company in 2026 — Contracts, Scope and Cost - figure 2

Designing acceptance criteria — never accept “99% accuracy”

With contract type decided and the estimate layers aligned, acceptance criteria come next. We consider this the section with the greatest practical impact.

When a proposal says “99% accuracy”, the buyer’s first question should not be “really?” It should be “99% of what?”

Change the denominator and you change the checking workload

Take a project that reads business forms with AI. “99% accuracy” has at least three readings.

DenominatorWhat “99%” meansHow it feels to the buyer
Per character99% of all characters recognised correctlyWith hundreds of characters per form, almost every form contains an error
Per field99% of all extracted fields extracted correctlyThe more fields per form, the lower the form-level pass rate
Per form99% of all forms contain not a single errorThe strictest definition. Few vendors will guarantee it

The same “99%” produces order-of-magnitude differences in how many items a person has to check. And proposals normally do not state the denominator at all.

Thought experiment: how the denominator changes the manual checking workload

The following is a thought experiment using assumed values, not real data. It also simplifies by assuming errors occur independently across fields. Real forms cluster errors in particular fields and particular form types, so this simplification does not match reality.

Assumptions (all hypothetical):

AssumptionValue
Forms processed per month1,000
Fields extracted per form20
Case A: per-field accuracy99% (= 1% error rate per field)
Case B: per-field accuracy99.9% (= 0.1% error rate per field)
Visual check time for one form containing at least one error3 minutes
Errors assumed independent across fields

Case A (99% per field)

  • Total fields per month: 1,000 forms × 20 fields = 20,000 fields
  • Expected erroneous fields: 20,000 × 1% = 200 fields
  • Probability all 20 fields in a form are correct: 0.99 to the power of 20 ≈ 0.8179 (about 81.8%)
  • Share of forms with at least one erroneous field: 1 − 0.8179 = 0.1821 (about 18.2%)
  • Forms needing a manual check: 1,000 × 18.2% ≈ 182 forms/month
  • Checking effort: 182 forms × 3 minutes = 546 minutes ≈ 9.1 hours/month

Case B (99.9% per field)

  • Expected erroneous fields: 20,000 × 0.1% = 20 fields
  • Probability all 20 fields in a form are correct: 0.999 to the power of 20 ≈ 0.9802 (about 98.0%)
  • Share of forms with at least one erroneous field: 1 − 0.9802 = 0.0198 (about 2.0%)
  • Forms needing a manual check: 1,000 × 2.0% ≈ 20 forms/month
  • Checking effort: 20 forms × 3 minutes = 60 minutes = 1.0 hour/month

Measured per field, “99%” and “99.9%” are 0.9 percentage points apart. Under these assumptions, though, the forms needing a manual check are 182 versus 20 and the checking effort is 9.1 hours versus 1.0 hour — a gap of roughly nine times. Raise the fields per form from 20 to 40 and the same calculation gives about 331 forms and about 39 forms, so the gap in absolute form counts widens from 162 to 292 (the ratio itself narrows slightly, from about 9.1 times to about 8.4 times). The more fields a form has, the more directly a misread denominator lands on the shop floor as workload.

The implication is clear. Do not let the contract say only “99% accuracy”. An accuracy metric with no stated denominator will always be read two different ways at acceptance.

Write acceptance criteria as business metrics, not “did it read”

The formulation we recommend defines acceptance in business metrics rather than AI performance metrics.

Poor wordingBetter wording
Accuracy of 99% or aboveAgainst the acceptance dataset (1,000 records listed in Schedule A, with the breakdown by type stated), no more than X forms require manual correction
Few misrecognitionsFor designated fields (amount, quantity, part number), forms containing one or more errors held to X% or below
Fit for practical useIn monthly closing, operator time on the process reduced by X hours or more compared with before implementation

Three points. First, fix the acceptance dataset in a schedule. Move the population and the numbers move with it. Second, weight the fields. An error in a part number or an amount has an entirely different business impact from an error in a remarks column. Third, write it in terms of the work that remains. Not “the read rate” but “how many items a human still has to touch” is the buyer’s real metric.

We go deeper into pinning down accuracy denominators, including how individual products approach it, in our AI-OCR comparison. Reading it before you draft acceptance criteria sharpens the questions you ask.

Write down what happens when acceptance fails

The other commonly forgotten piece is what happens if the acceptance criteria are not met. Put the following in the contract or the phase agreement.

  • Number of retries and the time allowed (how many rounds of remediation at no charge, over how many weeks)
  • What happens if the criteria are still not met after retries (price reduction, termination at that phase, not proceeding to the next phase)
  • What is handed over to the buyer at that point (interim deliverables, data, documentation, treatment of the trained model)
  • The deletion procedure for data the buyer supplied, and the evidence of deletion

The third is the important one. Even when acceptance is never reached, having the prepared data and the annotation results in your hands gives the next attempt a starting point. Leave it out and you reach the worst outcome: money paid, nothing retained.

IP allocation — who ends up owning the trained model

Alongside acceptance criteria, IP allocation is where disputes cluster. METI’s contract checklist calls for distinguishing foreground IP (results newly created) from background IP (pre-existing assets) and clarifying the position in advance.

Foreground IP and background IP

  • Background IP: assets each party held before the contract. On the vendor side, existing libraries, general-purpose models, in-house frameworks. On the buyer side, your operational data, existing systems, know-how.
  • Foreground IP: results newly created under this contract.

The tangle happens where the two mix. A trained model produced by additionally training the vendor’s general-purpose model on your shop-floor data — whose asset is it? If the contract has no answer, the project stops dead the moment you want to switch vendors after go-live, or roll the system out to another plant in the group.

Four things to treat separately

When you discuss IP allocation, do not lump everything together as “the AI model”. Split it four ways.

ObjectIssues
Raw data (your shop-floor data)Who owns it, how far the vendor may use it, retention period and deletion
Training dataset (after preparation and processing)The preparation effort may have been borne by the vendor. Set ownership and scope of use separately
Trained model (parameters)Ownership or licence. Exclusive or non-exclusive. Handed over on a vendor switch or not
Derived models (results of additional learning, distillation, transfer learning)The most commonly overlooked. This is where reuse on other clients’ projects is decided

The checklist states that development-type contracts need to define clearly the permitted scope and specific conditions of the vendor’s use of inputs. Read the other way round: write nothing and the scope stays undefined while the project proceeds.

The one sentence to confirm in the contract

There is one issue every manufacturing buyer should confirm explicitly.

Will the vendor be permitted to reuse the model trained on your shop-floor data — and its derived models — on other clients’ projects? If so, on what conditions?

Many AI development companies build accumulated general know-how into their business model, and there is nothing wrong with that in itself. The problem is the state in which the buyer believes a model was built exclusively for them while, contractually, the vendor may supply the same model to a direct competitor. In ASEAN, where competitors often sit in the same industrial estate, that difference is not something management can ignore.

Drafting options include:

  1. Prohibit outright: prohibit provision to third parties of the project data and of trained and derived models originating from it
  2. Permit outside your industry: prohibit provision only to specified industries or to a named list of competitors
  3. Exclusive for a period: exclusive for X years after go-live, converting to non-exclusive thereafter
  4. Permit abstracted know-how only: permit reuse of generalised methods and know-how, but not individual data or parameters

Option 1 is the safest but may raise the vendor’s price. Options 3 and 4 tend to be realistic landing points. Whichever you choose, we recommend putting the issue into the RFP. Raise it just before signature and you reopen price and schedule.

Specify the form of handover too

Writing “the model belongs to the buyer” means little if what actually gets handed over is left vague.

  • File format and version of the trained model
  • Training scripts, pre-processing code, hyperparameter settings
  • Documentation sufficient for you or a third party to carry out retraining
  • Dependent libraries and their licences (including open-source licence constraints)
  • Specification of the runtime environment needed for inference

If future insourcing or transfer to a third party is in view, this level of detail needs to be written down. On how to design the layer that connects your own data, our article on building RAG on factory knowledge is also useful.

Extra decisions when you contract from a Thai base

From here on the issues are specific to companies operating in Thailand and ASEAN. For the same AI project, where you place the contracting entity changes the tax and legal treatment.

Contract from the Japanese parent or from the Thai entity?

This is the entry decision. Contract through the Japanese parent and it is easier to keep everything within Japanese law and Japanese tax, but the plant that actually uses the system is in Thailand. Contract through the Thai entity and you run it on local budget and local approval, but Thai tax rules apply. According to JETRO, Thai corporate income tax has been set at 20% on a permanent basis for accounting periods beginning on or after 1 January 2016, and VAT is currently 7% (the statutory headline rate is 10%, reduced to 7% by royal decree).

Contracting with a vendor inside Thailand

Service fees paid to a service provider inside Thailand attract withholding tax (WHT). For payments by ordinary companies the rate is 3% (5% where the recipient is a non-permanent foreign company branch). Payments to companies are filed on form PND53, filed and paid by the 7th of the month following the month of payment.

There is also a measure worth knowing about: the reduced 1% e-Withholding Tax rate. The Thai cabinet approved an extension of the reduced electronic-withholding rate, so what had been due to end at the close of 2025 continues, with retroactive effect, to the end of 2027. The rate stays at 1%, down from the usual 5%, 3% and 2%, and the period runs 1 January 2026 to 31 December 2027. It applies to payments of assessable income to companies (excluding foundations and associations) and to individuals, made through the e-Withholding Tax system — including service fees, professional fees, rent, advertising fees, royalties, promotional expenses and prize money.

Alongside this, additional corporate income tax deductions were approved for investment in e-Tax Invoice / e-Receipt systems, investment in e-Withholding Tax systems, service provider fees relating to their use, and information system evaluation services through ETDA (same period, 1 January 2026 to 31 December 2027). These are an extension of measures that had lapsed on 31 December 2025.

Thought experiment: how withholding changes the amount remitted

The following is a thought experiment using an assumed amount. It does not represent actual transaction terms or a tax determination. Only the tax rates are published figures drawn from the sources.

Assumption (the amount is hypothetical): a Thai entity pays a service fee of 100,000 THB (excluding tax) to an AI development company inside Thailand.

ItemStandard WHT at 3%e-Withholding Tax at 1%
Service fee (excluding tax)100,000 THB100,000 THB
VAT at 7%+7,000 THB+7,000 THB
Amount withheld−3,000 THB (3%)−1,000 THB (1%)
Actual remittance to the vendor104,000 THB106,000 THB

The arithmetic: 100,000 + 7,000 − 3,000 = 104,000. And 100,000 + 7,000 − 1,000 = 106,000.

The total payable (service fee plus VAT) is the same; what changes is the split between what is withheld and remitted to the state and what goes directly to the vendor. From the vendor’s side more cash lands in hand, so in some cases e-Withholding Tax capability can be used as a negotiating point. Note that the treatment of VAT (whether and when input tax can be credited) and the refund or offset of withholding tax require case-by-case judgement — confirm with a tax specialist.

Contracting with an overseas vendor — reverse-charge VAT and PND54

If you contract with a Japanese AI development company, or with a vendor in Singapore, Europe or the US, there are additional issues.

Where a company outside Thailand supplies services from abroad and those services are used in Thailand, the supply is deemed to take place in Thailand and VAT at 7% arises. In that case the Thai entity receiving the service must file on behalf of the overseas supplier, using form P.P.36. This is the reverse charge.

Further, payments out of Thailand are described as highly likely to be subject to withholding tax (PND54) at the same time, and taking specialist advice in advance is recommended. The applicable rate and whether a tax treaty applies vary by case, so this article does not assert a figure.

The practical implication: an overseas vendor that looks cheaper may score differently once the filing work and cost arising on the Thai side are included. And these filing obligations are not the kind you can excuse by saying you did not know about them. At the negotiation stage, settle these three points with your own finance and tax team.

  1. Which entity contracts (Japanese parent or Thai entity)
  2. Where the “place of use” of the service is determined to be
  3. Whether VAT and withholding tax are inside the contract amount or borne separately (whether there is a gross-up clause)

The third is a drafting question. A contract that says only “amounts exclude taxes” leaves who bears them to be argued later.

Thailand’s AI Act is not yet enacted — which is exactly why it belongs in the contract

One further issue is time-sensitive. Thailand’s Electronic Transactions Development Agency (ETDA) published a new draft of the Draft Act on Artificial Intelligence on 2 July 2026 and ran a public consultation of about 30 days. That consultation window is understood to have closed in early August 2026. As of writing, the Act has not been enacted.

The Act is not yet enacted, and the timing of enactment is not fixed. So this article does not tell you to “comply with the Thai AI Act”. The draft does, however, contain elements a buyer can usefully act on now.

The draft borrows a risk-based structure from the EU AI Act while carrying Thailand-specific elements: strict liability for AI-related harm, an obligation for foreign providers to appoint a local representative, and a labelling obligation for AI-generated content. The risk tiers comprise prohibited-risk AI (manipulation of cognition and behaviour through subliminal techniques; AI systems producing unfair and widespread discrimination) and high-risk AI (to be designated by royal decree, covering systems affecting national security, health, the environment, energy, telecommunications and transport). Subsequent royal decrees may additionally require notification to, registration with, or licensing by the regulator before certain AI systems are deployed.

One practical conclusion follows. Write into the contract now who handles compliance if and when the Act is enacted. Clauses worth considering:

Possible requirementWhat to put in the contract
Local representative obligation for foreign providersIf you contract with an overseas vendor, who arranges the local representative and who bears the cost, should the obligation arise
Labelling obligation for AI-generated contentIf a labelling function has to be implemented, whether it is treated as additional development or as within the maintenance scope
Notification and registration of high-risk AIDivision of work in preparing filings if the system falls in scope, and the obligation to supply the necessary technical documentation
Change of risk classificationThe duty to consult if amendments to legislation or royal decrees change the classification, and the principle for allocating cost

Even a single sentence to the effect that “the parties will consult separately if and when the legislation is enacted” is far better than writing nothing. Ideally you also fix the framework for that consultation: who, by when, and on what basis costs are shared. Country-level differences in practice around AI adoption in Thailand are also covered in our article on AI implementation in Thailand.

On both tax and legislation, this article provides general information only. Confirm any actual decision with tax and legal specialists.

How to Choose an AI Development Company in 2026 — Contracts, Scope and Cost - figure 3

Ten questions to always ask an AI development company

Here is the whole discussion reduced to a question list you can use in a meeting. Each comes with why you ask it and what answer should worry you.

1. Which of the three checklist types (general-purpose AI service use / customized / new development) do you think this project falls into?

Why ask: different types raise different issues. The answer tells you how much contracting experience the vendor actually has.

Red flag: an instant “we’ll build everything bespoke”. Parts that an off-the-shelf service would cover may be getting pushed into new development.

2. Can we run this as separate contracts per phase, along the lines of exploratory staged development?

Why ask: this is the test of whether you can create a place to stop.

Red flag: pushing back with “a single contract works out cheaper”. You would be trading the freedom to withdraw for the discount.

3. If the assessment phase concludes the project is not feasible, what gets delivered?

Why ask: it tells you what you hold if things go badly.

Red flag: no clear answer, or “that doesn’t happen to us”.

4. For each phase, is the contract a contract for work or a quasi-mandate? If a quasi-mandate, performance-ratio or deliverable-completion?

Why ask: it forces the presence or absence of an obligation to complete into the open.

Red flag: vague framing like “it’s a quasi-mandate but effectively the same as a contract for work”. Legally they are not remotely the same.

5. What is the denominator of the accuracy figure in your proposal — per character, per field, or per form?

Why ask: as the thought experiment above showed, the denominator changes the required effort by an order of magnitude.

Red flag: no immediate answer, or “generally we get around 99%” with no denominator attached.

6. How will the acceptance dataset be decided? How many records, with what breakdown?

Why ask: acceptance does not work unless the population is fixed.

Red flag: trying to settle it with “we’ll run it and see if there are problems”.

7. If the acceptance criteria are not met, how many rounds of remediation over how many weeks are free of charge? And what happens after that?

Why ask: it checks the exit route when things go wrong.

Red flag: “we’ll keep going until it’s met”, with no time limit. Unlimited liability is never actually performed.

8. Will the trained model and derived models built on our data be reused on other clients’ projects?

Why ask: this is the core IP issue.

Red flag: stopping at “we may use it as generalised know-how” without going on to define the scope.

9. Is integration with our existing systems and permissions design included in this estimate? If not, what has to be fixed before you can quote it?

Why ask: it closes the gap where layer 4 goes missing.

Red flag: leaving integration out of scope with “your side handles that” and giving no conditions.

10. After go-live, who does the retraining, at what frequency, and with what team? And which entity is the contracting party — the Thai company or the Japanese one?

Why ask: it checks layer 5 and the entry point to the tax and legal issues at the same time.

Red flag: a sudden change of personnel when operations come up, or brand-new information such as “maintenance is a different company”.

Ask these ten in writing and get the answers in writing wherever you can. Verbal explanations do not survive into the contract. If the vendor attaches the answer document to the proposal, it becomes a shared reference point at acceptance.

Six common failure patterns

Here are the failures we see most often, in cause-and-countermeasure form.

Pattern 1: the project stalls at PoC and quietly goes dormant

Cause: the PoC exit condition was defined at the level of “try it and get a feel for it”. With no criterion for progressing to the next phase, nobody can make the call.

Countermeasure: agree in writing, before the PoC starts, which numbers have to reach what level to move into the development phase. Write the termination procedure for the other case at the same time.

Pattern 2: “accuracy” is read two different ways at acceptance

Cause: the contract contains only an accuracy metric with no denominator.

Countermeasure: fix the acceptance dataset in a schedule and write the criteria in business metrics (the count and hours of manual work that remain).

Pattern 3: the model works but never connects to the business systems

Cause: layer 4 (integration with existing systems and permissions design) was out of scope. After go-live, someone is still shuttling CSV files by hand.

Countermeasure: run the five-layer matching on the estimate and close the layer 4 gap before you place the order. Build the list and count of systems to integrate with yourself, in advance.

Pattern 4: accuracy drops six months in and nobody can fix it

Cause: layer 5 (operation and retraining) was not contracted for. The engineers who built it have moved to another project and no retraining procedure was left behind.

Countermeasure: include the additional learning phase in the contract from the start. Make handover of the scripts and documentation needed for retraining part of the acceptance criteria.

Pattern 5: you cannot switch vendors

Cause: ownership of the trained model is ambiguous, or the runtime environment is vendor-specific and the form of handover was never defined.

Countermeasure: put the handover list from the IP section into the contract. State background IP and foreground IP separately.

Pattern 6: a tax problem surfaces just before go-live

Cause: reverse-charge VAT and PND54 on payments to an overseas vendor come to light after signature. With no gross-up clause, the parties argue over who bears them.

Countermeasure: put the choice of contracting entity at the very front of the negotiation. Involve finance and tax from the estimate-comparison stage onwards.

What the six have in common is that none of them is a technical problem; all are problems of what was agreed in advance. Which is precisely why comparing lists of vendors does not prevent them.

A 90-day plan

Finally, a roadmap for execution. The following is an example of a standard sequence; it moves either way depending on project size and the state of your data.

WeekWhat to doOwnerOutput
Weeks 1–2Narrow the target process. List three processes that hurt, then pick oneBusiness function + ITTarget process decided
Week 3Run the 28 METI checklist items against your project and identify the ones you cannot answerIT + admin functionsList of open items
Week 4Determine which of the three types applies. Test first whether an off-the-shelf service is insufficientITType classification memo
Weeks 5–6Inventory the data. Establish in-house where the target data sits, how many records, over what period, with what gapsShop floor + ITData status sheet
Week 7Decide the contracting entity (Japanese parent / Thai entity). Bring finance and tax inAdmin functionsContracting entity decided
Week 8Write the RFP. State the staged-contract premise, your IP preferences and your approach to acceptanceIT + admin functionsRFP issued
Weeks 9–10Receive proposals from several vendors. Put the ten questions in writing and get written answersITProposals + answer documents
Week 11Match estimates across the five layers. Query empty layers in writing. Restate as a three-year totalIT + admin functionsLayer-by-layer comparison
Week 12Review the draft contract. Have lawyers and tax specialists check it. Draft the acceptance criteria scheduleAdmin functions + external specialistsContract draft finalised
Week 13Sign the assessment-phase contract. Final check of entry and exit conditionsManagementSignature and kick-off

Notice that within these 90 days, you are actually talking to AI development companies for two to three weeks. The rest is time spent deciding, internally, what you have to decide. Reverse the order and start with vendor visits, and you will have the vendor decide your open items for you — and those decisions will always land in the vendor’s favour.

If the data inventory in weeks 5–6 slips, everything slips. It is the hardest piece to size, so start it with time in hand. The specialist review in week 12 also depends on external availability, so make the approach early.

FAQ

What is the difference between an AI development company and an AI consulting firm?

There is no agreed industry definition, but in practice the emphasis differs. Firms that describe themselves as AI consulting tend to be strong at layer 1 — problem definition and data survey — judging which processes AI should be applied to and organising the investment case. AI development companies centre on layer 3, the model and application build. The problem is that layers 2 (data preparation), 4 (integration with existing systems) and 5 (operation) can fall into neither party’s remit and float. If you engage both, put on paper at the outset who holds which of the five layers. And if you contract consulting separately, check whether its deliverables come in a form the development company can actually pick up.

Why do AI development outsourcing quotes vary so much between companies?

The main driver is not day rates but difference in scope. As the estimate section showed, whether layer 4 (integration and permissions design) and layer 5 (operation and retraining) are included moves the total a great deal. A further large factor is different assumptions about layer 2. A vendor quoting on the premise that “the data is already prepared” and one quoting on the premise of “starting from correcting inconsistent notation” will produce completely different effort figures for the same project. When you compare, do not line up the totals — tag the lines to the five layers and level the scope first. Once the scope is levelled, the price gap is usually smaller than you expected.

Is it better to outsource custom AI development or hire AI developers in house?

Treating it as a binary is the wrong framing. A realistic split across the five layers keeps layer 1 (problem definition) and layer 5 (operation) close to your own organisation and sends the technically heavy parts of layers 2 to 4 outside. Two reasons. First, layer 1 does not get more accurate without your own understanding of the business. Second, if you outsource layer 5 entirely, detection of accuracy drift is delayed and you can do nothing the moment the contract lapses. Conversely, trying to build layer 3 in house from the start takes too long in recruiting and retaining people. Our article on AI insourcing support sets out the layer-by-layer division of responsibility in detail.

Can an AI system development contract be a contract for work?

It is not that it cannot be done, but if you ask for a contract for work covering the AI model itself, the general pattern is that the vendor declines or prices in a risk premium. Because AI accuracy depends on training data, vendors are said to tend to avoid contracts for work. Two designs are realistic. One is to separate the AI model from the surrounding system and carve only the surrounding system, where specifications can be fixed, into a contract for work. The other is the middle-ground deliverable-completion quasi-mandate. If you take the latter, note that consultation with a specialist on the definition of “completion” is treated as essential; if you cannot write that definition down, the deliverable-completion type is a formality only. Always confirm the final choice of contract type with a lawyer.

Which part of AI development cost varies the most?

Layers 2 and 4 — data preparation and integration with existing systems. Layer 2 varies because the true state of the data is unknown until someone starts work. Being told “we have three years of data” and then finding that the format changed midway, or that each operator followed different input rules, is a common story. Layer 4 grows more than linearly with the number of systems to integrate with and the specification of the interfaces. The best way to contain the variance in both is to make the assessment phase a standalone contract, establish the actual state of the data and the existing systems there, and then re-quote layers 3 onwards. Staged contracting does not only create the freedom to withdraw; it is also a device for improving estimate accuracy.

At what stage can we start an AI consulting conversation?

The concept stage is fine. In fact, waiting until the RFP is written can lock in assumptions about layers 2 and 4 that do not match reality. That said, three things prepared in advance make the conversation move much faster. First, candidate target processes (three listed, narrowed to one). Second, the current state of the data — where it sits, how many records, over what period, with what gaps. Third, a hypothesis about how you will measure the benefit. On the third, our article on measuring the effect and ROI of AI adoption is a useful starting point. With those three in hand, a first meeting can get as far as an initial feasibility view.

Summary

To restate the argument.

AI development proceeds by an inductive method grounded in training data, so the limits of accuracy may already be contained in that data, and setting an obligation to complete or a performance guarantee is difficult. The assumption that picking a good company means you get a finished system is structurally hard to sustain. Which is why, before comparing lists of vendors, you need to design four things.

  1. Contract type: contract for work or quasi-mandate. If a quasi-mandate, performance-ratio or deliverable-completion. Whether you can write down the definition of “completion”
  2. How to split the phases: four phases — assessment, PoC, development, additional learning — with entry and exit conditions on each. Create a place where you can stop
  3. Acceptance criteria: state the denominator of any accuracy figure, fix the acceptance dataset in a schedule, and write the criteria in business metrics (the count and hours of manual work that remain)
  4. IP allocation: distinguish foreground IP from background IP, and set terms separately for raw data, the training dataset, the trained model and derived models. Always confirm whether reuse on other clients’ projects is permitted

If you contract from Thailand, add to that the choice of contracting entity, the treatment of WHT, VAT and the reverse charge, and the allocation of work should the not-yet-enacted Thai AI Act come into force. Always confirm tax and legal points with specialists.

And as the 90-day roadmap shows, most of this design work can be finished inside your own organisation before you meet a single AI development company. Simply respecting that order prevents a large share of selection failures.

TOMAS TECH is based in Bangkok and builds IT and OT systems for the factories of Japanese manufacturers across Thailand and ASEAN. We are glad to talk at an early stage — “we are still at the concept stage and have not decided which process to target”, or “we would like help reading the estimates other vendors have sent us”. Use us as a sounding board before you choose a supplier. Get in touch through our contact form.

References