Blog

2026.09.19

AI PoC Exit Criteria for Production, Acceptance and Handover

AI PoC Exit Criteria for Production, Acceptance and Handover

An AI PoC does not end when the demo works. It ends when the people accountable for the business, technology, data and operations can use the same evidence to decide whether to continue, stop or move to production. This guide is written for manufacturers operating in Thailand and ASEAN. It turns AI PoC exit criteria, acceptance deliverables, production transition and operational handover into one decision package. The focus is not another product demo or generic vendor comparison; it is closing the investment decision without creating PoC fatigue.

The short answer: use four decisions and six separate gates

The exit criterion should not be one accuracy number. Assess business value, quality, critical risk, security/privacy/legal readiness, operability and economics separately. The steering group must record one of four decisions.

DecisionMeaningRequired next action
GoEvery mandatory gate passed and the production scope and funding are approvedBegin a limited release with monitoring and staged expansion
Conditional goValue is proven, but bounded remediation remainsRecord the condition, owner, due date, evidence and consequence of failure
HoldEvidence or a prerequisite is missingDefine one narrow validation cycle; prohibit an open-ended extension
StopValue, critical risk, operability or economics are unacceptableCapture the learning and close access, data, environments and contracts safely

A critical safety event, personal-data exposure or prohibited action cannot be offset by good average accuracy. A weighted score of 80 does not create a Go if a critical gate fails. Conversely, a PoC need not reach perfection. Residual risks that can be owned and controlled in limited production should be separated from risks that must be closed before release.

NIST describes the AI Risk Management Framework as voluntary and structures AI risk work around Govern, Map, Measure and Manage. NIST AI 600-1 applies that thinking to generative AI risks. Therefore, the sample thresholds in this article are not NIST, ISO or ETDA mandates. They are TOMAS TECH recommendations that an organization should tailor to the use case and impact.

Why PoC fatigue is usually an exit-design problem

PoC fatigue occurs when teams repeat demonstrations, data preparation and reviews but can neither promote nor stop the initiative. A common sequence is:

  1. The project starts with a broad goal such as “use AI to improve efficiency.”
  2. A polished demo increases expectations.
  3. Success metrics, excluded uses, representative data and critical failures remain undefined.
  4. “Accuracy,” “usefulness” and “safety” mean different things to each department.
  5. Small enhancements continue, but no decision record, production budget or operating owner exists.
  6. The budget year or project team changes, and the same PoC restarts under a new name.

A stronger model alone does not solve this. Before work begins, define what uncertainty must be reduced and which evidence will close the decision. If your unresolved question is who should build the system, first use our AI development company selection guide. This article begins after the supplier and PoC boundary are known.

A PoC buys evidence, not a cheap production system

The PoC tests hypotheses: the use case creates value; the necessary data exists; the risk can be controlled; and production benefit can exceed total cost. PoC code not being production code is not automatically a failure. However, a demo that leaves no evaluation data, configuration record, constraint list, migration estimate or operational plan creates little reusable value. Accept the reduction in uncertainty, not only the visible feature.

Fix five boundaries before the AI PoC begins

1. Decision statement

Write who will decide what and by when. For example: “By 30 November 2026, the Quality Director, Plant IT Manager and Business Sponsor will decide whether to deploy the maintenance-report assistant at one plant.” “Explore AI potential” is not a decision.

2. Included and excluded use

Define users, sites, languages, inputs, outputs, connected systems and authority. Drafting a report may be included while stopping equipment, evaluating employees or sending data externally without approval is excluded. An exclusion is a risk boundary, not necessarily a product defect.

3. Baseline

Measure current cycle time, throughput, errors, rework, delay and loss. Use the same definitions before and after the PoC. Preserve the median, 90th percentile and important segments, not only the average.

4. Test and prohibited data

Include normal cases, edge cases, missing data, long documents, languages, old forms and hostile inputs. Document what population the set represents. For personal data, customer secrets or drawings, record purpose, access, location, retention, deletion and cross-border processing. A technical pass is not legal approval.

5. Configuration freeze

Assign identifiers to the model, prompt, retrieval index, code, API, data version and parameters. If one changes during testing, rerun the affected evidence. Scores from an old build cannot accept a different production build.

Build the AI PoC exit scorecard with six gates

Gate 1: business value

Measure whether work improves, not merely whether the function responds. Choose two to four measures such as cycle time, capacity, first-pass completion, rework, delay, avoidable loss and adoption. Define the baseline, target, method, period and data owner.

“Reduce report time” is vague. “For 400 approved maintenance reports, reduce median time from information collection to review submission from 30 to no more than 18 minutes while keeping rework below 10%” is reproducible. Measure the whole workflow so that saved drafting time is not replaced by extra checking.

Gate 2: quality and effectiveness

Use task-specific measures. Classification may need precision, recall, F1 and different costs for false positives and negatives. Generated text may need factual consistency, source support, mandatory-field coverage, prohibited-content checks and readability. For retrieval, assess retrieval relevance separately from the final answer.

A single “95% accuracy” hides the rubric and denominator. Preserve the grading guide, gold-data author, disagreement process and retest rule. Evaluation platforms such as OpenAI Evals can structure criteria, data and graders, but a tool cannot decide which failure is acceptable for the business.

Gate 3: critical risk and safety

Treat critical risk as pass/fail outside any average score. List prohibited equipment actions, unsafe procedures, legal determinations, price approvals, personal-data exposure and confidential-data release. Test the expected refusal, escalation, human confirmation and audit record for each scenario.

No observed incident is weak evidence if the sample never challenged the boundary. Create negative tests from the risk analysis. For detailed agent test-case design, see our AI agent acceptance testing guide; here, the point is that critical evidence must be part of the PoC decision.

Gate 4: security, privacy and legal readiness

Review identity, least privilege, secrets, encryption, logging, vulnerability response, suppliers, retention, deletion, terms, intellectual property, personal data and cross-border processing. ETDA’s organizational generative-AI guidance likewise combines benefits, limitations, risk and governance.

The deliverable is not a one-line “PDPA compliant” claim. It is a data flow showing what data, whose data, which purpose, which recipient, which location, how long and how deletion works; plus responsibilities and open exceptions. Legal conclusions depend on the parties, sector, data and processing location and require case-specific review.

Gate 5: operability

A one-time success in a PoC environment is different from a service that can be supported when the developer is unavailable. Evaluate monitoring, alerts, first response, escalation, backup, recovery, rollback, model/prompt change, cost ceilings, user support and training.

An organization might recommend 99.5% availability, p95 response within eight seconds and acknowledgement of a critical alert within 15 minutes for a limited release. These are examples, not universal standards. A daytime internal assistant and a 24-hour production process require different service objectives.

Gate 6: economics and scale

Include production migration, monthly operation, model/API use, monitoring, reevaluation, data work, training, licensing, support and exit cost—not only the PoC invoice. Adjust labor-time savings for adoption and realizability; avoid counting the same quality benefit twice.

Scale also means more plants, languages, departments, data classes and permission models, not only ten times more requests. Do not multiply one plant’s value across ten plants without costing local formats, networks, owners and support windows.

AI PoC Exit Criteria for Production, Acceptance and Handover - figure 1

Do not let a weighted average hide a mandatory failure

The following is an example scorecard, not a standard.

GateExample exit conditionEvidence
Business valueMedian cycle time down at least 30%; rework no more than 10%Baseline and PoC measured with the same definition
QualityPrimary-task success at least 90%; key segment at least 85%Frozen evaluation set and human audit
Critical riskZero unacceptable failures in defined critical scenariosPass/fail; cannot be offset
Security etc.Zero unresolved high-risk issue; approved flow, retention and deletionSign-off by IT, legal and data owner
OperationsMonitoring, runbook, recovery and rollback demonstratedExercise records
EconomicsPayback within the approved limit in a conservative caseInput sheet, formulas and sensitivity analysis

Weighted scores can prioritize noncritical work—for example, business 30, quality 25, operations 20, economics 15 and scale 10. But freeze the rule that every mandatory gate must pass before using the total.

Make AI PoC acceptance criteria recalculable

Freeze denominator, period and segments

Ninety successes out of 100 carry different uncertainty from 900 out of 1,000. Do not remove timeouts, manual rescues or excluded cases from the denominator without showing them. Report total, excluded with reasons, success, partial success, failure and undecided. An overall 90% can still conceal 60% performance on Thai handwritten input, so set minimums for material segments.

Preserve the sample-size formula

A simple planning approximation for a proportion is n = z² × p × (1-p) / e². Using z=1.96, expected success p=0.90 and margin e=0.05 gives 1.96²×0.9×0.1÷0.05²=138.30, rounded up to 139. If the team chooses a simple 20% management buffer for invalid cases and segment coverage, 139×1.2=166.8, or 167. That is not the same as guaranteeing 139 valid cases when 20% may be invalid; that requirement is 139÷0.8=173.75, rounded up to 174. State which reserve definition is being used.

This assumes a simple random sample. A small population, class imbalance, correlated cases, rare catastrophic events or segment-specific confidence needs another design. Preserve assumptions, sampling date, random seed and data version.

Do not outsource acceptance to an LLM judge

Automated graders are useful for regression at scale, but they also vary and fail. Calibrate them against a human-reviewed gold set; record grader version, prompt and settings; inspect high-impact failures manually. An average model-judge score is not approval.

Never reuse a score after an untested change

Changes to the model, prompt, retrieval corpus, tool, API or guardrail can alter both quality and risk. Run impact analysis and the required regression set. Produce a difference list between the accepted PoC build and the production candidate.

Require 12 AI PoC acceptance deliverables

DeliverableRequired contentPrimary accepter
1. Decision summaryHypotheses, result, proposed decision, residual riskBusiness sponsor
2. Scope/exclusionsUsers, work, data, integration, prohibited useProcess owner
3. Architecture/BOMModels, APIs, data, environments, versions, licencesIT/architect
4. Data inventory/flowSource, rights, classification, location, retention, deletionData owner/legal
5. Evaluation planMetrics, denominator, data, grading, pass/fail, retestQA/process owner
6. Results/raw evidenceCase results, logs, failures, exclusions, calculation sheetQA
7. Risk/exception registerSeverity, control, owner, date, acceptanceRisk owner
8. Security evidenceAccess, secrets, findings, logging, supplier reviewSecurity
9. Economic modelInvestment, recurring cost, benefits, sensitivity, paybackFinance/sponsor
10. Production planGaps, phased rollout, rollback, promotion decisionIT/operations
11. Handover packageRunbook, monitoring, SLO, contacts, training, change processOperations owner
12. Closure evidenceFor Stop: access removal, data deletion and contract closureProject owner

Specify ownership and editable formats. A PDF alone does not let the customer recalculate economics or thresholds. Hand over the evaluation set, calculation sheet, configuration and runbook in reusable formats, while separately listing vendor background IP, third-party models and open-source licences.

AI PoC Exit Criteria for Production, Acceptance and Handover - figure 2

Put measurable acceptance wording in the RFP and contract

Replace “high AI accuracy” with the accepted version and environment; model and data cut-off; customer and supplier responsibilities; metric, threshold, denominator, segment floor and critical-failure definition; frozen test-data rules; remediation count and retest scope; conditional-acceptance terms; deliverables and IP; Stop-time return/deletion; and production-estimate assumptions.

Also define what happens if customer data is late, a foundation-model provider changes a version or an external API changes. The purpose is a controlled change process, not blame. Organizations still framing requirements may also consult our generative AI consulting guide, but the customer must retain acceptance accountability.

Perform a production-gap analysis before promotion

AreaTypical PoC stateRequired production state
DataManual sample uploadApproved integration, quality monitoring, retention/deletion
IdentityShared test ID, broad rightsIndividual/service identity, least privilege, review
PerformanceA few demo usersPeak load, rate limits, queue and capacity plan
AvailabilityDeveloper recovers manuallyMonitoring, on-call, recovery objective, fallback
ChangeImmediate developer editsRequest, regression, approval and rollback
CostFree tier or fixed sampleUsage ceiling, budget alert, allocation
ContractEvaluation termsProduction SLA, DPA, support and exit
UsersSkilled PoC participantsGeneral users, training, misuse and support

For every gap, choose “close before production,” “close during limited production” or “accept with an owner.” Do not wait for perfection, but do not silently move the gap to plant users.

Promote through shadow, internal pilot, limited production and expanded production. Each stage needs entry and rollback triggers. Examples include one critical failure, weekly cost above 120% of ceiling or a key segment below its quality floor for two weeks. Set the values for the business before the incident occurs.

Treat operational handover as a production acceptance test

Handover is complete when operations can inspect monitoring, triage, stop, roll back, recover and escalate without informal developer help.

Define RACI for each event: quality degradation, unsafe output, suspected personal data, API failure, cost spike, access change, model update, complaint and major incident. Include the time zones, languages and working hours of Thailand sites, headquarters, local suppliers and cloud providers.

Write runbooks for abnormal paths: alert meaning, screen or query, evidence capture, workaround, stop step, recovery criterion and escalation contact. Do not embed secrets; point to the vault and required role. Exercise rollback to a prior model, prompt, index or application. If an external model cannot be rolled back, test a pinned model, alternative provider, feature shutdown or manual process.

Google Cloud’s generative-AI operations guidance describes capturing production output and running continuous evaluation. Keep the PoC regression and critical-risk sets for operations, and govern whether production examples containing personal or confidential data may re-enter the evaluation system.

AI PoC Exit Criteria for Production, Acceptance and Handover - figure 3

A 30/60/90-day path to close the PoC

Days 0–30: design the hypothesis and evidence

Appoint decision, process, data and operations owners. Measure the baseline; define exclusions, critical failures and the economic formula. Approve data rights and security boundaries. Put deliverables, change control and Stop-time deletion in the contract.

Days 31–60: build and run an interim gate

Build the smallest version that can be frozen. Test normal, edge, negative and segmented cases. Separate remediable defects from broken assumptions. Estimate the production gap, monthly cost and support staffing early.

Days 61–90: independent retest and handover

Retest on a holdout not used in development. Exercise critical risk, security, rollback and incident response. Recalculate conservative, base and optimistic economics. Have operations demonstrate the runbook and record one decision.

Do not extend because the system “could get better.” State which uncertainty, which evidence, which budget and which decision date justify the extension.

Worked manufacturing example with explicit assumptions

The following numbers are hypothetical and illustrate recalculation. Assume an assistant saves 12 minutes on each of 4,000 monthly maintenance reports, adoption is 70%, relevant labor cost is THB 600/hour and the quality/realization factor is 0.85.

Time benefit = (12/60) × 4,000 × 0.70 × THB 600 × 0.85 = THB 285,600/month

Assume avoidable rework/loss of THB 90,000/month and run cost of THB 150,000/month.

Monthly net benefit = 285,600 + 90,000 - 150,000 = THB 225,600

With THB 1,800,000 in transition investment, simple payback is 1,800,000/225,600 = 7.98, or 8.0 months.

If adoption is 50% and the quality factor is 0.70, time benefit becomes THB 168,000, net benefit THB 108,000 and payback about 16.7 months. The steering group should use the approved conservative case, not only the demo case. Also test whether saved minutes become usable capacity; fragmented time or extra review may reduce realizability.

Record a one-page decision

FieldRecord
DecisionGo / conditional go / Hold / Stop
Accepted buildModel, application, prompt, data, evaluation date
Gate evidenceResult and evidence link for all six gates
Open itemsImpact, control, owner, due date
Production boundarySites, users, volume, rights, start date
Stop triggersQuality, incident, cost and outage triggers
FundingOne-time investment, monthly ceiling, contingency
Next reviewDate, required evidence and decision makers

Do not invent new criteria in the meeting. A conditional Go without an owner, due date, evidence and stop consequence is a Hold in disguise.

Common AI PoC mistakes

  • Accepting the executive demo: demos explain; acceptance proves with representative and negative evidence.
  • Using only average accuracy: critical failure and weak minority-language segments disappear; use pass/fail gates and segment floors.
  • Comparing PoC cost with benefit: add production, monitoring, support, reassessment and exit costs.
  • Leaving conditional items forever: ticket every condition with an owner, evidence and automatic escalation.
  • Hiding a Stop: an early Stop that disproves an assumption prevents future loss; preserve the reason and restart condition.
  • Delaying handover until after launch: make an operational exercise mandatory for Go.

FAQ about AI PoC exit criteria

When should AI PoC exit criteria be defined?

Before contracting or kickoff. At minimum, agree the decision, scope/exclusions, metrics, critical failures, test data and deliverables. Approve and version any later change before using it for acceptance.

What success percentage is enough for production?

There is no universal percentage. It depends on impact, human review and segment. Combine task quality with zero unacceptable critical failures, segment floors and residual-risk controls. Values in this article are recommendations, not standards.

Should acceptance include ROI?

Yes. Include migration, operation, monitoring, reevaluation, training and exit, not only PoC development. Adjust benefits for adoption and realizability and show sensitivity.

Should a failed PoC receive more development?

Only when a bounded additional test can reduce a named uncertainty. Freeze scope, cost, date and retest criteria. Stop when essential data does not exist, critical risk cannot be controlled or conservative economics fail.

Conditional Go versus Hold?

Conditional Go permits a safe limited production with owned, dated remediation and stop triggers. Hold means evidence or a prerequisite is insufficient to begin use.

What is most often missing during production transition?

An operations owner, monitoring, rollback, cost ceiling, reevaluation after model change and data deletion. Convert developer knowledge into runbooks and exercises.

Who owns PoC deliverables?

The contract decides. Separate customer-specific data, code, settings, calculations and documents from vendor background IP, third-party models and open source. State rights to use, modify, subcontract and retain after termination.

What happens to data after a Stop?

Separate evidence that must be retained from operational data that must be deleted. Cover environments, backups, logs, suppliers and evaluation services; record access removal, return, deletion and confirmation.

Conclusion: close the decision, not just the demo

Ending an AI PoC means closing six gates with recalculable evidence. A Go starts staged production and handover; a Hold names the missing proof; a Stop safely closes access, data and contracts. The most effective cure for PoC fatigue is to contract the exit at the beginning.

TOMAS TECH helps Thailand and ASEAN operations define AI PoC exit gates, acceptance deliverables, economic models, production gaps and handover requirements from the RFP stage. You can contact us even before selecting a product or supplier if you want the continue/stop/produce decision to be based on one auditable evidence set.

Official and primary sources