An AI PoC does not end when the demo works. It ends when the people accountable for the business, technology, data and operations can use the same evidence to decide whether to continue, stop or move to production. This guide is written for manufacturers operating in Thailand and ASEAN. It turns AI PoC exit criteria, acceptance deliverables, production transition and operational handover into one decision package. The focus is not another product demo or generic vendor comparison; it is closing the investment decision without creating PoC fatigue.
The short answer: use four decisions and six separate gates
The exit criterion should not be one accuracy number. Assess business value, quality, critical risk, security/privacy/legal readiness, operability and economics separately. The steering group must record one of four decisions.
| Decision | Meaning | Required next action |
|---|---|---|
| Go | Every mandatory gate passed and the production scope and funding are approved | Begin a limited release with monitoring and staged expansion |
| Conditional go | Value is proven, but bounded remediation remains | Record the condition, owner, due date, evidence and consequence of failure |
| Hold | Evidence or a prerequisite is missing | Define one narrow validation cycle; prohibit an open-ended extension |
| Stop | Value, critical risk, operability or economics are unacceptable | Capture the learning and close access, data, environments and contracts safely |
A critical safety event, personal-data exposure or prohibited action cannot be offset by good average accuracy. A weighted score of 80 does not create a Go if a critical gate fails. Conversely, a PoC need not reach perfection. Residual risks that can be owned and controlled in limited production should be separated from risks that must be closed before release.
NIST describes the AI Risk Management Framework as voluntary and structures AI risk work around Govern, Map, Measure and Manage. NIST AI 600-1 applies that thinking to generative AI risks. Therefore, the sample thresholds in this article are not NIST, ISO or ETDA mandates. They are TOMAS TECH recommendations that an organization should tailor to the use case and impact.
Why PoC fatigue is usually an exit-design problem
PoC fatigue occurs when teams repeat demonstrations, data preparation and reviews but can neither promote nor stop the initiative. A common sequence is:
- The project starts with a broad goal such as “use AI to improve efficiency.”
- A polished demo increases expectations.
- Success metrics, excluded uses, representative data and critical failures remain undefined.
- “Accuracy,” “usefulness” and “safety” mean different things to each department.
- Small enhancements continue, but no decision record, production budget or operating owner exists.
- The budget year or project team changes, and the same PoC restarts under a new name.
A stronger model alone does not solve this. Before work begins, define what uncertainty must be reduced and which evidence will close the decision. If your unresolved question is who should build the system, first use our AI development company selection guide. This article begins after the supplier and PoC boundary are known.
A PoC buys evidence, not a cheap production system
The PoC tests hypotheses: the use case creates value; the necessary data exists; the risk can be controlled; and production benefit can exceed total cost. PoC code not being production code is not automatically a failure. However, a demo that leaves no evaluation data, configuration record, constraint list, migration estimate or operational plan creates little reusable value. Accept the reduction in uncertainty, not only the visible feature.
Fix five boundaries before the AI PoC begins
1. Decision statement
Write who will decide what and by when. For example: “By 30 November 2026, the Quality Director, Plant IT Manager and Business Sponsor will decide whether to deploy the maintenance-report assistant at one plant.” “Explore AI potential” is not a decision.
2. Included and excluded use
Define users, sites, languages, inputs, outputs, connected systems and authority. Drafting a report may be included while stopping equipment, evaluating employees or sending data externally without approval is excluded. An exclusion is a risk boundary, not necessarily a product defect.
3. Baseline
Measure current cycle time, throughput, errors, rework, delay and loss. Use the same definitions before and after the PoC. Preserve the median, 90th percentile and important segments, not only the average.
4. Test and prohibited data
Include normal cases, edge cases, missing data, long documents, languages, old forms and hostile inputs. Document what population the set represents. For personal data, customer secrets or drawings, record purpose, access, location, retention, deletion and cross-border processing. A technical pass is not legal approval.
5. Configuration freeze
Assign identifiers to the model, prompt, retrieval index, code, API, data version and parameters. If one changes during testing, rerun the affected evidence. Scores from an old build cannot accept a different production build.
Build the AI PoC exit scorecard with six gates
Gate 1: business value
Measure whether work improves, not merely whether the function responds. Choose two to four measures such as cycle time, capacity, first-pass completion, rework, delay, avoidable loss and adoption. Define the baseline, target, method, period and data owner.
“Reduce report time” is vague. “For 400 approved maintenance reports, reduce median time from information collection to review submission from 30 to no more than 18 minutes while keeping rework below 10%” is reproducible. Measure the whole workflow so that saved drafting time is not replaced by extra checking.
Gate 2: quality and effectiveness
Use task-specific measures. Classification may need precision, recall, F1 and different costs for false positives and negatives. Generated text may need factual consistency, source support, mandatory-field coverage, prohibited-content checks and readability. For retrieval, assess retrieval relevance separately from the final answer.
A single “95% accuracy” hides the rubric and denominator. Preserve the grading guide, gold-data author, disagreement process and retest rule. Evaluation platforms such as OpenAI Evals can structure criteria, data and graders, but a tool cannot decide which failure is acceptable for the business.
Gate 3: critical risk and safety
Treat critical risk as pass/fail outside any average score. List prohibited equipment actions, unsafe procedures, legal determinations, price approvals, personal-data exposure and confidential-data release. Test the expected refusal, escalation, human confirmation and audit record for each scenario.
No observed incident is weak evidence if the sample never challenged the boundary. Create negative tests from the risk analysis. For detailed agent test-case design, see our AI agent acceptance testing guide; here, the point is that critical evidence must be part of the PoC decision.
Gate 4: security, privacy and legal readiness
Review identity, least privilege, secrets, encryption, logging, vulnerability response, suppliers, retention, deletion, terms, intellectual property, personal data and cross-border processing. ETDA’s organizational generative-AI guidance likewise combines benefits, limitations, risk and governance.
The deliverable is not a one-line “PDPA compliant” claim. It is a data flow showing what data, whose data, which purpose, which recipient, which location, how long and how deletion works; plus responsibilities and open exceptions. Legal conclusions depend on the parties, sector, data and processing location and require case-specific review.
Gate 5: operability
A one-time success in a PoC environment is different from a service that can be supported when the developer is unavailable. Evaluate monitoring, alerts, first response, escalation, backup, recovery, rollback, model/prompt change, cost ceilings, user support and training.
An organization might recommend 99.5% availability, p95 response within eight seconds and acknowledgement of a critical alert within 15 minutes for a limited release. These are examples, not universal standards. A daytime internal assistant and a 24-hour production process require different service objectives.
Gate 6: economics and scale
Include production migration, monthly operation, model/API use, monitoring, reevaluation, data work, training, licensing, support and exit cost—not only the PoC invoice. Adjust labor-time savings for adoption and realizability; avoid counting the same quality benefit twice.
Scale also means more plants, languages, departments, data classes and permission models, not only ten times more requests. Do not multiply one plant’s value across ten plants without costing local formats, networks, owners and support windows.

Do not let a weighted average hide a mandatory failure
The following is an example scorecard, not a standard.
| Gate | Example exit condition | Evidence |
|---|---|---|
| Business value | Median cycle time down at least 30%; rework no more than 10% | Baseline and PoC measured with the same definition |
| Quality | Primary-task success at least 90%; key segment at least 85% | Frozen evaluation set and human audit |
| Critical risk | Zero unacceptable failures in defined critical scenarios | Pass/fail; cannot be offset |
| Security etc. | Zero unresolved high-risk issue; approved flow, retention and deletion | Sign-off by IT, legal and data owner |
| Operations | Monitoring, runbook, recovery and rollback demonstrated | Exercise records |
| Economics | Payback within the approved limit in a conservative case | Input sheet, formulas and sensitivity analysis |
Weighted scores can prioritize noncritical work—for example, business 30, quality 25, operations 20, economics 15 and scale 10. But freeze the rule that every mandatory gate must pass before using the total.
Make AI PoC acceptance criteria recalculable
Freeze denominator, period and segments
Ninety successes out of 100 carry different uncertainty from 900 out of 1,000. Do not remove timeouts, manual rescues or excluded cases from the denominator without showing them. Report total, excluded with reasons, success, partial success, failure and undecided. An overall 90% can still conceal 60% performance on Thai handwritten input, so set minimums for material segments.
Preserve the sample-size formula
A simple planning approximation for a proportion is n = z² × p × (1-p) / e². Using z=1.96, expected success p=0.90 and margin e=0.05 gives 1.96²×0.9×0.1÷0.05²=138.30, rounded up to 139. If the team chooses a simple 20% management buffer for invalid cases and segment coverage, 139×1.2=166.8, or 167. That is not the same as guaranteeing 139 valid cases when 20% may be invalid; that requirement is 139÷0.8=173.75, rounded up to 174. State which reserve definition is being used.
This assumes a simple random sample. A small population, class imbalance, correlated cases, rare catastrophic events or segment-specific confidence needs another design. Preserve assumptions, sampling date, random seed and data version.
Do not outsource acceptance to an LLM judge
Automated graders are useful for regression at scale, but they also vary and fail. Calibrate them against a human-reviewed gold set; record grader version, prompt and settings; inspect high-impact failures manually. An average model-judge score is not approval.
Never reuse a score after an untested change
Changes to the model, prompt, retrieval corpus, tool, API or guardrail can alter both quality and risk. Run impact analysis and the required regression set. Produce a difference list between the accepted PoC build and the production candidate.
Require 12 AI PoC acceptance deliverables
| Deliverable | Required content | Primary accepter |
|---|---|---|
| 1. Decision summary | Hypotheses, result, proposed decision, residual risk | Business sponsor |
| 2. Scope/exclusions | Users, work, data, integration, prohibited use | Process owner |
| 3. Architecture/BOM | Models, APIs, data, environments, versions, licences | IT/architect |
| 4. Data inventory/flow | Source, rights, classification, location, retention, deletion | Data owner/legal |
| 5. Evaluation plan | Metrics, denominator, data, grading, pass/fail, retest | QA/process owner |
| 6. Results/raw evidence | Case results, logs, failures, exclusions, calculation sheet | QA |
| 7. Risk/exception register | Severity, control, owner, date, acceptance | Risk owner |
| 8. Security evidence | Access, secrets, findings, logging, supplier review | Security |
| 9. Economic model | Investment, recurring cost, benefits, sensitivity, payback | Finance/sponsor |
| 10. Production plan | Gaps, phased rollout, rollback, promotion decision | IT/operations |
| 11. Handover package | Runbook, monitoring, SLO, contacts, training, change process | Operations owner |
| 12. Closure evidence | For Stop: access removal, data deletion and contract closure | Project owner |
Specify ownership and editable formats. A PDF alone does not let the customer recalculate economics or thresholds. Hand over the evaluation set, calculation sheet, configuration and runbook in reusable formats, while separately listing vendor background IP, third-party models and open-source licences.

Put measurable acceptance wording in the RFP and contract
Replace “high AI accuracy” with the accepted version and environment; model and data cut-off; customer and supplier responsibilities; metric, threshold, denominator, segment floor and critical-failure definition; frozen test-data rules; remediation count and retest scope; conditional-acceptance terms; deliverables and IP; Stop-time return/deletion; and production-estimate assumptions.
Also define what happens if customer data is late, a foundation-model provider changes a version or an external API changes. The purpose is a controlled change process, not blame. Organizations still framing requirements may also consult our generative AI consulting guide, but the customer must retain acceptance accountability.
Perform a production-gap analysis before promotion
| Area | Typical PoC state | Required production state |
|---|---|---|
| Data | Manual sample upload | Approved integration, quality monitoring, retention/deletion |
| Identity | Shared test ID, broad rights | Individual/service identity, least privilege, review |
| Performance | A few demo users | Peak load, rate limits, queue and capacity plan |
| Availability | Developer recovers manually | Monitoring, on-call, recovery objective, fallback |
| Change | Immediate developer edits | Request, regression, approval and rollback |
| Cost | Free tier or fixed sample | Usage ceiling, budget alert, allocation |
| Contract | Evaluation terms | Production SLA, DPA, support and exit |
| Users | Skilled PoC participants | General users, training, misuse and support |
For every gap, choose “close before production,” “close during limited production” or “accept with an owner.” Do not wait for perfection, but do not silently move the gap to plant users.
Promote through shadow, internal pilot, limited production and expanded production. Each stage needs entry and rollback triggers. Examples include one critical failure, weekly cost above 120% of ceiling or a key segment below its quality floor for two weeks. Set the values for the business before the incident occurs.
Treat operational handover as a production acceptance test
Handover is complete when operations can inspect monitoring, triage, stop, roll back, recover and escalate without informal developer help.
Define RACI for each event: quality degradation, unsafe output, suspected personal data, API failure, cost spike, access change, model update, complaint and major incident. Include the time zones, languages and working hours of Thailand sites, headquarters, local suppliers and cloud providers.
Write runbooks for abnormal paths: alert meaning, screen or query, evidence capture, workaround, stop step, recovery criterion and escalation contact. Do not embed secrets; point to the vault and required role. Exercise rollback to a prior model, prompt, index or application. If an external model cannot be rolled back, test a pinned model, alternative provider, feature shutdown or manual process.
Google Cloud’s generative-AI operations guidance describes capturing production output and running continuous evaluation. Keep the PoC regression and critical-risk sets for operations, and govern whether production examples containing personal or confidential data may re-enter the evaluation system.

A 30/60/90-day path to close the PoC
Days 0–30: design the hypothesis and evidence
Appoint decision, process, data and operations owners. Measure the baseline; define exclusions, critical failures and the economic formula. Approve data rights and security boundaries. Put deliverables, change control and Stop-time deletion in the contract.
Days 31–60: build and run an interim gate
Build the smallest version that can be frozen. Test normal, edge, negative and segmented cases. Separate remediable defects from broken assumptions. Estimate the production gap, monthly cost and support staffing early.
Days 61–90: independent retest and handover
Retest on a holdout not used in development. Exercise critical risk, security, rollback and incident response. Recalculate conservative, base and optimistic economics. Have operations demonstrate the runbook and record one decision.
Do not extend because the system “could get better.” State which uncertainty, which evidence, which budget and which decision date justify the extension.
Worked manufacturing example with explicit assumptions
The following numbers are hypothetical and illustrate recalculation. Assume an assistant saves 12 minutes on each of 4,000 monthly maintenance reports, adoption is 70%, relevant labor cost is THB 600/hour and the quality/realization factor is 0.85.
Time benefit = (12/60) × 4,000 × 0.70 × THB 600 × 0.85 = THB 285,600/month
Assume avoidable rework/loss of THB 90,000/month and run cost of THB 150,000/month.
Monthly net benefit = 285,600 + 90,000 - 150,000 = THB 225,600
With THB 1,800,000 in transition investment, simple payback is 1,800,000/225,600 = 7.98, or 8.0 months.
If adoption is 50% and the quality factor is 0.70, time benefit becomes THB 168,000, net benefit THB 108,000 and payback about 16.7 months. The steering group should use the approved conservative case, not only the demo case. Also test whether saved minutes become usable capacity; fragmented time or extra review may reduce realizability.
Record a one-page decision
| Field | Record |
|---|---|
| Decision | Go / conditional go / Hold / Stop |
| Accepted build | Model, application, prompt, data, evaluation date |
| Gate evidence | Result and evidence link for all six gates |
| Open items | Impact, control, owner, due date |
| Production boundary | Sites, users, volume, rights, start date |
| Stop triggers | Quality, incident, cost and outage triggers |
| Funding | One-time investment, monthly ceiling, contingency |
| Next review | Date, required evidence and decision makers |
Do not invent new criteria in the meeting. A conditional Go without an owner, due date, evidence and stop consequence is a Hold in disguise.
Common AI PoC mistakes
- Accepting the executive demo: demos explain; acceptance proves with representative and negative evidence.
- Using only average accuracy: critical failure and weak minority-language segments disappear; use pass/fail gates and segment floors.
- Comparing PoC cost with benefit: add production, monitoring, support, reassessment and exit costs.
- Leaving conditional items forever: ticket every condition with an owner, evidence and automatic escalation.
- Hiding a Stop: an early Stop that disproves an assumption prevents future loss; preserve the reason and restart condition.
- Delaying handover until after launch: make an operational exercise mandatory for Go.
FAQ about AI PoC exit criteria
When should AI PoC exit criteria be defined?
Before contracting or kickoff. At minimum, agree the decision, scope/exclusions, metrics, critical failures, test data and deliverables. Approve and version any later change before using it for acceptance.
What success percentage is enough for production?
There is no universal percentage. It depends on impact, human review and segment. Combine task quality with zero unacceptable critical failures, segment floors and residual-risk controls. Values in this article are recommendations, not standards.
Should acceptance include ROI?
Yes. Include migration, operation, monitoring, reevaluation, training and exit, not only PoC development. Adjust benefits for adoption and realizability and show sensitivity.
Should a failed PoC receive more development?
Only when a bounded additional test can reduce a named uncertainty. Freeze scope, cost, date and retest criteria. Stop when essential data does not exist, critical risk cannot be controlled or conservative economics fail.
Conditional Go versus Hold?
Conditional Go permits a safe limited production with owned, dated remediation and stop triggers. Hold means evidence or a prerequisite is insufficient to begin use.
What is most often missing during production transition?
An operations owner, monitoring, rollback, cost ceiling, reevaluation after model change and data deletion. Convert developer knowledge into runbooks and exercises.
Who owns PoC deliverables?
The contract decides. Separate customer-specific data, code, settings, calculations and documents from vendor background IP, third-party models and open source. State rights to use, modify, subcontract and retain after termination.
What happens to data after a Stop?
Separate evidence that must be retained from operational data that must be deleted. Cover environments, backups, logs, suppliers and evaluation services; record access removal, return, deletion and confirmation.
Conclusion: close the decision, not just the demo
Ending an AI PoC means closing six gates with recalculable evidence. A Go starts staged production and handover; a Hold names the missing proof; a Stop safely closes access, data and contracts. The most effective cure for PoC fatigue is to contract the exit at the beginning.
TOMAS TECH helps Thailand and ASEAN operations define AI PoC exit gates, acceptance deliverables, economic models, production gaps and handover requirements from the RFP stage. You can contact us even before selecting a product or supplier if you want the continue/stop/produce decision to be based on one auditable evidence set.
Official and primary sources
- NIST AI Risk Management Framework
- NIST AI 600-1 Generative Artificial Intelligence Profile
- ISO AI management systems overview
- ETDA Generative AI Governance Guideline for Organizations
- ETDA guideline PDF
- Google Cloud: Deploy and operate generative AI applications
- OpenAI Evals guide
- Thailand National AI Strategy and Action Plan 2022–2027