What a factory should buy in an edge AI deployment is not a standalone “AI box.” The real purchase is an operable inference system: defined sensor inputs, a model validated on local data, performance that holds under worst-case load, a safe fallback when the system fails, controlled updates and rollback, and an operating model that can support multiple lines. This guide gives procurement, plant management, production engineering, quality, and IT/OT teams one evidence-based framework for moving from an RFP through a 90-day proof of concept (PoC), acceptance, and multi-site scale-out.
Define the purchase unit before deploying Industrial Edge AI
Industrial Edge AI runs inference near production equipment and can support inspection, maintenance, worker safety, and yield improvement. Local processing can reduce latency and bandwidth use, continue during a network disruption, and support data-sovereignty requirements. However, “works without the cloud” does not automatically mean “safe and supportable.” A factory also needs defined behavior when an input degrades, the model becomes stale, an operating-system update fails, storage fills, or a device must recover after an unplanned stop.
Thailand’s Board of Investment (BOI) reported that realized investment in the first half of 2026 reached THB 535.8 billion, with more than THB 255 billion in the second quarter. The same release states that AI-related equipment and infrastructure exceeded THB 127 billion. That AI-related figure is roughly half of the Q2 figure; it is not half of the first-half total. The direction is relevant, but it does not prove the economics of any individual factory project. A plant should decide with its own loss baseline and acceptance evidence.
Buy six acceptance layers, not a bundle of hardware
An RFP should require the following six layers as one system:
- Sensor/input contract: what the system receives, at what quality, frequency, and time reference.
- Worst-case latency and throughput: whether the takt requirement is maintained during peak input, logging, synchronization, and recovery.
- Model quality on local data and drift monitoring: whether errors can be understood by product, machine, defect class, and operating condition.
- Fail-safe boundary and manual fallback: who does what when AI is uncertain, unavailable, or wrong.
- Cybersecurity, patching, rollback, and recovery: whether packages, access, assets, backups, and restoration are controlled and tested.
- Fleet operations across lines and sites: whether deployment, versions, health, suspension, and audit can be managed consistently.
A strong demonstration model does not compensate for missing evidence in any of these layers. Defining this purchase unit also makes the responsibility split between hardware supplier, model provider, system integrator, and plant operations explicit.

Do not turn edge AI and cloud AI into a false choice
The useful question is not “edge or cloud?” but “where should each workload run?” Immediate image decisions, signals close to machine interlocks, and inference that must continue during a WAN outage are candidates for the edge. Long-term analytics, training, cross-plant comparison, model approval, and fleet-level monitoring may belong in a central platform or cloud. Our guide to edge computing architecture for factories explains the placement of compute across tiers. This article focuses specifically on buying, validating, and operating an inference workload in that architecture.
Give each workload these attributes in the RFP:
| Decision | What the RFP must state | Acceptance evidence |
|---|---|---|
| Allowed latency | Maximum time from input to usable result | Time-series log under peak load |
| Network outage | Functions and duration that must continue | WAN isolation and recovery test |
| Data egress | Raw images, features, results, and metadata allowed to leave | Network capture and configuration |
| Retraining | Location, approver, and permitted data | Model registry and approval history |
| Retention | Duration, capacity, and deletion rules | Capacity model and deletion log |
| Safety boundary | Actions the AI may directly influence | I/O drawing, stop test, and procedure |
Claims such as “the data stays on site” or “central management is included” are not evidence. Verify the actual traffic, the behavior during isolation, and the actions that central operators can execute.
Translate the six acceptance layers into RFP requirements
1. Sensor/input contract: stabilize inputs before optimizing the model
Factory inference can fail because of lighting, camera movement, a dirty lens, conveyor-speed variation, sensor clock drift, PLC tag changes, or an incorrect product master—not only because of the model. The input contract should therefore include the signal name, unit, allowed range, sampling interval, timestamp source, missing-value behavior, quality flags, calibration method, and change-notification process.
For vision, specify the field of view, workpiece position, exposure, illumination, trigger, lot association, and re-capture rule in addition to resolution. For time-series sensors, specify measuring range, calibration interval, clock synchronization, and outlier treatment. Inputs that breach the contract should become a distinct “input invalid” state rather than being passed silently to the model. A low-confidence AI result and a defective sensor require different corrective actions.
2. Latency and throughput: buy worst-case behavior, not an average
Vendor demonstrations often use one normal product, a short run, and an unloaded computer. Acceptance testing should combine maximum line speed, maximum camera count, image retention, log transfer, model switching, and post-restart catch-up. Measure input acquisition, preprocessing, inference, post-processing, and notification to the PLC or MES separately. Review tail latency and maximum observed delay, not only the mean.
A line rated at 600 parts per hour averages ten parts per minute, but arrivals may be bursty. Two cameras may trigger close together while the device is writing images and synchronizing logs. The RFP must say what happens when the queue reaches capacity: discard the oldest input, stop the line, reject the workpiece, or route it to manual inspection. Leaving this decision inside a supplier’s default configuration creates an uncontrolled production rule.
3. Model quality: measure local loss, not one global accuracy score
Accuracy, precision, or recall alone is not an acceptance decision. Examine confusion matrices by product, machine, tool, material lot, shift, illumination condition, and defect class. Translate false negatives and false positives into escape risk, scrap, reinspection work, or delayed output. For a deeper treatment of this use case, see our factory AI visual-inspection implementation guide.
Local evaluation data should include startup, end-of-run conditions, changeovers, post-cleaning operation, and minor equipment adjustments—not only clean samples selected for a PoC. Split training and evaluation so nearly identical images from one manufacturing lot do not leak into both. After acceptance, monitor input distribution, confidence, reinspection rate, operator corrections, and performance by product. If drift crosses a defined threshold, return to evaluation. Automatic retraining should not immediately push a new model to production; it still needs approved data, reproducible training, testing, approval, and staged deployment.
4. Fail-safe boundary: do not mistake AI for a safety function
The closer an AI result is to physical actuation, the stronger the boundary must be. An “OK” prediction should not be assumed to replace a safety PLC, emergency stop, guard, or legally required inspection. Separate levels such as advisory output, operator approval, and limited automatic sorting according to potential impact.
The RFP should specify the safe state for low confidence, missing input, model load failure, excessive temperature, full storage, clock mismatch, network loss, and an unavailable upstream system. The operating procedure must name the person who starts manual inspection, the allowed switch-over time, the treatment of in-process material, and reconciliation after recovery. “Switch to manual” is not executable unless people, gauges, work instructions, and capacity have actually been reserved.
5. Cybersecurity, patching, rollback, and recovery: accept a system that can be changed safely
Industrial Edge AI is a composite asset containing an operating system, containers, drivers, runtime, model, configuration, and certificates. The asset inventory should identify each version, owner, dependency, and support date. Require signature and hash verification for packages, role-based access, action logs, and controlled maintenance paths.
Update tests should include not only a successful upgrade but interrupted power, insufficient disk, incompatible dependencies, expired certificates, and rollback to the prior known-good state. The criterion is not “a backup exists.” It is that the team can restore the system within the agreed time, align the model and configuration versions, and resume equipment communication.
NIST SP 1800-41, published on 21 May 2026, is an initial public draft. It says OT/ICS operators need response and recovery plans because defense in depth cannot eliminate every risk. It is not a certification. Use it as a practical prompt for recovery plans, exercises, and evidence in the RFP.
6. Fleet operations: scale one success without losing control
A one-device PoC can fragment when deployed to ten devices or several plants. Versions, configurations, network policies, and local procedures start to differ. Fleet management should track device identity, site, hardware, OS, runtime, model, configuration, last contact, health, deployment history, and rollback history.
In April 2026, Siemens announced general availability of its Industrial AI Suite on Industrial Edge. In that same announcement, however, IEC 62443-4-2-certified security functions and air-gapped operation were roadmap items planned for the second half of 2026. Its published architecture describes vendor-agnostic equipment connections, local inference without cloud connectivity, and cross-site deployment, versioning, and monitoring through AI Asset Manager. This is a product example, not a universal procurement recommendation. Separate roadmap items from currently released functions, verify the procured version, and convert the required capabilities into vendor-neutral acceptance criteria.
NVIDIA IGX is another product example, positioned for industrial and medical edge use with emphasis on safety and security and options for a ten-year lifecycle and support. Long support can be valuable, but ten years should not become a universal rule. Select the required period from equipment life, model-change frequency, spares strategy, and downtime exposure.
Structure a 90-day edge AI PoC around decisions
The PoC is not meant to prove that a demonstration can work. Its job is to reveal why production deployment may fail while the cost of stopping is still low. Divide 90 days into four stages and put exit conditions at each stage.

Days 0–15: baseline the loss and boundary
- Align definitions for defects, downtime, manual-inspection time, reinspection, and scrap.
- Name the target line, products, exclusions, and accountable owners.
- Separate the result that AI produces from the actions equipment may execute.
- Draft the input contract and data-retention conditions.
- Measure the current process and freeze a baseline.
If there is no trustworthy baseline, there is no defensible benefit calculation. Where departments code the same event differently, postpone the technical PoC and first create one measurement definition.
Days 16–30: freeze data and architecture
- Review whether local data represent the real process, and agree labeling rules and exclusions.
- Build a communication matrix between edge, central management, cloud, MES, and PLC.
- Define offline functions, buffer capacity, and resynchronization order.
- Set version control and approvers for model, container, and configuration.
- Design access, maintenance, logs, and threat controls.
At this stage, reproducibility is more important than chasing an extra point of model accuracy. If the same input, version, and configuration cannot reproduce the same result, later fault analysis will be unreliable.
Days 31–60: test operating conditions and failure conditions together
- Evaluate all in-scope products, multiple shifts, changeovers, and post-cleaning operation.
- Reproduce peak load, burst input, and storage pressure.
- Inject WAN loss, central-management loss, time drift, and missing sensors.
- Confirm that operators can correct a decision and retain an audit record.
- Have the actual operating team switch to manual mode and recover using the procedure.
The deliberate stop is essential. If the plant cannot demonstrate continuity or transition to a safe state when AI is unavailable, it will not have a viable maintenance window after go-live.
Days 61–90: evidence acceptance, recovery, and scale-out
- Freeze an acceptance data set and repeat the six-layer measurements.
- Demonstrate a patch, a failed patch, rollback, and backup restoration.
- Deploy a controlled package to another line and review all differences.
- Hand over monitoring, first response, supplier escalation, and change approval.
- Hold a Go, Conditional Go, Retest, or Stop decision review.
Stop at day 90 if any critical layer lacks evidence. “The model should improve with more training” is not a decision. Extend conditionally only when the missing data, proposed change, owner, deadline, and retest criterion are explicit.
RFP clauses you can adapt directly
Scope and deliverables
- Inventory the production equipment, PLCs, cameras, sensors, networks, and upstream systems.
- State responsibility for hardware, OS, AI runtime, model, application, integrations, and support.
- Deliver architecture, I/O table, communication matrix, asset register, model register, SOPs, and test reports.
- Define ownership and use rights for training data, labels, models, configurations, and logs.
- Verify data deletion, hardware removal, and account revocation at the end of a PoC.
Performance and model quality
- Separate targets and minimum gates by product and defect class.
- Measure latency, pending items, and missing results at maximum input.
- Record human disagreement so that decisions can be reevaluated.
- Treat low confidence and “unable to decide” as valid output states.
- Define retraining, approval, release, and rollback after drift.
OT safety and continuity
- Document the path from AI output to equipment and every permitted action.
- Keep safety control, emergency stops, and guards independent of the AI path.
- Define state transitions after network loss, power loss, restart, and upstream outage.
- Reserve people, gauges, instructions, and capacity for manual fallback.
- Test resynchronization, duplicate prevention, and work-in-process reconciliation.
Cybersecurity and operations
- Require least privilege, named accounts, audit logs, and time-limited maintenance access.
- Require signed packages, a vulnerability process, and tested rollback.
- Confirm the update process for restricted or air-gapped networks.
- Assign severity, owner, response time, and escalation to each alert.
- Define configuration export, data return, and migration at contract exit.
Build the business case with a transparent hypothetical model
The following is a hypothetical worksheet, not a market average. Assume one line, two cameras, 600 parts per hour, 16 operating hours per day, and 300 days per year. That produces 5.76 million inspections per year.
| Input | Illustrative baseline | Your value |
|---|---|---|
| Lines | 1 | |
| Cameras | 2 | |
| Production speed | 600 parts/hour | |
| Operating time | 16 hours/day | |
| Operating days | 300 days/year | |
| Annual inspections | 5,760,000 | |
| Image size | Not assumed | |
| Share uploaded to cloud | Not assumed | |
| Network charge | Not assumed | |
| Network downtime | Not assumed | |
| Loss per downtime hour | Not assumed | |
| Edge equipment and support | Not assumed |
The inspection count is 600 × 16 × 300 = 5,760,000. For a cloud comparison, multiply image size by uploaded share and inspection count, then include storage, transfer, network upgrades, and the effect of an outage or manual fallback. For edge, include hardware, spares, energy, cooling, field service, patches, fleet management, and periodic model validation. Do not insert invented vendor prices or a universal saving percentage. Compare the same time horizon and risk boundary.
Benefits can include reduced escapes, lower scrap, less reinspection, earlier maintenance action, and shorter startup. Avoid double counting. If “fewer escapes” and “lower warranty expense” describe the same event, use one line or show the dependency. For each benefit, name the source log that will prove it during the PoC.
Use NIST AI RMF as a lifecycle and TEVV structure
The NIST AI Risk Management Framework is voluntary and does not certify a system. Its Govern, Map, Measure, and Manage functions can keep acceptance from becoming a model-only discussion:
- Govern: owners, approval authority, prohibited uses, change control, training, and audit.
- Map: process, users, impacts, failure scenarios, dependent equipment, and data flows.
- Measure: quality, latency, robustness, drift, security, and manual fallback.
- Manage: prioritize risk and decide whether to mitigate, accept, stop, or retest.
Map PoC evidence to these functions so model performance, plant operation, accountability, and recovery reach the same decision meeting. Every monitored metric needs an owner and an action threshold; otherwise it is only a dashboard.
Confirm BOI incentives case by case
A BOI announcement effective 3 January 2023 indicates that AI/machine learning, big data/data analytics, and qualifying factory-integrated digital technology may be included in specified efficiency and Industry 4.0 investment measures. Buying an edge device does not automatically make a project eligible. The business, technology scope, timing, expenditure, application route, and required outcomes can differ by case.
Do not put an assumed incentive into the base return as if it were guaranteed. Seek case-by-case confirmation from BOI or an appropriate adviser. Run technical acceptance and incentive eligibility as separate workstreams, then combine only verified results.
Score supplier proposals on evidence
Price-only comparison hides operating effort and downtime risk. The following 100-point example must be adjusted to the plant’s loss structure.
| Area | Example weight | Evidence |
|---|---|---|
| Input contract and shop-floor integration | 15 | I/O table, site check, invalid-input test |
| Performance and model quality | 20 | Local evaluation, load test, error analysis |
| Fail-safe and manual fallback | 20 | State transitions, stop demonstration, SOP |
| Security and recovery | 20 | Asset register, signing, rollback, restore test |
| Fleet operations | 15 | Deployment, monitoring, versions, site roles |
| Transfer, support, and total cost | 10 | SLA, training, migration, cost breakdown |
Award points to repeatable PoC evidence, not to “supported” in a proposal. Comparing business outcomes rather than product feature names also keeps the scorecard useful when products change.
Decide: Go, Conditional Go, Retest, or Stop

Go means all six minimum gates are met and each residual risk has an owner and monitoring method. Conditional Go is limited to a minor gap with bounded production impact, a compensating control, an owner, a due date, and a review date. Retest applies when a correction or additional representative data can produce a valid decision. Stop applies when the safety boundary is unclear, recovery cannot be demonstrated, the input contract is unstable, representative data do not meet required quality, or no operating team will own the system.
Ninety days is not a promise to launch. It is a deadline for obtaining evidence or stopping. Treating a Stop decision as loss avoidance prevents an endless PoC from consuming resources without improving production capability.
Common failure patterns
Awarding the project on demonstration accuracy
A curated demonstration does not prove input stability, peak performance, failure behavior, recovery, or fleet operation. Tie payment milestones to evidence in all six layers.
Leaving the AI-to-safety boundary implicit
“Connected to the PLC” is not a definition of automation. Document allowed commands, approvals, hardwired safety, and manual fallback separately and test them.
Treating offline inference as maintenance-free operation
Even an air-gapped inference node has certificates, logs, models, operating software, and vulnerabilities. Include controlled import, approval, rollback, and recovery in the scope.
Copying one PoC image to every line
Equipment, lighting, products, networks, and work practices differ. Separate a common approved baseline from site-specific configuration and perform a short acceptance test on each line.
Responding to every quality change with retraining
Retraining can hide a broken sensor, shifted fixture, or labeling change. Check the input contract, equipment, labels, and operating change before deciding that the model must change.
FAQ: deploying edge AI in a factory
Which process should be selected first for edge AI deployment?
Choose a process where loss is measurable, inputs and ground truth can be collected, and manual operation remains possible when AI is stopped. Vision inspection, condition monitoring, and safety assistance may qualify, but make the decision from loss, data, boundary, and fallback—not from the use-case name.
Is a 90-day edge AI PoC long enough?
It is long enough to assemble the six layers of production evidence and decide whether to proceed. It cannot prove every seasonal condition. If seasonality matters, accept the monitoring method and retain a longer observation gate in the production plan.
Can factory AI inference work without cloud connectivity?
It depends on the product and architecture. Local inference is available in real products, while air-gapped operation can still be a roadmap item for a particular product or version. Verify availability in the version being procured, including licensing, time synchronization, updates, log collection, and model distribution. Include a real WAN isolation test.
What is the most important item in an edge AI RFP?
No single clause is enough. Integrate all six acceptance layers. If the input contract and fail-safe boundary are vague, strong latency and accuracy numbers cannot establish production accountability.
What should edge AI MLOps manage?
It should manage not only a model file but also data definitions, training code, runtime, configuration, targets, approvals, monitoring, and rollback. A factory also needs links to device, product, line, and shift context.
Can this project receive a BOI incentive?
AI/ML and qualifying factory-integrated digital technology may be eligible under conditions, but there is no automatic treatment. Confirm the business, technology, cost, timing, and application requirements case by case with BOI or a qualified adviser.
Conclusion: buy a system that can stop, recover, and scale
The acceptance target for factory edge AI is an operable inference system, not an AI appliance. Test the input contract, worst-case performance, local model quality and drift, fail-safe behavior, patch and recovery path, and fleet operation through the RFP and a 90-day PoC. Inject failure, demonstrate manual fallback and rollback, and stop at day 90 when critical evidence is missing. That discipline turns a successful demo into a sustainable plant capability.
Even if your edge AI plan is still at the requirements stage, TOMAS TECH can help review existing equipment, networks, and quality data and shape the RFP, PoC, and acceptance gates. Contact TOMAS TECH with the target process and current operating challenge.
References
- Thailand BOI, “Realized Investment in Q2 Exceeds THB 255 Billion; H1 Reaches THB 535.8 Billion” (14 Aug 2026): https://www.boi.go.th/upload/content/PR131_2569.pdf
- Siemens, “Siemens Industrial Edge ecosystem strengthens data and AI integration” (21 Apr 2026): https://press.siemens.com/global/en/pressrelease/siemens-industrial-edge-ecosystem-strengthens-data-and-ai-integration
- Siemens, “AI Suite Architecture”: https://www.siemens.com/en-us/content/architecture-hub/ai-suite/
- IBM, “What is edge AI?” (updated 2 Apr 2026): https://www.ibm.com/think/topics/edge-ai
- NVIDIA Developer, “NVIDIA IGX”: https://developer.nvidia.com/igx
- NIST, “AI Risk Management Framework”: https://airc.nist.gov/airmf-resources/airmf/
- NIST, “SP 1800-41, Responding to and Recovering from a Cyber Attack,” Initial Public Draft (21 May 2026): https://csrc.nist.gov/pubs/sp/1800/41/ipd
- Thailand BOI, Announcement No. 15/2565, effective 3 Jan 2023: https://www.boi.go.th/upload/content/15_2565EN.pdf