When considering countermeasures for false detections in AI visual inspection, it is risky to simply adjust settings based on the explanation that “lowering sensitivity will reduce the rejection of good units.” False rejects—where good units are incorrectly classified as defective—lead to material loss and increased verification workload. Conversely, escapes—where defective units are incorrectly classified as good—can result in defective products reaching customers. On mass production lines in Thai factories, it is essential to measure both types of errors for each defect type and establish acceptance criteria that include re-inspection lanes and change management. This article does not cover general topics like lighting or training images, but focuses on the practical task of deciding “which judgment to adopt” after implementation, directly at the production site.
First, Separate False Detections, Escapes, and “Needs Verification” into Distinct Categories
Begin by defining what “positive” means. In this article, a “positive” is when the AI judges a unit as defective. In this context:
- A true defective correctly judged as defective is a true positive.
- A good unit incorrectly judged as defective is a false positive (false reject).
- A good unit correctly judged as good is a true negative.
- A defective unit incorrectly judged as good is a false negative (escape).
In manufacturing, false rejects correspond mainly to false positives, while escapes correspond to false negatives. However, in systems where humans re-inspect before final shipment, an AI’s false positive does not necessarily mean final disposal. Always record AI judgments, re-inspection results, and actual disposal or shipment in separate columns.
If you simply add “needs verification” cases to the defective count, the system may appear safe but yield will worsen. Conversely, if you add “needs verification” to the good count, throughput may increase but you lose visibility of escape risk. Always count “needs verification” as a third state, and compare it to the ground truth after human confirmation. Your judgment table should include at least: part ID, part number, lot, image timestamp, defect class, score, applied threshold, model version, AI primary judgment, human final judgment, gate operation, and final disposition. If you match only by timestamp, there is a risk of linking to the wrong workpiece image during continuous conveyance.
For example, hearing “99% accuracy” on a line inspecting 10,000 units per day is not enough for management decisions. If the actual defect rate is low, even a small number of false rejects will fill the reject bin mostly with good units. Conversely, even if overall accuracy is high, a single escape of a critical defect may fail to meet customer requirements. Rather than overall accuracy, first check the number of escapes by defect class, the false reject rate for good units, and the re-inspection workload. Separate your evaluation data population from your mass production data. You cannot call a threshold decision made on training images an “independent acceptance test” if you use the same images.
If you want to organize the conditions and equipment configuration for stable imaging, first refer to PoC, RFP, and Acceptance Testing for AI Visual Inspection Implementation in Thailand. This article focuses on how to translate scores from such equipment into quality actions.
List the “Cost of Misjudgment” for Each Defect Class
Before applying a single threshold value to all defects, create a defect ledger. Include: defect name, customer specification wording, limit samples, process of occurrence, action if escaped, action if a good unit is falsely rejected, time available for re-inspection, appearance likely to be confused, and approver. There is no reason to treat a chip affecting dimensions or strength and a color variation within tolerance with the same cost or threshold. If customer drawings or quality agreements define the judgment, those take precedence. Do not override specifications just because the AI score is high.
Even if you use terms like “critical,” “major,” or “minor” defect classes, do not set thresholds based on names alone. The general principle is to minimize escapes for critical defects and suppress false rejects for minor appearance differences. However, in practice, operation changes depending on the tolerance limit, identifiability, whether re-inspection is possible in later processes, and the likelihood of customer detection. For ambiguous boundaries like “this spot is OK for sample A but NG for sample B,” reach agreement between quality assurance and manufacturing, and store them as limit samples with version control. Do not fill undecided boundaries with majority-vote labels in training data, as this only hides ambiguity in the model.
A threshold is the boundary where “if the defect-likeness score is above this value, immediately reject.” However, a model output of 0.8 does not always mean an 80% probability. Uncalibrated scores should be treated as ranking values, and actions should be based on the actual confusion matrix. If you allow class-specific thresholds, also specify the priority when multiple classes are detected simultaneously. For example, if a critical class is “needs verification” and a minor class is “reject,” which action takes precedence? Each time you add a class, review these conflict rules.
Adjusting the threshold changes both detection rate and precision, and you can include costs in the evaluation function—this is explained in the official scikit-learn documentation. However, you cannot simply transfer their example optimum values to your production line. You must estimate the ground truth, escape costs, good unit loss, and re-inspection capacity for your own factory and products.
Convert Image Inspection Thresholds into Three Actions
With only binary “OK/NG” judgments, borderline units tend to be either all rejected or all accepted. In mass production, consider dividing the score range into three actions: automatic pass for the lower range, isolation and human re-inspection for the middle range, and automatic reject or stop for the upper range. The middle range is not a place to relax quality standards. It is where you design the physical destination, container, label, and authority so that products do not return to the shipping route until a final judgment is made. If humans simply look at the screen and let units pass because “the AI seems fine,” this cannot be called auditable re-inspection.

When mapping these three actions, link actual workpieces—good, borderline, and clearly defective—to their images. Show the imaging position on the conveyor, encoder or trigger, PLC judgment signal, pass route, verification route, and reject gate in a single timeline. If states exist only on the software screen, the location of physical items may be lost during power outages, communication failures, or jams. Also decide in advance how the equipment will handle “judgment incomplete,” “image missing,” or “model stopped” as abnormal modes.
The width of the middle range should be calculated backward from the site’s verification capacity. For example, if 10,000 units/day are inspected, 2% fall into the middle range, and each re-inspection takes 45 seconds, then the verification workload is 200 units × 45 seconds = 9,000 seconds, or 2.5 person-hours/day. This is a simple calculation that does not include breaks, transport, re-imaging, recording, or peak time concentration. If you do not separately verify how many units arrive per minute at peak and how long the isolation buffer can hold them, the line may appear feasible on a daily average but actually jam. If there are no verification staff on the night shift, do not change the middle range logic itself, but establish safe handling procedures for held items.
Lowering the threshold generally makes it easier to catch defects, but also increases false alarms for good units. Raising the threshold has the opposite effect. However, the curve changes depending on class distribution, image quality, and calibration status. Plot candidate thresholds on the x-axis and false reject rate for good units and escape rate for defects on the y-axis, separated by defect class and part number. There is no absolute “correct” threshold like “0.7 is best.” Compare candidate values using both an approved challenge set and actual production conditions.
Fix Boundaries Using Golden and Challenge Sets
A golden or challenge set is a fixed test group used to ask the same questions every time you change thresholds or models. If you only line up typical units that are easy to pass, you will not see the boundaries most likely to break with changes. For each defect class, include clear good units, good units within tolerance limits, defects just outside the limit, clear defects, and workpieces that are hard to image but valid. If possible, prepare samples that span part numbers, lots, materials, day/night shifts, post-equipment startup, and after jig changes. This is the process of overlapping “limit samples” and “images that confuse the AI,” but do not confuse the two. The former defines quality standards; the latter reveals model weaknesses.

Do not leave the ground truth for the set to a single person’s memory; store it in quality assurance approval records. If opinions are divided on a unit, do not simply decide by majority vote, but return it for review against customer standards or limit samples. Images with undetermined ground truth can be used for training, but do not include them in the denominator for acceptance metrics. Register image hashes, physical storage locations, lighting recipes, imaging versions, and judgment rationale to prevent content from being swapped under the same name. If you cannot store physical items long-term for confidentiality reasons, record that you cannot guarantee identity upon re-imaging.
Separate challenge sets that represent actual occurrence frequency from those intentionally enriched with rare but important defects. The latter are useful for confirming capability with critical defects, but do not directly convert their results into positive predictive value for mass production, as defect rates differ. When estimating re-inspection workload or false reject costs, use the actual part number composition and defect rate from mass production. You cannot infer “half the reject bin will be defective” from evaluation data artificially mixed with “50% defects.”
If you cannot collect enough real defective samples for a class, do not hide the low count. “Zero escapes in 0 cases” is not the same as “zero escape probability.” For example, even if you observe zero escapes in 20 defective samples, this does not prove a very low escape rate with high confidence for that class. Decide the required sample size and confidence interval with your quality department, and follow customer requirements if any. Avoid counting augmented images (rotated, duplicated) as independent samples. Images derived from the same physical item are not statistically independent.
Use a Cost Model to Compare False Rejects and Escapes in the Same Units
Compare threshold candidates using both counts and costs. A simple expected cost can be expressed as:
“Number of false rejects × unit cost of handling good units + number of escapes × unit cost of handling escapes + number of re-inspections × unit re-inspection cost + equipment downtime loss.”
Customer escape costs may include returns, sorting, logistics, replacements, line stoppages, and customer audits, but if the amount is uncertain, show it as an assumption. Do not relax quality assurance limits based solely on minimum cost; first compare costs among candidates that pass constraints such as escape upper limits.
The following is a hypothetical example to illustrate calculation methods. These are not actual TOMAS TECH results or industry averages. Assume 10,000 units/day, with 100 true defects and 9,900 good units. Candidate A detects 98 of 100 defects, falsely rejects 198 of 9,900 good units, and has 200 units in the middle range. Candidate B detects 95 of 100 defects, falsely rejects 99 of 9,900 good units, and has 100 units in the middle range. Do not assume that re-inspection always perfectly recovers errors. For simplicity, the table below shows residual counts after final judgment, assuming both escapes and false rejects actually occur. Assume unit cost for false reject is 300 THB, escape handling is 10,000 THB, re-inspection is 25 THB, and downtime loss is zero for both.
| Item (Hypothetical Example) | Candidate A | Candidate B |
|---|---|---|
| Final false rejects of good units | 198 units | 99 units |
| Final escapes of defective units | 2 units | 5 units |
| Re-inspections | 200 units | 100 units |
| Good unit loss | 59,400 THB | 29,700 THB |
| Escape handling | 20,000 THB | 50,000 THB |
| Re-inspection | 5,000 THB | 2,500 THB |
| Total | 84,400 THB/day | 82,200 THB/day |
In this scenario, Candidate B is 2,200 THB/day less expensive. However, if the contract requires escapes to be no more than 2 units/day, Candidate B is not acceptable even if the cost is lower. If the escape cost is changed to 20,000 THB, Candidate A becomes 104,400 THB/day and Candidate B 132,200 THB/day, reversing the result. If escape cost accuracy is low, present multiple scenarios (e.g., 5,000, 10,000, 20,000 THB) for decision-making. The unit cost for false rejects also differs between reworkable parts and those already destructively tested. Always note the source of unit costs, period covered, and the person responsible for calculations—this is more important than a polished ROI.
Expand the cost formula by defect class. Escape costs for critical defects are high, while re-inspection costs for minor appearance differences may be relatively high. If a single escape triggers sorting of an entire lot, you cannot express the impact with a per-unit cost. Separate worst-case lot handling from normal cases, and present both to management. Rather than fabricating prices to justify automation, disclose them as variables and show how conclusions change with different values—this aids decision-making.
A Re-Inspection Lane Is Not Complete Just Because “A Human Looks at It”
For re-inspection staff, display the physical workpiece, original image, AI-detected region, limit sample, and judgment rationale on the same screen or work order. However, if the AI’s answer is shown too prominently, operators may be biased by it. If necessary, have the human make an independent judgment first, then view the AI score. If the human’s final judgment differs from the AI, do not just mark it as “corrected”—record who overturned what and on what basis. Humans also make mistakes, so for critical defects, set up periodic re-audits or dual checks.
When designing the re-inspection lane, the destination of the workpiece is the top priority. Ensure that units sent to the middle range by the AI do not physically mix with good units in transport. Do not allow units to proceed to the next process until the terminal’s judgment is finalized. Warn operators if they open the wrong part number’s limit sample. Clearly separate the return destinations for good and defective units after re-inspection. Ensure that unprocessed units can be identified even during network outages. These checks involve not only the UI, but also PLCs, gates, barcodes, and container labels.
Designs that cross-check inspection results by part number or lot with downstream process conditions are covered in detail in AI Analysis of Quality Inspection Data. While that article focuses on root cause analysis, the accuracy of individual identification is equally important for auditing misjudgments. If you only collect images misjudged by AI later, but cannot link them to the original workpiece and human action, the data cannot be used for improvement.
In Thai factories, Japanese quality specifications, Thai shop floor instructions, and English equipment screens may coexist. Use language-independent unique IDs for defect codes, and link names and limit samples in all three languages to ensure stable explanations for shift changes and contractors. Do not rely on the nuance of translated terms like “scratch” or “stain”—always attach photos and measurement conditions. For lines with tight verification times, include inspector training workload, queue length during busy periods, and availability of backup staff in your capacity calculations.
Manage Model, Threshold, and Reference Samples as a Single Change
Even changing only the threshold without changing the model will alter the number of false rejects and escapes. Change records should include: current and proposed versions, reason for change, affected part numbers and defect classes, whether training data changed, old and new thresholds, challenge set results, replay results from mass production logs, re-inspection workload, rollback method, approver, and application time. If you change lighting recipes or camera positions, treat these as the same category since they alter the input distribution. For physical conditions related to images, refer to Lighting Selection for AI Visual Inspection; this article focuses on version management for the judgment side.

Before approval, conduct a side-by-side comparison of “new model vs. current model.” Run both versions on the same images and ground truth, and list, by individual unit, the defects newly detected, increased false rejects, and new escapes caused by the change. Comparing only averages can hide a single critical class worsening. If a borderline good unit that passed with the old version is rejected by the new version, do not just call it “safer”—evaluate re-inspection capacity and yield. Maintain separation between training and evaluation data, and avoid overfitting to the same evaluation set by repeatedly tuning while viewing results.
Mass production deployment does not have to be a big-bang switch. First, run shadow operation with both old and new judgments recorded in parallel, but let the old version control the gate. Then, enable the new rules for limited part numbers, lines, or time periods, and revert to the old version in case of anomalies. At this stage, check whether workpiece and log IDs match, whether there is enough time to reach the gate, and whether “needs verification” items are not accumulating. Even if images are the same, if a PLC processing delay or communication retransmission causes the wrong workpiece to be rejected, that is a system-level false reject, not reflected in the model’s score table.
After changes, monitor false reject rates, escape cases, “needs verification” rates, judgment delays, image loss rates, and human override rates by defect class, weekly or by lot. Decide in advance who investigates when thresholds are exceeded, under what conditions to stop shipments or revert to 100% manual inspection, and when to notify customers. The NIST AI RMF provides a framework for detecting performance changes from operational data and feeding measurement results back into risk management. While not a legal requirement for Japanese or Thai factories, it is a useful reference for change management design.
Contract Acceptance Criteria for FAT and SAT: Specify Both Numbers and Actions
Factory Acceptance Test (FAT) is performed in the vendor environment, and Site Acceptance Test (SAT) is performed on the actual line. Even if the same images are correctly classified in FAT, you cannot guarantee local transport speed, vibration, product orientation, lighting variation, or PLC connection. Separate the roles of both tests and specify pass/fail conditions in advance. Based on the PoC and RFP points in the comprehensive inspection machine guide, here we focus on items specific to misjudgment countermeasures.
| Acceptance Item | FAT Example | SAT Example |
|---|---|---|
| Judgment by defect class | Replay fixed challenge set by version and submit class-specific confusion matrix | Re-measure same metrics across specified part numbers and process conditions on actual lots |
| Thresholds and three-way split | Check that pass, needs verification, and reject actions match specifications just before/after each boundary value | Confirm physical destination, buffer, operator terminal, and return location for actual items |
| Traceability | Match images, part IDs, model version, threshold, and judgment logs | Match PLC signals, gate actions, actual workpiece IDs, and disposition |
| Exception handling | Simulate image loss, no score returned, communication loss, restart | Test power recovery, continuous input, shift change, carryover of held items |
| Change and recovery | Confirm rollback to old version, authority, and audit logs | Confirm limited operation, rollback, and isolation of unprocessed items |
A single line such as “accuracy ≥99%” is not sufficient as a numerical acceptance criterion. For example, as a contract proposal: zero escapes in the challenge set for critical defects, escape rate for major defects below an agreed upper limit, false reject rate for good units below a part number-specific upper limit, and “needs verification” rate below shift-specific processing capacity. However, the meaning of “zero cases” depends on sample size and set composition. Do not use example numbers as guaranteed values for your actual factory; always agree based on customer requirements, process risk, and available ground truth. Metrics only become contractual when you specify measurement period, denominator, exclusion conditions, and retest conditions.
Tests often overlook how to handle cases where the inspection target cannot be properly imaged. An image judged “good” by the model is not the same as a workpiece for which no image was obtained. Isolate or stop for imaging failures using a dedicated code, and do not mix them into the normal good unit rate. When the number of “needs verification” items exceeds a certain amount, perform load testing to ensure the gate can correctly handle all target workpieces. Even if the reject signal is timely, if pneumatic gate actuation time or spacing between items is insufficient, the physical item may end up in the wrong bin.
ISO 2859-1:2026 specifies the AQL-based attribute sampling method for lot acceptance inspection. However, the AQL table in this standard does not automatically determine AI image-level thresholds or escape limits by defect class. Lot acceptance and AI judgments on individual items are different units of decision. If you use sampling for customer acceptance inspection, follow that contract, but define AI equipment performance acceptance separately. Clearly distinguish this in RFPs and meeting minutes to prevent vendors from claiming “AQL compliance means AI misjudgments are not a problem.”
Implementation Sequence: What to Agree in 30 Days, What to Prove in Mass Production
The most important thing to avoid is approving the equipment demo first and only later investigating “how the factory actually handled this defect.” If the demo’s ground truth labels differ from the site’s shipping criteria, no matter how high the score, it is not evidence for implementation approval. Also separate whether rework is possible in-process and whether shipment is possible at final inspection. Workpieces that can be sent to “needs verification” during rework may require unsealing costs after packaging. Even with the same image, the optimal action changes depending on process stage. Include imaging position, branching point, and downstream salvageability in your implementation plan.
In the first week, quality assurance, manufacturing, equipment, and IT should narrow the scope to one part number and one process, and finalize the defect ledger and limit sample versions. Next, decide the identification method for actual workpieces and the data columns from AI primary judgment to final disposition. Even if costs cannot be calculated precisely at this stage, decide which costs are significant and who will estimate them. Without this agreement before equipment proposals, even a successful model demo will stall at “what to mark as NG” on site.
In the second week, separately collect images representative of mass production distribution and a challenge set including borderline and rare defects. Split data by individual or lot, ensuring that burst images of the same workpiece do not end up on both training and evaluation sides. Once the quality department approves the ground truth, propose multiple threshold candidates by defect class and simultaneously review false rejects of good units, escapes of critical defects, and “needs verification” counts. Do not mix “assumptions” and “actual measurements” in the same row of your calculation sheet—separate them by column or color.
In the third week, check the actual flow of workpieces in the re-inspection lane and test whether people and buffers are sufficient under peak load. Include machine branching signals, terminal confirmation operations, lot holds, carryover of unprocessed items across shifts, and communication failures. Even if the system can display “needs verification,” if boxes are mixed up on site, the design has failed. Especially on lines with mixed part numbers, confirm during operation that the active part number recipe matches the image metadata.
In the fourth week, assemble the FAT/SAT acceptance table, change record, rollback procedure, and mass production monitoring screen as a set. At mass production start, record the approved version, start time, target lot, and monitoring manager. For the first few lots, conduct full or high-frequency manual audits to build evidence for trusting the AI from actual lots. Do not immediately feed misjudgments found here into training; reflect them in a new version only after ground truth confirmation and change procedures. You can both accelerate improvement and avoid unauthorized changes to accepted performance.
Daily Review Agenda: Focus on “Why Actions Changed” Rather Than Just “Yesterday’s Accuracy”
In short daily reviews at the start of mass production, do not just report yesterday’s accuracy rate. Lay out a few actual false rejects of good units, all actual escapes of defective units, reasons for returning “needs verification” items to good, images that could not be judged, and mismatches between gate and log. For each, tentatively identify whether the cause lies in quality standard interpretation, imaging conditions, model identification, PLC or transport operation, or human action. If the cause is in imaging, increasing training images will not prevent recurrence. If the cause is inconsistency in limit samples, you cannot align site opinions with threshold alone.
After the review, always decide “who will check what by when.” For example, quality assurance determines the ground truth for borderline units, equipment staff check lighting current for the same time period, IT matches image IDs and PLC reject signals, and manufacturing measures the maximum number of items waiting for re-inspection. Even if the same false reject occurs the next day, you can proceed with improvement if causes and actions are recorded. Simply accumulating misjudged images in a folder does not prompt action by the necessary stakeholders. Weekly, review the number of cases by defect class and changes in part number composition to distinguish apparent improvement due to denominator changes from real improvement.
Minimum Request Sheet to Vendors
When requesting a quote, specify not only “want to reduce false detections” but also the deliverables you want returned during testing. At a minimum, request: per-unit score CSV, correspondence table of images and ground truth, confusion matrix by defect class, counts of false rejects, escapes, and “needs verification” for each candidate threshold, delay time distribution, behavior during communication loss, gate operation logs, model and recipe versions, retest methods, and time to revert to old version. Cost estimates must include not only cameras and GPUs, but also isolation conveyors, branching gates, re-inspection terminals, MES integration, maintenance, on-site training, and sample storage—specify who provides each. Also confirm before contract how many times retraining or readjustment is included if acceptance tests fail.
The factory should prepare product specifications, limit samples, defect classes, required processing speed, physical transport diagrams, and current reject/re-inspection counts. If sufficient data is not available, clearly state this and plan to collect it during PoC. Do not contract for accuracy guarantees first if you have no past images—this leads to mismatched expectations. Do not assume that reference numbers from vendor case studies can be reproduced with your own materials, surfaces, defect frequencies, and imaging conditions; require challenge tests with actual items. With such a request sheet, you can compare different proposals on the same basis.
At the final meeting before mass-production approval, aim to answer five questions in writing. First, who approved which limit sample as the controlled master? Second, how many defect escapes and false rejects of good units were observed for each defect class against that master? Third, which container receives review items, who makes the final decision within what time, and what prevents an unresolved item from shipping? Fourth, can the model, threshold, lighting and part-number recipe versions be reconstructed later for each individual unit? Fifth, when an image is missing or communications fail, which physical units are held and who authorizes recovery? A high average accuracy cannot carry mass-production quality risk if these questions remain unanswered. Once they are answered, even a model with room for improvement may be operated safely within a restricted scope.
Clarify responsibility and cost after acceptance as well. Does maintenance cover extra training to reduce false rejects? Who pays to revalidate the system if a customer changes a limit sample? Who owns the images and decision logs, and for how long are they retained? Do not reuse acceptance results from the first part number without testing when expanding to another. Surface reflection, geometry, conveying posture and defect frequency can change. Evaluate at least the differences, and approve defect classes and boundary samples again. Recording these decisions on a change form lets the factory keep improving while explaining exactly what quality assurance approved.
Map the approval path for exceptions before starting operation. If a critical defect may have escaped, equipment staff should not lower the threshold and keep running on their own. Hold the affected lot and have quality assurance determine the scope. If false rejects suddenly increase, production staff should not narrow the review band on their own either. Separate changes in the images or ground truth from a conveyor identification error, then decide whether to restore the approved old settings or return to manual inspection. Keep the first contact, deputy, night-shift authority and customer-notification rule accessible on paper so physical products remain controlled during a terminal or network failure.
In the monthly review, the key question is whether both escapes and false rejects were reassessed after every change, not how many times the threshold moved. If apparent accuracy improved while there were too few rare defects to test, record the evidence gap. Rank improvement work by customer impact, review labor, loss of good parts and downtime. Only with these records can the team explain what is the same and what is different when expanding to another part number.
Frequently Asked Questions
Can false rejects in AI visual inspection be eliminated by raising the threshold?
With the same model and images, raising the threshold generally reduces the number of good units falsely classified as defective. At the same time, the number of defective units incorrectly classified as good (escapes) may increase. Adjust only after checking the curve by defect class, upper limits for critical defects, and re-inspection lane capacity. If image conditions change and the score distribution itself shifts, do not try to control it with threshold alone—investigate images and equipment status.
How many test images are needed to reduce misjudgments in visual inspection?
There is no universal fixed number. It depends on the defect classes you want to verify, acceptable escape rates, confidence level, and variation in part numbers or lots. The evidential strength of “zero escapes in 20 defective samples” is not the same as “zero escapes in thousands of independent defective samples.” Decide on a statistical plan and ground truth definition in advance, and do not inflate independent counts by burst imaging of the same workpiece.
Can I leave image inspection threshold setting to the vendor?
Vendors can provide candidate values and performance curves. However, acceptable escapes, good unit loss, and re-inspection workload depend on factory and customer conditions. Quality assurance approves limit samples and defect classes, manufacturing confirms processing capacity, and equipment/IT verifies operation and logs—then jointly decide the adopted values. Ensure that authority to change settings and approval history are under factory control in the contract.
Can I use only overall accuracy as the acceptance criterion for AI inspection?
It is not recommended. On lines where good units are the majority, high overall accuracy can hide escapes of critical defects. Combine escapes by defect class, false rejects of good units, “needs verification” rate, actual gate mix-ups, image loss, and recovery conditions. Clearly state denominators and test counts, and separately report FAT image replay results and SAT actual line results.
Summary
Countermeasures for false detections in AI visual inspection are not just about tweaking model scores. They require fixing quality standards by defect class, separately counting false rejects of good units and escapes of defective units, isolating “needs verification” items physically, and managing threshold changes with version control. By evaluating separately with challenge sets and mass production distributions, clarifying cost assumptions, and verifying everything from image judgment to actual gate operation in FAT and SAT, you will be able to explain the trade-off between false rejects and escapes.
If your Thai factory is seeing an increase in false rejects on existing lines, or you are at the stage of creating an acceptance table for new AI inspection equipment, you can consult with us via Contact Us even if you do not yet have a defect ledger or current judgment logs. We can start by organizing the target part numbers, defects you do not want to escape, and available time for re-inspection together.
References
- scikit-learn: Tuning the decision threshold for class prediction — Explains separating scores from actions and how threshold changes affect metrics.
- scikit-learn: Post-tuning the decision threshold for cost-sensitive learning — Example of including costs in the evaluation function. Factory costs in this article are hypothetical.
- NIST AI Risk Management Framework Core — Framework for measuring performance changes from operational data and feeding results back into management.
- ISO 2859-1:2026 — AQL-based attribute sampling for lot acceptance. Distinct from AI image-level judgment criteria.
- ISO 9001:2026 overview — Overview of the 2026 quality management system. Do not claim specific FAT/SAT numbers in this article as ISO requirements.
- ISO 9001 Auditing Practices Group: Monitoring and measuring resources — Audit explanation that effective measurement resources support valid results.