Before procuring annotation for AI visual inspection, a factory needs to settle more than the number of images to label. How will an inspector handle a scratch at the limit sample? Who resolves conflicting judgments? Which version of the dataset receives the ruling? If these questions remain open, a large labeled dataset may simply encode inconsistent inspection practice. This guide helps quality assurance, production engineering and IT teams in Thailand specify a defect taxonomy, adjudication process, golden set, version control, vendor RFP and acceptance test.
For cameras, lighting and the inspection cell as a whole, see our factory AI visual inspection guide. Our high-speed inline inspection guide focuses on image capture and cycle time. Here we focus on the meaning assigned to the captured image. A repeatable decision standard should precede the collection of more training images.
Define the deliverable before starting visual inspection annotation
A label is more than a mark on an image
Annotation attaches a human-defined meaning to an image or an object within it in machine-readable form. For a factory, a file tagged “NG” is rarely sufficient. The record may need to connect the source image, part and process IDs, defect type, position, extent, severity, final disposition, reason, annotator, reviewer and version of the applicable inspection criterion. A customer specification or drawing can change the disposition of an identical image. If the image alone becomes the master record, the reason for that change disappears.
Separate three questions: Is a mark visible? What kind of mark is it? May the part be shipped? A faint scuff can be detected and classified while still falling within an approved limit sample. Combining observed defect and shipment disposition in one label makes it impossible to tell whether the model learned the physical phenomenon or the current factory acceptance rule. Put observation, severity and disposition in separate fields in both the annotation schema and RFP.
Define the unit of work. One photograph does not always equal one product. A part can have front and rear views, a frame can contain several cavities, and a decision can apply to a lot. Distinguish image ID, workpiece ID, view, capture time and lot ID. State what counts as one annotation task. These definitions determine the denominator for prices, double reviews and missing-record checks. Settling them after production has started invites disputes about delivery volume.
Choose classification, boxes or segmentation for the decision
Image-level classification records whether a frame is acceptable but does not locate the defect. A bounding box shows location, though it may include a lot of background for a thin scratch or irregular contamination. Pixel-level segmentation supports area and boundary measurements at greater annotation cost and with more subjective edge decisions. The choice should follow the acceptance metric. A shipment decision may primarily require image-level disposition; comparison with a length limit may require a position and contour.
Separate defects in the product from failures of image acquisition. Blur, inadequate exposure, misplacement and lens contamination should not silently become either “good” or “scratch.” Use states such as UNJUDGEABLE and RECAPTURE. Putting every poor image in the NG bucket distorts defect reporting; putting it in the good bucket hides a camera problem. Specify quarantine and manual review when the equipment cannot recapture the part.

Build a defect taxonomy from the inspection standard
Reconcile names before creating classes
One linear mark may be called a scuff, scratch or streak on the shop floor and “scratch” in a customer report. Quality assurance should compare drawings, customer requirements, inspection sheets, physical rejects and limit samples to create a term dictionary. Check whether one term covers several phenomena or several terms describe the same phenomenon. The goal is to preserve local vocabulary while mapping synonyms to a stable decision code.
Keep the hierarchy as simple as the job permits. A first level might distinguish linear marks, spots, chips, contamination, missing components and capture failures. Add a second level only where material or mechanism changes the required action. Asking people to distinguish dozens of categories at once usually increases disagreement. “Other” needs a required explanation and a periodic review; otherwise the taxonomy’s blind spots remain invisible. Any threshold for creating or merging classes belongs to the factory’s risk assessment, not a generic blog benchmark.
OK, NG and REVIEW are decision states, not defect types. A visible mark can have a defect type while its disposition is OK. An unresolved limit case goes to REVIEW. Do not use one REVIEW reason for both customer approval and recapture. Distinct reason codes let the team see whether the model is uncertain, the rule is unsettled or the input image is unusable.
Preserve boundary cases as text and images
The annotation guide should say what to label and what to exclude. If a molded part’s gate mark is permissible, show how to distinguish it from contamination or a chip, with allowed locations and appearances. A reflection on a glossy part can resemble a scratch from one camera angle. If one image cannot settle it, define a recapture or human decision route. Pair each limit-sample photograph with a magnified region, lighting condition, part number and reason for its disposition.
Document multi-defect cases too. May boxes overlap? Is only the principal defect recorded? When do two nearby marks become one region? Should a cast shadow be included at the boundary? These rules look less dramatic than model accuracy, but they make consistent annotation possible across shifts and vendors. Deliver the instruction, positive examples, counterexamples and revision history as a versioned unit.
In a Thai plant using Japanese, Thai and English, keep label codes independent of language. Display names may be translated, while IDs, definitions and example images remain shared. A translated word for “stain” may not have precisely the same scope. Have inspectors using both languages judge the same examples and send disagreements back to the definition. The final ruling should belong to the quality owner, not to the translator alone.
Measure disagreement and set an adjudication route
Do not rely on one overall agreement rate
In an initial pilot batch, ask at least two people to annotate the same images independently, without seeing each other’s answers. Inspect results by defect class, part number, lighting and boundary case. An overall agreement figure can look strong because easy good parts dominate while critical defects remain disputed. Record disposition agreement, class agreement, spatial overlap and unjudgeable rate separately. Agree in advance how bounding boxes or regions will be compared.
Do not settle every disagreement with a majority vote. If two reviewers say OK while a third identifies a chip prohibited by the customer specification, quality assurance must examine the part, rule and image. CVAT’s consensus annotation documentation explains how multiple annotations can be merged and how low-consensus results can enter review. The tool’s vote cannot replace accountable quality judgment.
Keep an adjudication log with the original decisions, final decision, decider, criterion version, evidence and need for a guide revision. Repeated disagreement may indicate a missing definition, inadequate limit sample, unstable image capture or an unresolved product requirement. After changing the guide, decide how much previously labeled material must be re-reviewed. Applying the new rule only to future images can leave the same defect coded differently across periods.
Name review authority and stop conditions
Separate annotator, first reviewer and final adjudicator. A vendor’s internal approval alone cannot establish conformity to the factory’s inspection criterion. Factory QA should own high-risk and boundary decisions; production engineering should investigate capture-related questions; IT should preserve data versions and audit logs. Name deputies and a holding area for images awaiting a ruling.
Set stop conditions before volume work. A new defect type, conflicting interpretations between departments or a changed capture setup can justify pausing a batch and updating the guide. The appropriate threshold depends on the product and inspection risk, so put the decision rule in the RFP for both sides to agree. “Fix it later” without a bound makes the scope of rework unknowable.

Maintain a golden set as the evaluation yardstick
Include boundary and acquisition cases
A golden set is a collection of images decided with evidence by authorized reviewers. It is more than labels supplied by one inspector. Include clear good and bad examples, borderline parts, visually similar defects, bad captures and relevant variation across part number, lot, shift and lighting. Scarce defects may have too few real examples; show that gap instead of filling it silently with synthetic images. Our manufacturing synthetic-data guide discusses where generated images can help and where they still need validation.
CVAT’s quality-control guide describes ground-truth jobs and validation sets used to compare annotations. The operational point is to keep creation and approval of reference answers separate from routine labeling. Record who approved them and when. Quietly changing the golden set mid-project makes earlier and later batch scores incomparable. If updates are needed, create a new version, document the reason and retain the old evaluation result.
Keep evaluation examples separate from training data. Near-duplicate images of the same workpiece in train and acceptance sets can overstate performance. Decide the split unit using lot, workpiece, capture date, camera or part number as appropriate. State which products and conditions are outside the test; do not generalize a pass result beyond them. If duplicates or unusable images are discovered, apply a documented exclusion rule and preserve counts.
Preserve uncertainty in the reference answer
Do not force every image into OK or NG. “Requires physical inspection,” “awaiting customer approval” and “cannot be decided from image” can remain separate states. Distinguish the frozen, approved acceptance set from a holding set where the rule is being developed. Retain the latter as input to a future limit-sample revision or capture improvement.
Also separate model evaluation from annotation quality evaluation. A mismatch with the golden set may arise from the model, camera, label or rule. When a reference answer is challenged, do not rewrite it to match a model prediction. Send it to an independent quality review. The NIST AI RMF Measure playbook links reliable evaluation to data quality and representativeness. For a factory, that principle becomes a record of operating conditions and decision authority.
Version images, labels and criteria together
One “dataset v2” label is insufficient
Version at least the raw images, preprocessing, taxonomy, instructions, individual labels, golden set and train/validation/acceptance split. A model version alone cannot reproduce an experiment if the training labels were later corrected. MLflow Dataset Tracking describes lineage between input datasets and experiments. Even without adopting that particular tool, require a fixed record of which image set, label revision and rule version produced each result.
Do not overwrite source images. Store a stable image ID, checksum, capture conditions and import date. For brightness correction, cropping or rotation, retain a link to the raw image and the processing version. Avoid making the filename the only link to the part: a naming change can break the association. Use separate stable IDs for image and workpiece, independent of where files are stored.
For label edits, retain before and after values, reason, operator, approver and time. Flips between OK and NG deserve particular attention because they can affect past model training and shipment decisions. When a customer criterion changes, preserve the rule in force at the time of the original decision. Record a reassessment under the new rule as another version; QA determines whether it applies retrospectively and to which lots.
Verify export and portability in the pilot
Seeing labels in a vendor’s browser is not a deliverable. Specify export format for images and annotations, taxonomy, attributes, coordinate system, orientation, ID mapping and review history. CVAT’s dataset export guide shows that format and image inclusion depend on tool and plan. During the pilot, export real data, import it into the factory environment and check that the decisions survive the round trip.
If different language versions of the guide are maintained in separate files, give them a shared version identifier and translation approval date. A revised Japanese guide and an old Thai guide must not be used in the same batch. The version visible to vendor annotators is part of the audit record. Put the same identifier in both data delivery and commercial documents so a later dispute can be reconstructed.
What an annotation vendor RFP should contain
Make deliverables and responsibilities explicit
Start with product and view, expected defects, image format, owner of the acceptance criterion, work location, confidentiality and delivery unit. Responsibility changes depending on whether the plant supplies selected images or the vendor also handles capture. A vendor cannot create perfect traceability after receiving images without workpiece IDs. List missing input fields and assign responsibility for resolving them.
Attach the label schema: stable class IDs, definitions, included and excluded examples, rules for position and area, multiple defects, OK/NG/REVIEW, reasons for unusable images and references to limit samples. Distinguish issues that may be resolved after kickoff from those affecting price and acceptance that must be fixed before award. If the factory criterion is still unsettled, procure a small guide-design phase before quoting production labeling.
For quality control, specify how double reviews are sampled, independent judgment, review roles, adjudicator, ownership of the golden set and rework conditions. A single vendor claim such as “99% accuracy” means little without its denominator, good-part prevalence, treatment of critical defects and spatial scoring rule. Ask for separate prices for guide creation, pilot, production labeling, review, adjudication, correction and export. This exposes rework hidden behind a per-image price.
Security and exit conditions belong in the contract as well. Images may expose a customer drawing or unreleased product. Agree on storage location, access, subcontracting, secondary use for model training and evidence of deletion. Applicable law and customer contract terms must be assessed for the specific project; a universal retention period cannot be assumed here. Specify how the factory receives the raw images, labels, guides and history when changing tools or suppliers.
Compare bids using the same pilot set
Give each bidder the same images and provisional instructions. A test of only easy good parts will hide quality differences. Include borderlines, multiple defects, poor captures and part-number variation, and allow the supplier to answer “cannot determine” with reasons. The quality of its questions can be more informative than a forced label on every image. Give all bidders the same answer to questions that arise during the test.
Use their disagreements to repair your own guide before negotiating price. Preserve the first and resubmitted batch as separate versions. Compare not only price but boundary judgments, audit trail, rework turnaround, export, multilingual operations, storage controls and support response. Weight these criteria according to the factory’s critical defects and contract needs.

Acceptance should test reproducibility as well as count
Inspect quantity, quality and transfer separately
The required number of records is not enough if images are duplicated, IDs missing or labels outside the schema. First check file counts, unique image IDs, workpiece mapping, corrupt files, out-of-scope images and required-field completeness automatically. Next score a frozen golden set and an independent sample of reviewed images. Finally export and re-import the dataset with guide version and adjudication history to see whether the factory can reproduce the same decisions after handover.
Set metrics for the actual risk: disposition errors, misses of critical defects, class confusion, location deviation and appropriate referral to REVIEW. A high average agreement figure does not automatically make one missed critical defect acceptable. In another task, correct defect presence may matter more than a small box displacement. Agree on metrics, thresholds, retest rules, sampling unit and decision owner before production. Do not move the target after seeing delivery results.
If the supplier challenges a reference answer, have an independent quality owner inspect the physical part, criterion and limit sample. If that answer changes, record the version, affected scoring range and recalculation. When test images are later released for supplier training, reserve a different acceptance set next time. Repeating the same known test does not measure quality on new batches.
Monitor label quality after go-live
Acceptance is not the end. A product revision, mold replacement, lighting or lens change, new defect or changed customer limit sample can invalidate the guide’s scope. Collect operator corrections to model results with reasons, but do not automatically treat them as new ground truth; send them through QA adjudication. Concentration of corrections in a part number, shift or camera calls for investigation of both imaging and the guide.
In a periodic review, inspect REVIEW reasons, unresolved cases, disagreement by class, rework time and guide revisions. Choose measures that serve the line rather than tracking every possible number. On an anomaly, identify affected dataset and model versions, then evaluate the lots and decisions that may need recheck. This trace between image, label, model and shipment decision turns visual inspection into a maintainable quality process.
Use a small trial batch to expose operating gaps before procurement
Separate images from decision information before labeling
A trial batch for a prospective vendor should not be an arbitrary extract of stored images. Record product variant, camera position, lighting condition, surface material and production time, then select cases across these conditions. Include clear good parts, clear defects and borderline cases, as well as blur, glare and misalignment that make an image unsuitable for judgment. Forcing an unusable image into either the good or defective class hides an acquisition problem inside the training data. Give recapture cases a separate reason code and name the person responsible for obtaining a usable image.
Also prevent the answer from being inferred from filenames or production history. If defective parts come from a separate folder, an annotator or tool could classify them without inspecting the defect. Give labelers only the information they would have in production, while keeping the additional information needed by quality assurance for acceptance. Record which evidence supported each decision. Preserve the part identifier and capture time for traceability, but keep fields intended to be blinded out of the labeling interface.
Turn disagreement into a specific rework instruction
When two reviewers assign different defect names to the same image, distinguish a data-entry mistake from a missing taxonomy rule, poor image acquisition or ambiguity in the inspection standard. Returning every disagreement as an annotator training failure leaves the customer’s acceptance rule unresolved. Quality assurance should inspect the physical part or obtain additional images and record the applicable product specification and clause in the decision log. If the specification still cannot decide the case, do not certify a provisional label as training truth. Hold it for agreement among design, quality and the customer-facing owner.
A rework instruction needs the image ID, old and new labels, reason for change, taxonomy version and the scope of similar images to recheck. Correcting one image alone can leave earlier batches inconsistent under the same rule. The authorized owner must choose whether a revised rule applies retrospectively or from the next batch and record that boundary. Evaluate a vendor on its ability to discover, explain, correct and resubmit disagreements, as well as on labels that agreed at first pass.
Test whether the factory can take over after acceptance
Final acceptance should go beyond a demonstration on a vendor’s own screen. A factory employee should open the delivered images, annotations, taxonomy and decision guide in another environment and reproduce defect classes and locations for selected samples. Check the mapping between image and label IDs, the coordinate origin and units, and the retained reasons for excluded images. Confirm that the factory can still read required data and audit logs after the vendor’s accounts are disabled.
New scratch patterns, lighting replacement and product specification changes will arise after go-live. Before retraining only the model, decide whether the taxonomy or golden set also needs revision. If so, preserve results under the old version and rerun evaluation under the new one before switching. The factory should be able to explain why this month’s metric differs from last month’s. Including this handover in the purchase order makes operating effort visible alongside the price per labeled image.
Frequently asked questions
How many images do we need to start annotation for AI visual inspection?
There is no universal minimum. List defect classes, part numbers, capture conditions and acceptance boundaries first. Use a small pilot to find ambiguous instructions, then plan volume based on the real examples available in each condition. Adding thousands of easy good images will not resolve a missing definition for a rare critical defect. The RFP should specify an initial batch and a process for deciding the next tranche.
Can majority vote resolve conflicting defect labels?
It can assist on defined, low-risk cases. Boundary decisions, customer requirements, critical defects and unusable images should go to an authorized QA adjudicator. CVAT and Label Studio’s inter-annotator agreement tutorial illustrate how low-agreement work can be routed to review. Preserve the reason and rule version, not merely the final vote.
How does a golden set differ from ordinary training data?
Training data teaches the model. A golden set consists of evidence-backed answers approved by the responsible team and used to assess annotation or model outputs. Keep the sets separated to avoid over-optimistic evaluation. Reference answers may be corrected, but each correction needs a new version, reason and assessment of its effect on prior scores.
What determines annotation outsourcing cost?
Cost depends on more than image count: classes, spatial precision, borderline share, double review, expert adjudication, multilingual instructions, rework, delivery format and security requirements all matter. Separate guide creation and pilot from production pricing. Ask about REVIEW and specification-change charges. Compare suppliers on the same pilot images instead of treating a published generic price as a quote for your plant.
Conclusion
The first asset to procure is a repeatable method for applying the plant’s acceptance criterion, rather than a large pile of labels. Separate observations from shipment decisions, document boundaries, resolve independent judgments, freeze the golden set and version images, labels, criteria and evaluation splits. Translate those decisions into the RFP and acceptance test. Then the factory retains the reasoning even when it changes vendors.
If your Thailand plant is defining an image-inspection annotation brief or pilot batch, contact TOMAS TECH with your part numbers, defect examples and current limit samples. The taxonomy and acceptance criteria can be reviewed before vendor award.