If an RFP for robot learning data collection specifies only hours of video or numbers of frames, the buyer may receive data that works in a demonstration but cannot reproduce a capability on the factory floor. What a factory needs to procure is a reproducible mission, independent external ground truth, episodes that retain failures, an untouched holdout, and an operating loop that reaches controlled retraining. This guide translates those requirements into an RFP, a 90-day PoC and acceptance evidence for factories in Thailand and ASEAN.
“Collect more data” is not a procurement specification
Video hours, frame counts and episode counts are convenient quotation units. They are not evidence of capability. One hour of repeated success with the same workpiece and fixed lighting is not equivalent to one hour that covers pose, grasp point, reflection, friction, delay, stop and recovery. A large dataset can also be unusable when camera frames, joint states and commands do not share a reliable clock: perception and action may appear in the wrong causal order.
*The Robot Data Factory*, submitted to arXiv on 15 September 2026, argues that the critical resource for physical AI is not raw robot data alone but robot experience that preserves observations, actions, embodiment, context and outcomes. It proposes reproducible missions, skill curricula, synchronized multimodal sensing, external ground truth, data pipelines and living benchmarks. The paper is a preprint and had not been peer reviewed when this article was prepared. Buyers should use it as a useful design hypothesis for managing experience as a production process, not as a final standard or certification.
An RFP should therefore start with six questions rather than an hour target:
- What business mission runs from which start state to which end state?
- Who decides success, failure, interruption and unrecoverable outcome, using which measurement?
- Which measure is sufficiently independent of the robot’s own estimate?
- Which variations may appear in training, and which remain hidden in holdout?
- How are robot, software, model, tool, fixture, workpiece and sensor versions traced?
- When a new production failure appears, who classifies it, approves learning and authorizes redeployment?
Without these answers, the deliverable is a pile of files rather than a learning asset that can reproduce an accepted capability.
Translate the Robot Data Factory hierarchy into contract deliverables
The paper’s mission–task–skill–episode–dataset–benchmark–capability hierarchy is also a practical procurement breakdown. Skipping directly to “model accuracy” or “success rate” lets each vendor choose a different denominator and failure policy.
| Level | Factory decision | Acceptance evidence |
|---|---|---|
| Mission | Production objective, start/end and forbidden states | Mission contract, process map, video |
| Task | Transfer, grasp, position, load, inspect and related steps | Task state machine and I/O list |
| Skill | Reusable reach, grasp, place or recovery behavior | Skill version, preconditions, stop conditions |
| Episode | Observations, actions and outcome of one attempt | Episode manifest, synchronized logs, video |
| Dataset | Curated episodes and approved splits | Dataset card, lineage and rights schedule |
| Benchmark | Fixed conditions and evaluation procedure | Test protocol, raw results, rerun instructions |
| Capability | Repeatable business ability within a stated envelope | Holdout/accept results, constraints, open issues |
“Place a box on a shelf” is not yet a mission. It becomes one after the contract defines box dimensions, mass, center of gravity, surface, shelf height, allowed pose, graspable faces, surrounding equipment, starting location, obstacles, replacement attempts, timeout and recovery after a safety stop. Success is not the robot emitting a completion flag. External evidence should show that the box is inside tolerance and that equipment I/O and inventory state are consistent.
The hierarchy also reduces lock-in. Require the mission contract, episode schema, dataset version, evaluation protocol and failure taxonomy—not only trained weights. These artifacts preserve comparison when the factory changes a model or robot. If the success rule exists only inside a supplier interface and raw episodes or decision evidence cannot be exported, the factory does not own its improvement loop.
Freeze reproducibility and prohibited states in a mission contract
A mission contract is more than a prose use-case description. It defines the boundary conditions needed to repeat the experiment.
| Area | Fields to freeze or record |
|---|---|
| Robot | Model, serial, controller, firmware, tool, TCP and payload settings |
| Work | Part number, instance, lot, dimensions, mass, surface, tolerance, defect state |
| Fixture | Drawing revision, coordinate, fastening, wear, datum and replacement history |
| Sensor | Model, serial, sample rate, exposure, range, calibration and clock source |
| Environment | Lighting, background, floor, relevant temperature/humidity and network |
| Human | Role, intervention envelope, access conditions and teleoperator qualification |
| Control | Software/model/prompt/policy version, random seed and safety configuration |
| Outcome | Success, business rejection, safety stop, technical failure, recovery, unknown |
Separate controlled variation from fixed conditions. Training may deliberately vary pose or lighting, while the holdout plan reserves particular combinations as unseen. Selecting only attractive episodes after collection inflates the reported score and hides operational exposure.
Every mission-contract change should have a version, reason, affected datasets, required reevaluation and approver. A fixture moved by a millimetre, a replacement camera, a new gripper-pad material or a changed PLC handshake may alter the distribution. Recording the revision makes its effect measurable instead of turning it into an unexplained model regression.
Make ground truth independent and synchronize the sensors
If a model reports “grasp succeeded” and its own confidence becomes the label, the project counts the same error twice. External ground truth must be sufficiently independent for the acceptance question. Placement may use fixture sensing or an external camera; insertion may combine force, displacement, PLC completion and downstream inspection; logistics may combine weight, barcode and WMS state.
Independent does not necessarily mean expensive. The measurement must define its target, resolution, calibration, uncertainty, blind spots and missing-data treatment, and it must not rely only on the robot’s declaration. Ground truth itself can be wrong. Keep ambiguous episodes as unknown instead of forcing them into a success/failure binary.
Camera, depth, force/torque, joints, commands, PLC, safety events and operator inputs need a shared time reference. Do not accept “synchronized” as a one-word requirement. Specify clock source, permitted offset, drift check, drop detection and interpolation rules. If action and consequence are reversed by timestamps, the model can learn an effect as though it were a cause. Missing samples should remain visible through quality flags and reasons, not concealed with zeros.

An episode must preserve causality, failure and recovery
An episode manifest should join the precondition, observations, command, executed action, state transition, intervention and result. A video filename is insufficient. At minimum, require:
- mission/task/skill identifiers and versions;
- episode ID, site, cell, robot, tool, fixture and workpiece instance/lot;
- every sensor stream ID, timestamp, calibration reference, drop and saturation flag;
- command, executed action, controller mode and autonomy/assistance level;
- teleoperator, device, latency, intervention start/end and reason;
- software, model, configuration, dataset and simulator versions where applicable;
- outcome, failure class, near miss, safety stop, recovery and retry relationship;
- ground-truth measure, adjudicator, adjudication time and reason for unknown;
- consent, privacy, retention, reuse scope and export permission.
Failure taxonomy should include perception, object identification, grasp, contact, path, equipment handshake, timeout, human intervention, safety stop and business-state inconsistency—not merely software exceptions. Preserve context before a failure and the recovery action. Success-only episodes do not teach the robot when to stop, ask for assistance or back away and regrasp.
Do not create hazards merely to collect failures. Design safe fault injection and boundary cases within the approved hazard analysis, speed/force/space limits, safety functions and stop procedures. A near miss here means recording precursors such as an early stop, low confidence, unexpected contact or manual takeover; it does not mean staging an almost-accident.
Lineage should trace raw data through derived labels, filtered datasets, training runs, models and deployment. If a labeling defect is found later, the factory must be able to identify the affected models. Include checksums, immutable raw records, change history, dataset cards and propagation of deletion requests.
Teleoperation training data: useful, but not self-explanatory
Teleoperation training data can capture expert demonstrations, recovery moves and intent labels. It is misleading when operator differences or network latency disappear from the metadata. Fully manual control, shared control, human approval of an autonomous proposal and intervention only on failure are different kinds of experience. Record the control mode and assistance level, and preserve both the policy proposal and the human-executed action.
Operator identity is needed to detect bias, not to score an employee. Training on one expert may encode that person’s speed, viewpoint and habits. Manage qualification, training date, device, relevant handedness, session length and fatigue controls, then test reproducibility across operators. When video or audio contains workers, badges, screens or customer information, contract the purpose, access, retention, masking, cross-border transfer and secondary use.
Collect pause, abort, regrasp, undo, takeover and “do not act,” not only successful demonstrations. Preserve a pre-intervention circular buffer so analysts can see when instability began. Store episode-level latency and missing measurements rather than only an average; otherwise a network artifact can be mislabeled as a robot-skill limitation.
The delivery should include raw control streams, transformed trajectories, transformation code, coordinate frames, filters, resampling and exclusion rules. A supplier that provides only smoothed paths may remove contact, hesitation and intervention boundaries that carry important learning information.
When synthetic data for robots helps—and when it does not
As discussed in our guide to synthetic data for manufacturing AI, simulation can expand rare poses, lighting, occlusion, background and sensor perturbations. “Cheaper than real data” and “unlimited” are not acceptance arguments. Simulation offers designed coverage, but it can reproduce a simulator defect at industrial scale.
Contact, friction, slip, deformable materials, cables, wear, backlash, latency, human response and safety behavior require validation against physical measurement. Preserve simulator name/version, asset origin, physics parameters, sensor model, domain-randomization range, seeds, render settings and generation code. The delivery must state which physical measurements calibrated parameters and which holdout assessed the sim-to-real gap.
Instead of selecting a synthetic percentage first, define its role for each failure class: supplement occlusions that are difficult to repeat safely, create visibility labels, or initialize policy exploration. Final acceptance should return to a physical-cell holdout where the business decision demands real evidence.
NIST’s Physical AI and Data Generation for Robotics project highlights the need for metrics and test methods that consider the relationship among the algorithm, robot system and task, and reports work on a mixed physical/simulated testbed. This supports task-specific hybrid evaluation; it does not say that simulation replaces the factory.
Contract the DEPLOY–MEASURE–LEARN–REPEAT loop
A central Robot Data Factory idea is that a dataset is not a one-time endpoint. Production experience feeds the next controlled improvement. Put a gate between each stage:
- DEPLOY: Release only an approved model, configuration and safety setting to a defined cell.
- MEASURE: Use external criteria to measure denominators, success, failure, intervention, unknown, drift and cycle time by mission.
- LEARN: Classify failures and approve whether the response is new data, corrected labels, a skill change or a mechanical/process change.
- REPEAT: Redeploy only after regression, holdout and acceptance, retaining rollback to the approved prior version.

Do not feed every production log automatically into training. Human-rescued episodes, equipment faults, incorrect ground truth, personal data and out-of-contract work can be mixed together. Operate a retraining candidate queue in which data, process, safety and AI owners approve use. The question is not whether the new dataset is larger, but which named failure it is intended to reduce.
The model registry should link training dataset, code, parameters, evaluation, limitations, approval, deployed cells and rollback. Minimum operating metrics are mission-specific attempts, externally verified completion, failure class, manual intervention, safety stop, unknown outcome, missing data and distribution shift. A global success rate can improve merely because the workload contains more easy parts.
Compare vendors on reproducibility, not collection volume
Normalize quotations by deliverables, evidence, rights and change terms—not robot-hour price alone.
| RFP topic | Required answer | Acceptance evidence |
|---|---|---|
| Mission | Start/end, variation, forbidden states, recovery | Versioned mission contract |
| Instrumentation | Sensors, calibration, clocks, missing data | Calibration records and sync test |
| Episode | Schema, failures, intervention, lineage | Manifest and sample replay |
| Ground truth | Independence, resolution, uncertainty, unknown | Measurement study and adjudication procedure |
| Teleoperation | Operator, mode, latency and rights | Session log and raw/control transformation |
| Synthetic data | Simulator, assets, parameters and validation | Simulation card and real-gap report |
| Split | TRAIN/HOLDOUT/ACCEPT isolation | Hashes, access log and freeze record |
| Evaluation | Denominator, classes, threshold and repeats | Rerunnable test harness |
| Safety | Risk, limits, stops and change control | Risk assessment and test record |
| Operations | Monitoring, retraining, approval and rollback | Runbook, registry and RACI |
| Rights | Ownership, license, retention and export | Contract plus deletion/export test |
Separate pricing for instrumentation, mission/episode design, collection operations, annotation/QA, simulation, training, test harness, safety integration, site trials and operating handover. “AI development—one lot” hides the boundary between adding data and correcting software. Specify change rates for mission revision, sensor replacement, relabeling and reevaluation, not only additional hours.
A polished vendor video is an entry point, not evidence. Ask whether the protocol can be rerun with another fixture, workpiece instance, operator and shift; whether failures remain visible; and whether the chain from raw record to reported result is exportable. A 2026 FANUC America announcement about robotics, automation and physical AI is useful market context, but it is not independent proof for a buyer’s mission.
In an announcement dated 19 August 2026, Hexagon Robotics and Schaeffler described a Humanoid Gym using Train–Validate–Deploy, imitation learning and repeated execution on representative manufacturing applications. They said the initiative supports a planned rollout of at least 1,000 AEON humanoids over coming years. This is the companies’ plan, not 1,000 completed deployments. The transferable lesson is the separation of training and validation before production release, not the headline number.
Agility Robotics stated on 15 September 2026 that Digit had accumulated more than 65,000 hours of real-world operation. It also stated that Digit 5 can repeatedly lift up to 50 lb (22.7 kg) and uses a 90-minute runtime battery that charges in 9 minutes. These are vendor-published figures. They do not establish results with a different payload, floor, workflow or safety configuration. Convert catalogue claims into acceptance conditions for the buyer’s mission, tool, grasp pose, stop frequency, handoff and charging plan.
Separate safety scope, Thai editions and incentive scope
ISO 10218-2:2025 addresses the design, integration, commissioning, operation, maintenance and decommissioning of industrial robot applications and robot cells. It is not a robot-learning-data standard, nor should it be represented as a comprehensive standard for humanoids in general. Its stated scope leaves areas such as some mobile-platform integrations outside coverage. Determine applicable standards and additional risk assessment for the specific robot, mobility, human proximity, tool and process.
Thailand’s TISI มอก. 3950 เล่ม 2-2567 came into force on 8 February 2025, but it is an identical adoption of ISO 10218-2:2011. It is not the same edition as ISO 10218-2:2025. A contract should name the standard number, edition, scope, whether it is a legal or customer requirement, and who resolves edition differences. Referring technically to the newer international edition does not automatically replace the requirements applicable in Thailand. Confirm the project with local safety and legal specialists.
Thailand BOI’s 2026–2027 page discusses investment in automation, robots, AI/ML and Big Data under a measure for the automotive industry. Do not present it as an entitlement for all manufacturing. Confirm industry eligibility, equipment/software scope, timing and application conditions project by project. The technical case should remain viable without assuming an incentive.
Design separate TRAIN, HOLDOUT and ACCEPT zones
Randomly splitting frames lets neighboring frames from nearly identical episodes leak into training and holdout. Split on the variation that must generalize: workpiece instance and lot, fixture, operator, shift, site, camera, software version or collection period.
- TRAIN: Developers may inspect it and use it for fitting, tuning, failure analysis and additional collection.
- HOLDOUT: Hidden during tuning and evaluated a limited number of times with a frozen protocol. Any adjustment after inspection creates a new version.
- ACCEPT: The buyer runs the approved physical-cell conditions, workpieces, fixtures, shifts and evidence procedure for the release decision.

Do not make overall success rate the only threshold. Fix denominators by mission and report critical failures separately. Distinguish manual intervention, safety stop, unknown, cycle timeout, wrong object, wrong placement, inconsistent equipment state and successful recovery. Repeat across the approved variation envelope and expose the worst class, not just the average.
The acceptance evidence pack should contain protocol version, dataset hashes, model/configuration, risk controls, preconditions, every trial, raw results, failure videos, ground-truth values, exclusions, missing data, operator, timestamps, retests and sign-off. A smooth demonstration video is not reproducible evidence. Deliver the test harness and aggregation code and verify that the factory can regenerate the same tables from the same inputs.
As in our guide to AI PoC exit criteria, the decision should be deploy, extend with a defined question, or stop. An average may pass while one critical failure remains; deployment can then be restricted to approved products or cells. If the root cause is sensing, tooling, fixture design or process logic rather than data scarcity, correct the physical system instead of buying more training.
An example 90-day PoC
The following is a TOMAS TECH planning example, not a universal industry timetable. Adjust it for equipment lead time, safety review, product mix and shifts.
| Period | Main work | Exit evidence |
|---|---|---|
| Days 1–15 | Mission, risk boundary, rights, sensor/clock/calibration plan, acceptance protocol and holdout governance | Approved mission contract and test plan |
| Days 16–30 | Instrument one training cell; dry cycles; schema, lineage, ground truth and replay | Repeatable episode replay and adjudication |
| Days 31–60 | Collect normal, failure, intervention and recovery; validated synthetic additions; freeze dataset | Dataset card, failure coverage and frozen version |
| Days 61–75 | Train/tune on TRAIN; formal HOLDOUT with no adjustment; classify every failure | Holdout report and improve/stop decision |
| Days 76–90 | ACCEPT on approved fixtures, workpieces and shifts; safety/operations review and handover | Deploy/extend/stop decision, runbook and registry |
Days 1–15 fix who owns success and failure before collection accelerates. Agree on denominator and categories so that safety stops or human intervention cannot be removed later simply to improve the score. Define who holds the holdout and when it can be opened.
Days 16–30 instrument one cell and one mission. Dry cycles expose clock offsets, missing signals, calibration references and replay defects. Scaling before one episode can be replayed only scales the defect. If the mission includes handoffs to mobile robots, conveyors or software, use the boundary approach in our mobile manipulator, cobot and AMR selection guide and include equipment I/O plus MES/WMS state.
Days 31–60 follow an approved variation matrix and collect failures, pauses, aborts and recoveries alongside normal runs. Any synthetic addition must name the coverage gap and its physical validation. Freeze the dataset at day 60; later additions belong to the next version.
Days 61–75 tune only with TRAIN and execute HOLDOUT formally. If the team changes parameters after seeing the result, that holdout is opened and cannot be called unseen evidence again. Classify each failure as data, model, sensor, tool, fixture or business-rule related.
Days 76–90 let the factory lead ACCEPT in representative conditions. After a pass, operators should demonstrate monitoring, retraining-candidate approval, emergency response, rollback, supplier escalation and data deletion/export. The PoC deliverable is not only a model; it is a safe, repeatable improvement process.
Frequently asked questions
Is a Robot Data Factory a data center?
Not in the sense used here. The unreviewed preprint submitted on 15 September 2026 describes infrastructure and methodology for continuously generating, validating and reusing robot experience. Mission, sensing, external ground truth, pipelines, benchmarks and DEPLOY–MEASURE–LEARN–REPEAT are the central elements, not storage alone.
How many hours of robot data collection are enough?
There is no universal number. Need depends on missions, workpiece variation, failure frequency, sensor rate, policy and required confidence. Define a coverage matrix and holdout protocol, then ask which uncertainty or failure class each increment reduces. Hours can cap commercial scope, but capability evidence should determine acceptance.
Should teleoperation training data contain only successful demonstrations?
No. Safely include pauses, aborts, regrasps, intervention, recovery and “do not act.” Retain operator, control mode, assistance level, device, latency, coordinate transformation and filtering. Do not confuse one expert’s style or a network delay with robot capability.
Can synthetic data for robots replace physical data?
It can expand rare geometry, lighting, occlusion and sensor perturbation. Contact, slip, deformable material, wear, latency, human response and safety behavior require physical validation. Version the simulator, assets, parameters, randomization and seeds, then return to physical holdout and acceptance.
What is the most important separation in robot PoC acceptance?
TRAIN, HOLDOUT and ACCEPT. Prevent leakage by object, fixture, operator, shift, site, version and time—not merely random frames. If developers inspect a holdout and tune against it, treat it as opened and create a new evaluation version.
Does 65,000 hours of field operation remove the need for a local PoC?
No. The figure is an Agility Robotics statement about Digit’s real-world operation. It does not guarantee a result for the buyer’s workpiece, payload, tool, floor, handoff, safety configuration or cycle target. Use it in supplier qualification, then conduct mission-specific holdout and acceptance.
Does ISO 10218-2:2025 cover all humanoid safety?
No. It focuses on integration of industrial robot applications and cells and has scope limitations. Identify additional standards and risk assessment for humanoids, mobile platforms, human proximity, tools and process hazards. Thailand’s TISI 3950 Part 2-2567 is an identical adoption of ISO 10218-2:2011, not the 2025 edition.
Summary
Do not procure robot learning data collection as hours or frames. Procure reproducible missions; synchronized sensors and independent ground truth; episodes containing success, failure, intervention and recovery; traceable lineage; bounded roles for teleoperation and synthetic data; leak-resistant TRAIN/HOLDOUT/ACCEPT splits; and an operating DEPLOY–MEASURE–LEARN–REPEAT contract. Structured this way, an RFP and PoC move acceptance from “the demo worked” to “the factory can regenerate the evidence and improve safely.”
When you are defining the collection scope, RFP, 90-day PoC or holdout acceptance for a robot-learning project, talk with TOMAS TECH. We can structure the plan around your mission, installed equipment, safety boundary and available logs—starting from reproducible acceptance evidence rather than a target data volume.
References
- Haddadin et al., *The Robot Data Factory*, arXiv:2609.16705 (submitted 15 September 2026; unreviewed preprint)
https://arxiv.org/abs/2609.16705
- NIST, *Physical AI and Data Generation for Robotics*
https://www.nist.gov/programs-projects/physical-ai-and-data-generation-robotics
- Hexagon, *Towards factory deployment: How AEON is trained to perform* (19 August 2026; company announcement)
- Agility Robotics, *Agility Unveils Digit 5 Humanoid Robot Built for Cooperatively Safe Work at Scale* (15 September 2026; company announcement)
- NVIDIA Developer, *Isaac GR00T*
https://developer.nvidia.com/isaac/gr00t
- ISO, *ISO 10218-2:2025 — Robotics — Safety requirements — Part 2: Industrial robot applications and robot cells*
https://www.iso.org/standard/73934.html
- TISI, *มอก. 3950 เล่ม 2-2567*
https://a.tisi.go.th/t/?n=8107
- FANUC America, *FANUC America Brings Robotics, Automation, Physical AI and CNC Innovation to IMTS 2026*
- Thailand BOI, automation/robotics/AI measure for the automotive industry, 2026–2027