An OT cyber drill should not end when a security team detects an alert. Its decisive outcome is whether a factory can isolate an environment that may be compromised, rebuild trusted control functions, and resume production without reinfection or bypassing safety, quality, and traceability controls. The factory must also be able to explain every decision with evidence. In Thailand, local production, maintenance, IT, regional headquarters, equipment vendors, and outside responders may all participate. Decision rights and bilingual communication are therefore as important as technical recovery procedures.
This guide combines OT incident response, factory ransomware protection, OT backup recovery, and factory cyber exercises into one 90-day program. All timelines, roles, gates, and scorecards below are illustrative assumptions for a model case, not legal requirements or factual industry benchmarks. The feasible test scope must be selected for each site after considering safety, quality, production commitments, warranties, and operating constraints.
The success metric is safe, provable restoration
Many exercises declare success when a SOC detects suspicious traffic and calls the plant. OT recovery becomes difficult after that moment. Reconnecting too early may reinfect rebuilt assets. An old PLC project may omit a safety interlock. A booting HMI does not prove that recipes, alarms, time synchronization, quality decisions, or lot genealogy are correct.
A drill should therefore produce evidence-based answers to these questions:
- Which assets were isolated, when, and on what evidence?
- Was the selected backup an approved configuration and checked for malicious content?
- Were dependencies among the PLC, HMI, engineering workstation, historian, and network configuration respected?
- Were emergency stops, guards, and process limits retested?
- Were approved recipes, inspection logic, and traceability restored?
- Who authorized connection from the clean environment to production?
- What monitoring and observation showed that reinfection had not occurred?
NIST SP 800-82 Rev. 3 explains that OT security must address distinctive performance, reliability, and safety requirements. From that perspective, rebuilding computers is not the end state. Recovery is complete only when physical risk is controlled and production, safety, and quality owners have sufficient evidence to authorize operation.
Define the minimum safe operating state first
Before writing the attack scenario, define a minimum safe operating state. This is the restricted state in which the plant can make a limited product, at a limited rate, on specified equipment, perhaps with approved manual workarounds, while preserving safety and quality. It is not merely degraded operation; it is the boundary for release.
An illustrative model for one line is shown below.
| Dimension | Example minimum safe state | Evidence before release |
|---|---|---|
| Safety | Emergency stops, guards, and critical interlocks run approved logic | Inspection record, PLC comparison, witness sign-off |
| Quality | Only approved recipes; first-piece and added sampling performed | Recipe revision, inspection result, quality approval |
| Production | One line, one product, low-speed start | Work order, speed setting, assigned observer |
| Traceability | Lot, material, equipment, and time recorded in MES or controlled fallback | Genealogy test and time-sync record |
| Connectivity | Only explicitly required flows allowed | Temporary flow list, firewall delta, approval |
| Monitoring | Reinfection indicators watched against predefined stop criteria | Preserved logs and monitoring record |
The values are site-specific. If controlled paper records cannot support the process, MES or historian recovery may be mandatory before any release. If a validated manual method exists, restricted production may start earlier. The critical point is to settle the rule before an incident.
For prevention and asset context, see the Thailand factory OT security guide. For boundaries and permitted communications, see the industrial network construction guide.
A 90-day factory cyber exercise model
Ninety days is an illustrative program length. A site may need longer when maintenance windows are rare or backups are unproven. A mature site may compress some phases.
| Period | Work | Exit condition |
|---|---|---|
| Days 1–15 | Select process, critical assets, dependencies, safe-state criteria, owners | Scope and release criteria approved |
| Days 16–30 | Build scenario, contact tree, decision rights, evidence forms, backup register | Safety, quality, and production approve plan and stop rules |
| Days 31–45 | Run tabletop decision and communication exercise | Conflicts, delays, and unresolved decisions recorded |
| Days 46–65 | Perform functional restore in an isolated environment | System rebuilt and configuration/data verified |
| Days 66–80 | Conduct scoped partial operational test | Restricted physical operation meets safety and quality conditions |
| Days 81–90 | Complete after-action review and improvement plan | Each action has owner, date, and closure evidence |

Days 1–15: narrow the scope and map dependencies deeply
Do not begin with the whole factory. Start with one critical line, one representative product, and one PLC/HMI pair. The inventory should include firmware, PLC project, HMI application, recipes, licenses, service accounts, certificates, time source, backup location, restore software, cables, and vendor contacts—not only model and IP address.
Trace dependencies through power, switching, name resolution, identity, historian, MES, ERP, and quality instruments. A good PLC backup may be useless if the license service or proprietary driver is unavailable. Document shutdown order separately from restoration order.
Days 16–30: approve scenario and abort criteria
A model scenario is ransomware activity on a maintenance engineering workstation, followed by suspicious authentication attempts against a shared folder and HMI. Never use real malware. Controllers inject simulated logs, calls, screenshots, and file events.
Abort criteria are essential. Stop the drill if unexpected equipment motion, safety mismatch, quality deviation, unauthorized traffic, backup damage, or a genuine production emergency occurs. Safety and production authorities should have clear power to stop the exercise.
Days 31–45: test decisions at the tabletop
A tabletop is a facilitated discussion. Participants state whom they contact, what they isolate, what evidence they preserve, and when they escalate. They do not restore equipment. The objective is to test whether decisions connect under time pressure, not whether a procedure document exists.
CISA Cybersecurity Scenarios include ICS compromise and critical-manufacturing material. CISA CTEP Package Documents provide planning, evaluation, and after-action templates. These are inputs to adapt to the plant; filling out a template does not demonstrate readiness.
Days 46–65: perform functional OT backup recovery
Use an isolated test network, equivalent hardware, or an approved test environment and actually restore from backup. Seeing a file in storage is not a recovery test. Confirm that restore tools, credentials, encryption keys, PLC/HMI revisions, and dependency services work.
NIST SP 1339 says OT backup management should be integrated into change management, performed regularly, tested, and reviewed in recovery exercises. A successful backup job therefore does not prove recoverability. The exercise must trace whether post-change backups match approved configurations.
Days 66–80: test part of the physical process
A partial operational test uses the restored configuration under restrictions such as no load, low speed, representative workpiece, short duration, and extra observers. Validate safety interlocks, alarms, recipes, inspection, traceability, and stop procedures in sequence.
This does not mean every exercise must shut down live production. When operating constraints are severe, use FAT equipment, spare PLCs, a controlled testbed, or part of a planned maintenance stop. Record what could not be tested as residual risk and schedule it for an appropriate opportunity.
Build an isolated clean recovery zone
Restore into a clean recovery zone separated from the normal production network. Manage a validation switch, rebuild workstation, backup media, malware-checking capability, and configuration comparison tools there. Do not automatically rely on the same identity infrastructure that may be compromised.

A controlled recovery flow is:
- Select a backup created before the suspected compromise and linked to a change record.
- Verify media hash, custody, encryption, and readability.
- Restore from a clean workstation into the isolated zone.
- Compare PLC logic, HMI screens, recipes, and network configuration with approved baselines.
- Check for malicious indicators, unnecessary accounts, unexpected startup items, and unknown communications.
- Function-test safety and quality controls and approve any variance.
- Present an evidence pack to the release authority before production connection.
- Increase monitoring after connection and predefine conditions for immediate re-isolation.
The CISA StopRansomware Guide supports restoration from offline, encrypted backups and cautions against reinfecting clean systems. It cannot define the plant-specific PLC order or quality release. The drill must prove those local steps.
Assign decision rights by event, not by title
A contact list that says owner is insufficient. Decide who may stop a line, isolate an asset, preserve evidence, contact outside parties, authorize a restore, permit restricted operation, and approve full production. The following is an illustrative model.
| Decision | Responsible | Final authority | Required consultation | Informed |
|---|---|---|---|---|
| Isolate suspected asset | OT/IT responder | Incident commander | Production, maintenance | Plant manager, SOC |
| Stop line | Production lead | Plant manager or delegate | Safety, quality, OT | Regional HQ |
| Acquire evidence | IT/OT evidence lead | Incident commander | Legal/HR when relevant | Management |
| Begin restore | OT recovery lead | Incident commander | Vendor, safety | Production, quality |
| Release minimum safe operation | Production | Plant manager | Safety, quality, OT | HQ, customer interface |
| Return to full production | Production | Plant manager | Quality, safety, IT/OT | Relevant stakeholders |
Legal, regulatory, insurance, and customer notifications vary by jurisdiction, contract, and incident. This table is not legal advice. Confirm the decision path with counsel and applicable agreements.
Make the evidence chain the center of recovery
Evidence is not only for later investigation. It supports safe release. Connect original alert data, timestamps, isolation actions, network deltas, backup IDs and hashes, restore logs, configuration comparisons, interlock results, first-piece inspection, connection approval, and monitoring results in one timeline.

If system clocks differ, record the offset. A screenshot alone is weak: also capture who performed the action, from which workstation, under which change or approval number, and why an exception was accepted. Store the evidence outside the affected environment with defined access and retention.
| Assessment | Evidence | Result |
|---|---|---|
| Isolation completeness | Firewall/switch delta and traffic capture | Confirmed / unknown |
| Backup trust | Revision, hash, change ID, scan result | Confirmed / unknown |
| Safety functions | Interlock test record | Pass / fail |
| Quality functions | Recipe comparison and first-piece result | Pass / fail |
| Traceability | Forward/backward test of a trial lot | Pass / fail |
| Reinfection monitoring | Endpoint, network, and system logs | Clear / investigate |
| Authorization | Commander, safety, quality, production sign-off | Complete / incomplete |
Unknown is a valid result. It should block the relevant release decision, not be converted into a convenient pass.
Design bilingual and regional-HQ communication
A technically correct recovery may stall because words are interpreted differently. Use structured status fields: event time and timezone, line, observed facts, unknowns, isolation completed, next decision, authority, customer impact, safety impact, quality impact, and next update time.
A Thai-English or Thai-Japanese glossary should distinguish isolate, shutdown, emergency stop, safe state, restore, release, suspected, and confirmed. Restore complete is not the same as production release. For critical decisions, the approver should repeat the decision and conditions through closed-loop communication, even when an interpreter participates.
Headquarters needs a decision-ready summary, not a dump of logs. The site needs the reason and deadline for HQ decisions. If corporate email and chat may be unavailable, inject that failure and test a phone tree and offline contact sheet.
Ransomware injects for a factory exercise
To avoid a scripted performance, an exercise controller can release model injects such as:
- a night login from a shared HMI account;
- a PLC project timestamp that conflicts with change records;
- encryption of an online backup share;
- readable offline media whose latest equipment change is uncertain;
- HQ demanding full shutdown while local production proposes restricted operation;
- a customer asking for shipment and genealogy evidence;
- an unavailable equipment-vendor specialist;
- an unknown DNS request after restoration.
The final inject tests whether the team rushes to reconnect or returns the asset to isolation. The goal is not heroic speed. It is disciplined reduction of uncertainty while protecting safety and quality.
Keep the three exercise levels separate
Tabletop discussion
Tests people, authority, communication, priorities, and alternatives. It does not prove that equipment can be restored.
Functional restore test
Requires responders to use tools and rebuild in isolation. It tests credentials, media, restore duration, and configuration differences. It provides limited proof about the physical process.
Partial operational test
Validates equipment, safety, quality, and traceability under controlled conditions. Scope follows risk and production constraints. Record untested elements as residual risk.
CISA guidance on ICS incident-response capability discusses realistic, difficult scenarios, recurring exercises, and partial or full tests where feasible. Define discussed, restored, and operated safely as separate exit conditions.
Measure quality as well as time
Recovery time matters, but optimizing only for speed encourages skipped checks. The following metrics are illustrative assumptions, not benchmarks.
| Metric | What it reveals | Bad shortcut |
|---|---|---|
| Alert to commander assignment | Speed of command | Declaring cause without evidence |
| Isolation decision to verification | Containment effectiveness | Power-off without safety review |
| Time to select restore point | Backup-register quality | Choosing latest automatically |
| Time to approve configuration | Repeatability of clean restore | Ignoring unexplained deltas |
| Time to restricted release | Cross-functional readiness | Omitting safety/quality approval |
| Missing evidence count | Explainability | Recreating records from memory |
| Recurrence of unknown traffic | Reinfection risk | Shortening observation period |
In a model scorecard, any failed safety or quality gate should prevent release regardless of total points. Scores help prioritize improvement; they must not average away critical conditions.
Turn the after-action report into improvement
An AAR states what occurred, what was expected, why the difference existed, who will fix it by when, and what proves closure. Convert findings into system improvements, not blame.
| Finding | Action | Closure evidence |
|---|---|---|
| PLC backup revision unknown | Add backup and hash registration to change closure | Change ticket and restore-test result |
| Contact list obsolete | Monthly validation and alternates | Validation log |
| Quality release missing | Add first-piece inspection and authority to recovery plan | Revised procedure and exercise record |
| Rebuild workstation depends on production identity | Prepare independent clean workstation and credentials | Asset record and boot test |
| HQ decision delayed | Agree delegation conditions and time limit | Approved decision matrix |
Retest actions in the next exercise. Updating a document is not closure until the process works. If a finding needs major investment, report temporary controls, residual risk, and a budget decision date.
Read regional developments as a readiness signal
On 15 September 2026, Yokogawa Engineering Asia announced an Industrial Cyber Resilience Center in Singapore, positioned as a hub serving Southeast Asia, Oceania, and Taiwan and offering OT readiness and capability development. This is a company announcement, not independent proof of its marketing claims. It is, however, one current signal that regional services for OT resilience are expanding.
Using an outside center or vendor does not transfer the plant’s decision rights. Define scope, data movement, log retention, remote access, warranty constraints, confidentiality, local response, and exercise deliverables before engagement. A vendor may provide a restore tool, but it cannot automatically own the site’s safety and quality release.
Common failures and corrections
- Ending at the SOC alert. Extend the exit criteria through restore, quality release, restricted operation, and reinfection monitoring.
- Checking only that backups exist. Restore in isolation and verify licenses, versions, and dependencies.
- Trusting the newest backup. Select a trusted point using compromise time, change history, hashes, and scans.
- Exercising only OT. Include safety, quality, production, IT, management, HQ, and vendors where relevant.
- Avoiding all physical tests—or forcing unsafe live tests. Combine testbeds, spares, maintenance windows, and limited operation, and record residual risk.
- Leaving improvements undated. Assign owner, due date, closure evidence, and retest date.
Pre-exercise checklist
- Scope and exclusions are visible on a diagram.
- Safety, quality, production, and OT approve the minimum safe operating state.
- Exercise director, commander, stop authority, and release authority are named.
- Simulated actions are separated from actions that affect real equipment.
- Abort criteria and transition to real-incident handling are defined.
- Backup ID, revision, hash, change ID, and restore tool are checked.
- Clean recovery zone and evidence storage are ready.
- Bilingual status format and alternative communications are tested.
- Safety, quality, and traceability acceptance tests are prepared.
- AAR owner, action register, and retest date are scheduled.
FAQ: OT incident response and factory cyber exercises
What does an OT cyber drill include?
It should progress through a tabletop for decisions and communication, a functional restore test in isolation, and a scoped partial operational test of equipment, safety, quality, and traceability. The three levels have different purposes; tabletop discussion alone does not prove recoverability.
What should a factory restore first after ransomware?
There is no universal order. Pre-map safety functions, control dependencies, identity, networking, PLC/HMI, historian, quality, and traceability. Restore what is necessary for the minimum safe operating state without immediately reusing potentially compromised identity services or online shares.
How often should OT backup recovery be tested?
Frequency depends on equipment change, criticality, maintenance windows, and risk. There is no universal interval. At minimum, important changes should trigger backup/update checks, and periodic recovery exercises should review actual restoration capability.
Must the factory stop a live line for testing?
Not always. Use risk assessment and combine testbeds, spare PLCs, FAT environments, planned downtime, low speed, or no-load operation. Record items that only a live line can validate as residual risk and schedule them appropriately.
Is recovery time enough to score a factory cyber exercise?
No. Also assess isolation, backup trust, safety, quality, traceability, evidence completeness, reinfection indicators, and authorization. A fast restore that skipped safety or quality checks is not a pass.
How should overseas HQ and vendors participate?
Include them where real incidents require decisions, remote access, restoration, or customer communication. Exercise time-zone and language delay. Preserve clear production-release and safety/quality authority under company rules and contracts.
Conclusion: prove recovery with safety and evidence
The value of an OT cyber drill is not making an alert flash. It is demonstrating that the plant can restore production without unsafe shortcuts. Define the minimum safe state, allocate decision rights, test backups in a clean zone, progress from tabletop to functional and partial operational tests, and preserve evidence for safety, quality, traceability, and reinfection monitoring. Then close after-action improvements through retesting.
TOMAS TECH can support scoping, the 90-day exercise plan, OT backup recovery tests, and bilingual communication design while the program is still at the planning stage. To discuss an approach shaped around a Thai factory’s equipment and operating constraints, use the contact page.
References
- Yokogawa Engineering Asia company announcement, 15 September 2026: https://en.prnasia.com/releases/apac/yokogawa-launches-industrial-cyber-resilience-center-to-advance-ot-cyber-readiness-in-southeast-asia-547836.shtml
- NIST SP 800-82 Rev. 3: https://www.nist.gov/publications/guide-operational-technology-ot-security
- NIST SP 1339, OT Backup Quick Start Guide: https://csrc.nist.gov/pubs/sp/1339/final
- CISA Cybersecurity Scenarios: https://www.cisa.gov/resources-tools/resources/cybersecurity-scenarios
- CISA CTEP Package Documents: https://www.cisa.gov/resources-tools/resources/ctep-package-documents
- CISA ICS Cybersecurity Incident Response Recommended Practice: https://www.cisa.gov/sites/default/files/2023-01/final-RP_ics_cybersecurity_incident_response_100609.pdf
- CISA StopRansomware Guide: https://www.cisa.gov/stopransomware/ransomware-guide
This article provides general technical planning information. It is not legal advice, a compliance guarantee, or a recovery guarantee from any product or vendor.