Blog

2026.09.16

OT Cyber Drill: Prove Safe Factory Recovery in 90 Days

OT Cyber Drill: Prove Safe Factory Recovery in 90 Days

An OT cyber drill should not end when a security team detects an alert. Its decisive outcome is whether a factory can isolate an environment that may be compromised, rebuild trusted control functions, and resume production without reinfection or bypassing safety, quality, and traceability controls. The factory must also be able to explain every decision with evidence. In Thailand, local production, maintenance, IT, regional headquarters, equipment vendors, and outside responders may all participate. Decision rights and bilingual communication are therefore as important as technical recovery procedures.

This guide combines OT incident response, factory ransomware protection, OT backup recovery, and factory cyber exercises into one 90-day program. All timelines, roles, gates, and scorecards below are illustrative assumptions for a model case, not legal requirements or factual industry benchmarks. The feasible test scope must be selected for each site after considering safety, quality, production commitments, warranties, and operating constraints.

The success metric is safe, provable restoration

Many exercises declare success when a SOC detects suspicious traffic and calls the plant. OT recovery becomes difficult after that moment. Reconnecting too early may reinfect rebuilt assets. An old PLC project may omit a safety interlock. A booting HMI does not prove that recipes, alarms, time synchronization, quality decisions, or lot genealogy are correct.

A drill should therefore produce evidence-based answers to these questions:

  • Which assets were isolated, when, and on what evidence?
  • Was the selected backup an approved configuration and checked for malicious content?
  • Were dependencies among the PLC, HMI, engineering workstation, historian, and network configuration respected?
  • Were emergency stops, guards, and process limits retested?
  • Were approved recipes, inspection logic, and traceability restored?
  • Who authorized connection from the clean environment to production?
  • What monitoring and observation showed that reinfection had not occurred?

NIST SP 800-82 Rev. 3 explains that OT security must address distinctive performance, reliability, and safety requirements. From that perspective, rebuilding computers is not the end state. Recovery is complete only when physical risk is controlled and production, safety, and quality owners have sufficient evidence to authorize operation.

Define the minimum safe operating state first

Before writing the attack scenario, define a minimum safe operating state. This is the restricted state in which the plant can make a limited product, at a limited rate, on specified equipment, perhaps with approved manual workarounds, while preserving safety and quality. It is not merely degraded operation; it is the boundary for release.

An illustrative model for one line is shown below.

DimensionExample minimum safe stateEvidence before release
SafetyEmergency stops, guards, and critical interlocks run approved logicInspection record, PLC comparison, witness sign-off
QualityOnly approved recipes; first-piece and added sampling performedRecipe revision, inspection result, quality approval
ProductionOne line, one product, low-speed startWork order, speed setting, assigned observer
TraceabilityLot, material, equipment, and time recorded in MES or controlled fallbackGenealogy test and time-sync record
ConnectivityOnly explicitly required flows allowedTemporary flow list, firewall delta, approval
MonitoringReinfection indicators watched against predefined stop criteriaPreserved logs and monitoring record

The values are site-specific. If controlled paper records cannot support the process, MES or historian recovery may be mandatory before any release. If a validated manual method exists, restricted production may start earlier. The critical point is to settle the rule before an incident.

For prevention and asset context, see the Thailand factory OT security guide. For boundaries and permitted communications, see the industrial network construction guide.

A 90-day factory cyber exercise model

Ninety days is an illustrative program length. A site may need longer when maintenance windows are rare or backups are unproven. A mature site may compress some phases.

PeriodWorkExit condition
Days 1–15Select process, critical assets, dependencies, safe-state criteria, ownersScope and release criteria approved
Days 16–30Build scenario, contact tree, decision rights, evidence forms, backup registerSafety, quality, and production approve plan and stop rules
Days 31–45Run tabletop decision and communication exerciseConflicts, delays, and unresolved decisions recorded
Days 46–65Perform functional restore in an isolated environmentSystem rebuilt and configuration/data verified
Days 66–80Conduct scoped partial operational testRestricted physical operation meets safety and quality conditions
Days 81–90Complete after-action review and improvement planEach action has owner, date, and closure evidence
OT Cyber Drill: Prove Safe Factory Recovery in 90 Days - figure 1

Days 1–15: narrow the scope and map dependencies deeply

Do not begin with the whole factory. Start with one critical line, one representative product, and one PLC/HMI pair. The inventory should include firmware, PLC project, HMI application, recipes, licenses, service accounts, certificates, time source, backup location, restore software, cables, and vendor contacts—not only model and IP address.

Trace dependencies through power, switching, name resolution, identity, historian, MES, ERP, and quality instruments. A good PLC backup may be useless if the license service or proprietary driver is unavailable. Document shutdown order separately from restoration order.

Days 16–30: approve scenario and abort criteria

A model scenario is ransomware activity on a maintenance engineering workstation, followed by suspicious authentication attempts against a shared folder and HMI. Never use real malware. Controllers inject simulated logs, calls, screenshots, and file events.

Abort criteria are essential. Stop the drill if unexpected equipment motion, safety mismatch, quality deviation, unauthorized traffic, backup damage, or a genuine production emergency occurs. Safety and production authorities should have clear power to stop the exercise.

Days 31–45: test decisions at the tabletop

A tabletop is a facilitated discussion. Participants state whom they contact, what they isolate, what evidence they preserve, and when they escalate. They do not restore equipment. The objective is to test whether decisions connect under time pressure, not whether a procedure document exists.

CISA Cybersecurity Scenarios include ICS compromise and critical-manufacturing material. CISA CTEP Package Documents provide planning, evaluation, and after-action templates. These are inputs to adapt to the plant; filling out a template does not demonstrate readiness.

Days 46–65: perform functional OT backup recovery

Use an isolated test network, equivalent hardware, or an approved test environment and actually restore from backup. Seeing a file in storage is not a recovery test. Confirm that restore tools, credentials, encryption keys, PLC/HMI revisions, and dependency services work.

NIST SP 1339 says OT backup management should be integrated into change management, performed regularly, tested, and reviewed in recovery exercises. A successful backup job therefore does not prove recoverability. The exercise must trace whether post-change backups match approved configurations.

Days 66–80: test part of the physical process

A partial operational test uses the restored configuration under restrictions such as no load, low speed, representative workpiece, short duration, and extra observers. Validate safety interlocks, alarms, recipes, inspection, traceability, and stop procedures in sequence.

This does not mean every exercise must shut down live production. When operating constraints are severe, use FAT equipment, spare PLCs, a controlled testbed, or part of a planned maintenance stop. Record what could not be tested as residual risk and schedule it for an appropriate opportunity.

Build an isolated clean recovery zone

Restore into a clean recovery zone separated from the normal production network. Manage a validation switch, rebuild workstation, backup media, malware-checking capability, and configuration comparison tools there. Do not automatically rely on the same identity infrastructure that may be compromised.

OT Cyber Drill: Prove Safe Factory Recovery in 90 Days - figure 2

A controlled recovery flow is:

  1. Select a backup created before the suspected compromise and linked to a change record.
  2. Verify media hash, custody, encryption, and readability.
  3. Restore from a clean workstation into the isolated zone.
  4. Compare PLC logic, HMI screens, recipes, and network configuration with approved baselines.
  5. Check for malicious indicators, unnecessary accounts, unexpected startup items, and unknown communications.
  6. Function-test safety and quality controls and approve any variance.
  7. Present an evidence pack to the release authority before production connection.
  8. Increase monitoring after connection and predefine conditions for immediate re-isolation.

The CISA StopRansomware Guide supports restoration from offline, encrypted backups and cautions against reinfecting clean systems. It cannot define the plant-specific PLC order or quality release. The drill must prove those local steps.

Assign decision rights by event, not by title

A contact list that says owner is insufficient. Decide who may stop a line, isolate an asset, preserve evidence, contact outside parties, authorize a restore, permit restricted operation, and approve full production. The following is an illustrative model.

DecisionResponsibleFinal authorityRequired consultationInformed
Isolate suspected assetOT/IT responderIncident commanderProduction, maintenancePlant manager, SOC
Stop lineProduction leadPlant manager or delegateSafety, quality, OTRegional HQ
Acquire evidenceIT/OT evidence leadIncident commanderLegal/HR when relevantManagement
Begin restoreOT recovery leadIncident commanderVendor, safetyProduction, quality
Release minimum safe operationProductionPlant managerSafety, quality, OTHQ, customer interface
Return to full productionProductionPlant managerQuality, safety, IT/OTRelevant stakeholders

Legal, regulatory, insurance, and customer notifications vary by jurisdiction, contract, and incident. This table is not legal advice. Confirm the decision path with counsel and applicable agreements.

Make the evidence chain the center of recovery

Evidence is not only for later investigation. It supports safe release. Connect original alert data, timestamps, isolation actions, network deltas, backup IDs and hashes, restore logs, configuration comparisons, interlock results, first-piece inspection, connection approval, and monitoring results in one timeline.

OT Cyber Drill: Prove Safe Factory Recovery in 90 Days - figure 3

If system clocks differ, record the offset. A screenshot alone is weak: also capture who performed the action, from which workstation, under which change or approval number, and why an exception was accepted. Store the evidence outside the affected environment with defined access and retention.

AssessmentEvidenceResult
Isolation completenessFirewall/switch delta and traffic captureConfirmed / unknown
Backup trustRevision, hash, change ID, scan resultConfirmed / unknown
Safety functionsInterlock test recordPass / fail
Quality functionsRecipe comparison and first-piece resultPass / fail
TraceabilityForward/backward test of a trial lotPass / fail
Reinfection monitoringEndpoint, network, and system logsClear / investigate
AuthorizationCommander, safety, quality, production sign-offComplete / incomplete

Unknown is a valid result. It should block the relevant release decision, not be converted into a convenient pass.

Design bilingual and regional-HQ communication

A technically correct recovery may stall because words are interpreted differently. Use structured status fields: event time and timezone, line, observed facts, unknowns, isolation completed, next decision, authority, customer impact, safety impact, quality impact, and next update time.

A Thai-English or Thai-Japanese glossary should distinguish isolate, shutdown, emergency stop, safe state, restore, release, suspected, and confirmed. Restore complete is not the same as production release. For critical decisions, the approver should repeat the decision and conditions through closed-loop communication, even when an interpreter participates.

Headquarters needs a decision-ready summary, not a dump of logs. The site needs the reason and deadline for HQ decisions. If corporate email and chat may be unavailable, inject that failure and test a phone tree and offline contact sheet.

Ransomware injects for a factory exercise

To avoid a scripted performance, an exercise controller can release model injects such as:

  • a night login from a shared HMI account;
  • a PLC project timestamp that conflicts with change records;
  • encryption of an online backup share;
  • readable offline media whose latest equipment change is uncertain;
  • HQ demanding full shutdown while local production proposes restricted operation;
  • a customer asking for shipment and genealogy evidence;
  • an unavailable equipment-vendor specialist;
  • an unknown DNS request after restoration.

The final inject tests whether the team rushes to reconnect or returns the asset to isolation. The goal is not heroic speed. It is disciplined reduction of uncertainty while protecting safety and quality.

Keep the three exercise levels separate

Tabletop discussion

Tests people, authority, communication, priorities, and alternatives. It does not prove that equipment can be restored.

Functional restore test

Requires responders to use tools and rebuild in isolation. It tests credentials, media, restore duration, and configuration differences. It provides limited proof about the physical process.

Partial operational test

Validates equipment, safety, quality, and traceability under controlled conditions. Scope follows risk and production constraints. Record untested elements as residual risk.

CISA guidance on ICS incident-response capability discusses realistic, difficult scenarios, recurring exercises, and partial or full tests where feasible. Define discussed, restored, and operated safely as separate exit conditions.

Measure quality as well as time

Recovery time matters, but optimizing only for speed encourages skipped checks. The following metrics are illustrative assumptions, not benchmarks.

MetricWhat it revealsBad shortcut
Alert to commander assignmentSpeed of commandDeclaring cause without evidence
Isolation decision to verificationContainment effectivenessPower-off without safety review
Time to select restore pointBackup-register qualityChoosing latest automatically
Time to approve configurationRepeatability of clean restoreIgnoring unexplained deltas
Time to restricted releaseCross-functional readinessOmitting safety/quality approval
Missing evidence countExplainabilityRecreating records from memory
Recurrence of unknown trafficReinfection riskShortening observation period

In a model scorecard, any failed safety or quality gate should prevent release regardless of total points. Scores help prioritize improvement; they must not average away critical conditions.

Turn the after-action report into improvement

An AAR states what occurred, what was expected, why the difference existed, who will fix it by when, and what proves closure. Convert findings into system improvements, not blame.

FindingActionClosure evidence
PLC backup revision unknownAdd backup and hash registration to change closureChange ticket and restore-test result
Contact list obsoleteMonthly validation and alternatesValidation log
Quality release missingAdd first-piece inspection and authority to recovery planRevised procedure and exercise record
Rebuild workstation depends on production identityPrepare independent clean workstation and credentialsAsset record and boot test
HQ decision delayedAgree delegation conditions and time limitApproved decision matrix

Retest actions in the next exercise. Updating a document is not closure until the process works. If a finding needs major investment, report temporary controls, residual risk, and a budget decision date.

Read regional developments as a readiness signal

On 15 September 2026, Yokogawa Engineering Asia announced an Industrial Cyber Resilience Center in Singapore, positioned as a hub serving Southeast Asia, Oceania, and Taiwan and offering OT readiness and capability development. This is a company announcement, not independent proof of its marketing claims. It is, however, one current signal that regional services for OT resilience are expanding.

Using an outside center or vendor does not transfer the plant’s decision rights. Define scope, data movement, log retention, remote access, warranty constraints, confidentiality, local response, and exercise deliverables before engagement. A vendor may provide a restore tool, but it cannot automatically own the site’s safety and quality release.

Common failures and corrections

  1. Ending at the SOC alert. Extend the exit criteria through restore, quality release, restricted operation, and reinfection monitoring.
  2. Checking only that backups exist. Restore in isolation and verify licenses, versions, and dependencies.
  3. Trusting the newest backup. Select a trusted point using compromise time, change history, hashes, and scans.
  4. Exercising only OT. Include safety, quality, production, IT, management, HQ, and vendors where relevant.
  5. Avoiding all physical tests—or forcing unsafe live tests. Combine testbeds, spares, maintenance windows, and limited operation, and record residual risk.
  6. Leaving improvements undated. Assign owner, due date, closure evidence, and retest date.

Pre-exercise checklist

  • Scope and exclusions are visible on a diagram.
  • Safety, quality, production, and OT approve the minimum safe operating state.
  • Exercise director, commander, stop authority, and release authority are named.
  • Simulated actions are separated from actions that affect real equipment.
  • Abort criteria and transition to real-incident handling are defined.
  • Backup ID, revision, hash, change ID, and restore tool are checked.
  • Clean recovery zone and evidence storage are ready.
  • Bilingual status format and alternative communications are tested.
  • Safety, quality, and traceability acceptance tests are prepared.
  • AAR owner, action register, and retest date are scheduled.

FAQ: OT incident response and factory cyber exercises

What does an OT cyber drill include?

It should progress through a tabletop for decisions and communication, a functional restore test in isolation, and a scoped partial operational test of equipment, safety, quality, and traceability. The three levels have different purposes; tabletop discussion alone does not prove recoverability.

What should a factory restore first after ransomware?

There is no universal order. Pre-map safety functions, control dependencies, identity, networking, PLC/HMI, historian, quality, and traceability. Restore what is necessary for the minimum safe operating state without immediately reusing potentially compromised identity services or online shares.

How often should OT backup recovery be tested?

Frequency depends on equipment change, criticality, maintenance windows, and risk. There is no universal interval. At minimum, important changes should trigger backup/update checks, and periodic recovery exercises should review actual restoration capability.

Must the factory stop a live line for testing?

Not always. Use risk assessment and combine testbeds, spare PLCs, FAT environments, planned downtime, low speed, or no-load operation. Record items that only a live line can validate as residual risk and schedule them appropriately.

Is recovery time enough to score a factory cyber exercise?

No. Also assess isolation, backup trust, safety, quality, traceability, evidence completeness, reinfection indicators, and authorization. A fast restore that skipped safety or quality checks is not a pass.

How should overseas HQ and vendors participate?

Include them where real incidents require decisions, remote access, restoration, or customer communication. Exercise time-zone and language delay. Preserve clear production-release and safety/quality authority under company rules and contracts.

Conclusion: prove recovery with safety and evidence

The value of an OT cyber drill is not making an alert flash. It is demonstrating that the plant can restore production without unsafe shortcuts. Define the minimum safe state, allocate decision rights, test backups in a clean zone, progress from tabletop to functional and partial operational tests, and preserve evidence for safety, quality, traceability, and reinfection monitoring. Then close after-action improvements through retesting.

TOMAS TECH can support scoping, the 90-day exercise plan, OT backup recovery tests, and bilingual communication design while the program is still at the planning stage. To discuss an approach shaped around a Thai factory’s equipment and operating constraints, use the contact page.

References

This article provides general technical planning information. It is not legal advice, a compliance guarantee, or a recovery guarantee from any product or vendor.