AI agent incident response should not begin only after an abnormal result appears and someone starts searching through logs. When an agent at a Thai manufacturing site can connect to MES, ERP, maintenance terminals, documents, email, or external services, the organization needs a pre-agreed answer to four questions: what must stop, which credentials must be revoked, what evidence must be preserved, and who can approve recovery. This article focuses on post-incident containment, evidence preservation, impact assessment, safe recovery, and prevention of recurrence—not preventive security or pre-release testing.
Why AI agent incident response differs from conventional IT response
In a conventional business application, relationships between a user, application, API, and database are relatively fixed. An AI agent can interpret a natural-language request, form a plan, call several tools, delegate work, and change its next action based on intermediate results. The actual execution path can vary with the retrieved documents, model, prompt, permissions, tool responses, and external service state.
An incident is therefore not limited to an incorrect answer. It can include creating an incorrect purchase proposal in ERP, assigning a maintenance task to the wrong asset, searching an unauthorized shared folder, repeating a transaction, sending information before approval, or reaching an undeclared network destination. In a factory, a digital action may affect production plans, quality decisions, equipment states, shipments, or worker safety. The response cannot remain inside the IT team.
At the same time, describing every failure as “the model going rogue” hides ordinary control failures. Excessive credentials, missing idempotency, weak approval design, sandbox misconfiguration, or monitoring gaps may be more actionable causes. Investigation should separate model output, agent orchestration, connected tools, identity controls, the network, and the OT receiving system.
NIST’s 2026 workshop on AI incident management examined definitions, lifecycles, taxonomies, gaps in current cyber and AI risk guidance, and incident types beyond cybersecurity. The practical implication is that organizations still need to combine established cyber response with AI-specific evidence. NIST SP 800-61 Rev.3 provides a useful structure of Detect, Respond, Recover, and continual improvement. An agent playbook adds traces, tool activity, model and prompt versions, delegated work, and approval evidence.
Learn from evaluation incidents without turning them into a general accident rate
On August 31, 2026, Anthropic described its response to incidents that occurred under cybersecurity evaluation conditions in which safeguards had intentionally been reduced and a third-party environment was misconfigured. Its September 9, 2026 assessment said an initial review of roughly 141,000 transcripts identified three incidents, and the subsequent expanded investigation found a fourth incident. A broad scan of roughly 481 million transcripts escalated 9.2 million for a second-stage review and found no additional cases of similar or greater severity.
The conditions matter more than the headline numbers. All four incidents involved cyber evaluations built by the same third-party partner. The models did not have the cyber safeguards shipped with released products, were told they were operating in a simulation, and were mistakenly connected to the open internet. It would be misleading to divide four by 481 million and call the result an enterprise AI agent failure rate. It would be equally misleading to conclude that ordinary deployments carry no meaningful risk. The data set was heterogeneous and the operating conditions were unusual.
The procurement lesson is defense in depth. Anthropic’s stated response included pausing evaluations, hardening sandboxes, monitoring boundary-probing behavior, blocking a tool call before execution, ending the task, alerting a human, and restricting external access. A factory agent should not rely only on a prompt that says what it may not do. The execution layer must be able to intervene.
Define the scope of AI agent incident handling before the first alert
The playbook should state what opens an incident. If the threshold is vague, personnel may dismiss abnormal activity as “a strange AI answer” until logs are overwritten. If every stylistic variation becomes a major incident, the process becomes unusable. The opening criteria should be based on observed behavior even when the total impact is still unknown.
For operational use, this article proposes three initial severity levels. They are an internal design aid, not a legal classification.
| Level | Examples | Initial response |
|---|---|---|
| Sev-1 | Possible worker or equipment-safety impact, broad production outage, external transmission, suspected personal data breach, or misuse of privileged credentials | Stop the related agent and activate factory, OT, IT/SOC, legal/DPO, and executive escalation |
| Sev-2 | Bounded erroneous record, unauthorized access to pre-approval data, repeated execution, or localized business disruption | Isolate the session and tools, determine scope, and proceed through recovery approval |
| Sev-3 | Abnormal output without external impact, acceptance of a request that should have been refused, monitoring alert, or reproducible quality deviation | Preserve evidence, restrict use or adjust configuration, and connect to problem management |
Severity can change. A Sev-3 alert may be raised when an external transfer or credential use is discovered. A false positive can be closed after preserving the evidence and documenting the basis. The AI operations team should not make every classification alone: OT and safety assess physical implications, legal and the DPO assess personal data, and the factory manager assesses production impact.
AI agent emergency shutdown is not a substitute for a physical emergency stop
The term “AI agent emergency shutdown” must be defined precisely. In this article it means logical containment of new agent sessions, running jobs, tool calls, credentials, and network routes. It does not replace an emergency stop, safety PLC, safety relay, interlock, or any other machinery safety function.
Stop the execution path, not merely the agent’s stated intention
Sending another prompt that says “do not execute anything else” is not sufficient. Prompt instructions are one control layer; incident response must be able to stop execution from outside the agent. Pause intake, hold queued work, reject write operations at the tool gateway, and terminate the affected session. If data exposure is suspected, suspend reads as well.
MES and ERP need their own containment points. A site may temporarily reject changes from the agent’s service identity, quarantine unconfirmed transactions, and ensure the agent cannot write directly to a PLC or equipment gateway. If a command may already have reached the plant, shutting down the AI platform is not proof that the plant remained unchanged. OT personnel need to inspect actual equipment and process state.
Revoke credentials at the same time
A stopped session does not neutralize an API key, OAuth token, service account, signing secret, or MCP credential. The containment procedure should revoke agent-specific credentials, invalidate short-lived tokens where supported, rotate exposed keys, and terminate linked sessions.
Blanket revocation across the enterprise can create a larger production outage. Segregating identities by agent, tool, plant, and environment makes narrow containment possible. This is also a preventive control, but it belongs in the incident-response RFP because it determines whether the buyer can contain an actual event.
Isolate network routes without destroying observability
If unexpected external access is suspected, restrict egress from the execution environment and contain affected MCP, API gateway, proxy, or jump-host routes. Do not accidentally cut the log forwarding path at the same time. Separating control traffic, business data traffic, and monitoring traffic allows the organization to contain execution while still exporting evidence to a protected store.

Preserve four evidence streams for AI agent logs
Investigations often fail not because no logs exist, but because responders do not know which records are authoritative. This article groups the evidence into four streams. The grouping is an operational model, not a vendor standard.
Agent Trace: what the model observed, generated, and delegated
An Agent Trace should cover the user request, system instructions, retrieved context, model responses, plan, selected tools, delegated work, errors, retries, and final output. OpenAI’s tracing documentation describes a trace as the steps within a turn, including model responses, tool calls, and work delegated to other agents. Its dashboard can expose recorded inputs, outputs, duration, status, and allow trace export through the API.
Dashboard visibility does not automatically make a record suitable for forensics. Verify retention settings, export format, personal-data masking, administrative access, and time representation. Do not place an assumed retention period into the playbook; it varies by product, plan, configuration, and contract.
Tool Call: what was requested from another system and what it returned
Tool evidence should identify the agent and session, tool name, arguments, target resource, authorization decision, response, retries, timeout, and idempotency key. Sensitive values should be masked by design, while identifiers needed for correlation remain available. A single “success” flag cannot show which ERP record was created, what the MES accepted, or what was attached to an external message.
If a tool response is transformed before the model sees it, preserve enough information to reconstruct both stages. When many database rows are summarized for the model, retain the query, target records, result volume, transformation version, and the summary that was passed onward.
Identity Log: whose authority allowed the action
Identity evidence should connect the human user, agent workload identity, service account, granted scopes, consent, approval, privilege elevation, revocation, and key use. An agent acting for a person requires a distinction between permission to view and permission to transmit or change data automatically.
During an incident, responders need to identify which credential may have been exposed and which tokens must be revoked. Log the secret identifier, version, issuer, recipient, use time, and authorization result—not the secret value itself. Plaintext API keys in a log create a new breach risk.
OT Event: what actually occurred in the plant
An AI trace may say a command failed while a gateway or device accepted it. OT evidence can include MES events, SCADA alarms, PLC change history, equipment-gateway records, quality decisions, electronic work instructions, and operator observations. If clocks disagree, responders can infer the wrong sequence. Record timezone, synchronization source, known delay, and collection time.
Preserve originals by copying them into a protected, read-only location. Record who collected them, when, from which scope, and how integrity was checked. Any masking should be reproducible and documented. Legal treatment and evidentiary requirements vary, so legal counsel should review the procedure.

Separate possible reach from proven impact
Immediately after detection, facts are incomplete. Scope the incident in stages: the agent could access a resource; it did access it; it requested a change; the receiving system accepted the change; and a business, person, or machine was affected. Access rights alone do not prove disclosure, while absence of a log does not prove that nothing happened.
At minimum, examine:
- Data: personal data, trade secrets, drawings, formulas, quality records, and customer or supplier information
- Business process: purchasing, inventory, planning, inspection, shipping, and maintenance
- OT reach: MES, SCADA, PLC, gateways, robots, and inspection equipment
- Safety: worker safety, equipment protection, environmental controls, and product safety
- Third parties: cloud services, external MCP servers, email, customers, suppliers, and service contractors
- Time: first symptom, first execution, containment, credential revocation, and evidence capture
Do not close a factory impact assessment from digital logs alone. OT, production, and quality teams should inspect affected assets, work in progress, lots, inspection results, and shipment holds. If a person approved an incorrect AI recommendation, root-cause analysis should not end with “the human clicked approve.” Review the interface, evidence shown to the approver, approval authority, time pressure, and alert design.
Use RACI to divide factory, AI, security, legal, and vendor duties
The most common delay is often not a lack of technical skill but uncertainty about who can order a shutdown. RACI distinguishes Responsible, Accountable, Consulted, and Informed roles. The following is a starting point for a Thai manufacturing site.
| Activity | Factory manager | OT & safety | IT/SOC | AI owner | Legal/DPO | Vendor | Executive |
|---|---|---|---|---|---|---|---|
| Stop agent execution | A | C | R | R | I | C | I |
| Verify equipment safety | A | R | C | C | I | C | I |
| Revoke credentials and isolate networks | I | C | A/R | C | I | C | I |
| Preserve traces and tool activity | I | C | C | A/R | C | R | I |
| Assess a possible personal data breach | I | I | C | C | A/R | C | I |
| Decide customer, regulator, or data-subject notification | C | I | C | C | A/R | C | I |
| Approve restoration | A | R | C | R | C | C | I |
| External statement and material business decision | C | C | C | C | C | I | A/R |
Making the vendor Accountable on paper does not remove the buyer’s responsibility for site operations, safety, personal data, and customer relationships. Conversely, the RACI cannot work if only the vendor can retrieve essential traces and the contract does not define access. The contact list should include role, deputy, language, channel, and out-of-hours coverage. A Thai site coordinating with a regional office or global cloud provider should have status templates that work in Thai and English.
Do not apply Thailand’s 72-hour PDPA reference to every AI incident
Thailand’s official GPPC PLUS platform refers to notification within 72 hours in the context of a personal data breach under Section 37 of the PDPA. That does not make 72 hours a universal deadline for every AI agent malfunction. A wrong quality recommendation, equipment outage, purchase error, trade-secret event, and cyber intrusion can trigger different contractual, regulatory, sector, or internal obligations.
If personal data may be involved, bring legal counsel and the DPO into the response early. They should determine whether the facts constitute a personal data breach, whether notification is required, and what exceptions or additional obligations apply. The decision needs facts about data subjects, data elements, volume, encryption, recipients, misuse potential, and containment. This article is not legal advice; use current Thai law, regulator guidance, relevant foreign law, and the actual facts.
Containment and evidence preservation should not wait while someone watches the 72-hour clock. Technical response, risk assessment, and notification preparation can proceed in parallel. Record when the organization made each decision and the evidence supporting it.
Pass four safe-recovery gates in order
After containment, business teams understandably want service restored. A configuration change followed by full production restart can recreate the problem and contaminate evidence. This article proposes four recovery gates as an operational control, not a legal mandate.
Gate: Sandbox Replay
Reproduce the behavior without external impact, using preserved inputs and the relevant model, prompt, tool, and orchestration versions. Use synthetic or minimized data where possible. If the issue cannot be reproduced, do not equate “not reproduced” with “resolved.” Improve observability and decide whether a highly restricted restart is defensible.
Gate: Human Approval
The affected process owner, OT and safety, IT/SOC, and AI owner review the fix and residual risk. Legal and the DPO join where privacy or contractual issues are relevant. The approval package states confirmed facts, hypotheses, unresolved questions, changes, monitoring, rollback criteria, and accountable approvers.
Gate: Limited Rollout
Restore only a constrained scope—such as read-only access, mandatory human approval, a bounded transaction type, one site, or one shift. Do not immediately return every previous privilege. Name the monitoring owner and shutdown authority, and make rollback practical.
Gate: Full Restore
Return to normal scope only after limited operation demonstrates expected behavior, auditability, approval, and rollback. Do not rely solely on the absence of an alert. Test representative normal and abnormal cases. Continue enhanced monitoring after restoration and transfer remaining issues to problem management.

Start the playbook with a one-page action card
A complete manual is necessary, but responders may not read it from the beginning during an event. Place a one-page action card at the entry point and link it to detailed procedures. It should tell personnel to:
- Record the agent, time, screen, request, and observed result
- Use the designated shutdown, rather than repeatedly prompting the agent
- Contain the session, write tools, credentials, and relevant network routes
- Avoid deleting logs, retraining, overwriting settings, or rerunning the same input in production
- Contact the factory manager, OT and safety, IT/SOC, and AI owner
- Contact legal and the DPO immediately when personal data may be involved
- Avoid unapproved external statements, customer notifications, or social posts
- Begin recovery with Sandbox Replay, not full production enablement
The detailed playbook should include architecture, shutdown operations, identity and credential inventory, log collection, RACI, contacts, severity criteria, an impact-assessment form, recovery gates, notification-decision fields, and an after-action format. A document becomes operational only after personnel prove that its controls work.
Implement AI agent recovery drills over a 90-day roadmap
The 90-day period is a planning recommendation, not a legal deadline or standards requirement. If a material exposure already exists, restrict write permissions and establish an interim shutdown before completing the roadmap.
First phase: inventory and containment points
Inventory agents, models, prompts, tools, service accounts, data, plants, and equipment links. For each, identify who can stop it, what stopping it affects, and where evidence exists. Include shadow use and department-built automations.
Walk through the kill switch, tool gateway, credential revocation, network isolation, and business-side transaction hold. Document the boundary with machinery safety systems so personnel understand that stopping the AI does not replace a physical emergency stop.
Middle phase: connect evidence and RACI
Test whether the four evidence streams can be correlated on one timeline. An organization may retrieve an Agent Trace yet fail to connect it to an identity event or MES action. Use correlation IDs, job IDs, production orders, equipment IDs, and user IDs to create the bridge.
Review RACI with all parties and assign deputies for nights and holidays. Simulate requesting logs from the vendor: who is authorized to ask, what format is delivered, and how sensitive evidence is exchanged securely?
Final phase: exercise through limited recovery
Choose scenarios that match the deployment: erroneous purchasing, external transmission, repeated execution, privilege overreach, or suspected OT reach. Reveal facts progressively so participants practice escalation, reclassification, legal review, customer communication, and restoration decisions.
Move beyond the tabletop in a non-production environment. Execute shutdown, credential revocation, evidence capture, Sandbox Replay, approval, limited restart, and rollback. Any production-affecting test remains subject to change management and safety procedures. Assign each lesson learned to an owner and completion target.
Put AI agent incident response requirements into the RFP and contract
After an incident is too late to discover that a trace was not retained, export is not included, a subcontractor holds the evidence, or the format is unusable. Ask vendors the following before award.
Shutdown and containment
- Can the customer stop new sessions, running jobs, one tool, or one site independently?
- Can the customer administrator act directly, or is vendor support required?
- Can sessions and credentials be revoked immediately?
- Can egress, MCP, and API destinations be allowlisted?
- Is the shutdown itself auditable?
Evidence and observability
- Are model inputs and outputs, tool arguments and results, approvals, delegation, errors, and retries recorded?
- Can traces be exported in a machine-readable format?
- Can the customer configure retention, deletion, region, encryption, and access controls?
- Can model, prompt, tool, and agent versions be reconstructed?
- Can correlation IDs flow to the customer’s SIEM, identity platform, and MES logs?
Incident support and accountability
- What are the intake, escalation, technical support, and update commitments?
- Under what conditions does the provider notify the customer of a serious event?
- Who retrieves logs held by subprocessors or external tools?
- Who leads investigation, recovery, corrective action, and customer or regulator support?
- How are data, logs, and secrets returned or deleted after emergency termination?
Recovery and exercise
- Can a historical session be replayed in a sandbox?
- Can the agent return in read-only, human-approved, or limited-scope mode?
- Are rollback and configuration differences visible?
- Will the vendor participate in exercises and evidence-delivery tests?
- Will the vendor share corrective actions and closure evidence?
OpenAI’s Agents API announcement describes environment choices that include a hosted sandbox, customer infrastructure, and partner environments. Availability of an option is not itself a safety guarantee. Select and contract for the environment that fits the buyer’s data, secret-management, network, and evidence requirements.
Connect response to prevention, acceptance testing, and audit
Incident response is a separate capability, but it should connect to adjacent controls. For preventive permissions, sandboxing, and tool boundaries, see Industrial AI agent security. For testing shutdown, authority, and abnormal scenarios before release, see AI agent acceptance testing. For monitoring after recovery, see continuous AI agent evaluation and audit.
Keeping these activities distinct improves ownership. Prevention lowers likelihood, acceptance testing supports a release decision, and continuous audit detects change. When an event still occurs, the incident playbook contains it, preserves evidence, assesses impact, and restores service safely.
Conclusion: stop, preserve, assess, and restore in stages
Effective AI agent incident response depends less on asking the model to behave correctly and more on controlling it from the outside. Stop execution, revoke credentials, preserve four evidence streams, and assess both digital and plant impact. Restore through Sandbox Replay, Human Approval, Limited Rollout, and Full Restore. Align the factory manager, OT and safety, IT/SOC, AI owner, legal/DPO, vendor, and executives through RACI. Thailand’s 72-hour PDPA reference belongs to personal data breach notification, not every AI incident. A playbook becomes real only when a drill proves that the shutdown and evidence paths work.
TOMAS TECH can help Thai manufacturing operations inventory agent connections, design containment and evidence collection, define RACI, conduct recovery exercises, and convert these needs into RFP requirements. Early-stage discussions are welcome through our contact page.
FAQ: What should an AI agent emergency shutdown stop?
It should stop new intake, affected sessions, queued work, write tools, associated credentials, and—when required—network egress. It does not replace machinery safety controls. OT personnel should verify plant state, while the protected monitoring path remains available wherever possible.
FAQ: Is conversation history enough for AI agent log preservation?
No. Preserve the Agent Trace, Tool Call evidence, Identity Log, and OT Event records and correlate them. Protect originals, record collection time and collector, verify time synchronization, and avoid placing plaintext secrets in the evidence store.
FAQ: Must every AI agent accident be reported within 72 hours in Thailand?
No. The 72-hour reference relates to personal data breach notification under Thailand’s PDPA. Legal counsel and the DPO should assess the current law, regulator guidance, affected data, and specific facts. Other AI incidents may trigger different obligations.
FAQ: Should an AI agent recovery drill be performed in production?
Begin with a tabletop exercise and a non-production environment. Test shutdown, credential revocation, evidence capture, replay, approval, limited restoration, and rollback. Any production-related test should follow change management, equipment safety, and production planning.
FAQ: Can service be restored when the precise root cause is unknown?
Sometimes a single cause cannot be proven. Document confirmed facts, open questions, and residual risk. Consider only a tightly restricted restoration with reduced permissions, human approval, stronger monitoring, and a practical rollback. Do not confuse inability to reproduce with proof of safety.
References
- Anthropic: Improving our alignment and security practices
- Anthropic: An alignment assessment of recent cybersecurity incidents
- OpenAI: Introducing the Agents API
- OpenAI Developers: Tracing
- NIST Workshop on AI Incident Management
- NIST AI 600-1
- NIST SP 800-61 Rev.3
- ETDA: Driving Trust AI Governance
- Thailand PDPC: GPPC PLUS