An AI audit is not a one-time pass/fail exercise before launch. A production AI agent acquires a different risk profile whenever its model, system prompt, connected tools, permissions, reference data, or operating procedure changes. The real requirement is not merely to “measure model accuracy periodically.” It is to show which business scenario, configuration and evidence were evaluated, by whom, and who reauthorized the resulting version. This guide explains how companies operating or procuring AI agents in Thailand can connect continuous evaluation, evaluator independence, red teaming, evidence and release gates in one operating process.
Executive answer: audit the business configuration, not the model alone
A model name cannot reproduce the audit scope. The same model presents very different consequences in a read-only FAQ and a procurement agent that can create an ERP purchase-order draft. The same connector changes its risk when permission moves from read to update. A new source-data snapshot can alter retrieval, justification, language handling and exceptions even when the model is unchanged.
Use the following controlled evaluation unit:
business scenario × tool × permission × data × model/prompt version
Treat that combination as one auditable configuration. Record the difference between the approved baseline and the candidate release. The difference then determines whether production monitoring is sufficient, targeted regression is required, or an independent external assessment should be repeated. This prevents shortcuts such as “the model did not change” when the tool scope changed, or “it was only a prompt edit” when an important exception disappeared.
Independence is also not a binary question of whether an outside firm was hired. Outsourcing every daily update is impractical, while letting developers approve every high-impact change creates an obvious conflict. Allocate three lines according to risk: developer self-evaluation, second-line challenge by business and risk functions, and independent third-party assessment.
Why an accepted AI agent must be evaluated again
Acceptance testing remains essential, but it establishes that a named configuration met requirements under defined data and test conditions at a point in time. It does not prove that later versions are equivalent.
A provider may update model behavior or routing. A prompt cleanup may remove an edge condition. A connector release may expose an additional operation. A role edit may expand the population the agent can read. A new product master or operating procedure may alter retrieval evidence. Each change can appear small in isolation but change the agent’s effective business capability when combined.
The NIST AI RMF Core frames risk management as continuous throughout the lifecycle and says AI systems should be tested before deployment and regularly in operation. Measure 2.4 covers production monitoring of the system and its components. The NIST Playbook also identifies changed operating conditions, data drift and model drift as reasons to reassess whether metrics and controls remain appropriate.
Therefore, do not rely only on a calendar that reruns an unchanged test suite. Combine scheduled review with event-driven evaluation. At minimum, begin change-impact analysis when there is a change to:
- the model, model settings, routing or fallback;
- the system prompt, guardrail, policy or retrieval template;
- a tool, API, MCP server, connector or execution environment;
- a role, scope, approver, site or data population;
- training data, RAG data, product master, policy or procedure;
- a material wrong answer, unauthorized attempt, failed stop, evidence gap, complaint or near miss; or
- the business scenario, affected people, jurisdiction or outsourced provider.
The seven gates for AI agent acceptance testing address the pre-release decision. This article starts after that point: preserve the accepted baseline, retest what the change can affect, and reauthorize the candidate configuration.
Build a configuration register that can reproduce the target
Continuous evaluation needs a register that connects each result to the production configuration it represents. Its purpose is not to create another generic asset inventory. It is to make the test and authorization traceable.
Business scenario
“Customer service” or “procurement” is too broad. Define the starting condition, inputs, required outcome, forbidden actions, exceptions, approval point and end state. A testable scenario might be: “read an approved request for quotation, identify candidate suppliers and prepare a purchase-order draft, but do not send it.” That scope makes both success and unauthorized action observable.
Tool
Record the tool version, provider, allowed operations, input and output schemas, target, credential type, timeout, retry and fallback. Even if the connector code does not change, a revised function description exposed to the agent can change tool selection and therefore belongs in scope.
Permission
Do not simply copy a human role name. Specify read, create, update, send and approve capabilities; target tenant, site, customer and state; any value or time boundaries; expiry; delegated execution; and approver. Pair every permitted success case with a denial case outside the authorized boundary.
Data
Record dataset or index identifiers, snapshots, extraction time, language, scope, owner and quality checks. If production data cannot be copied into an audit environment, preserve test-input hashes, retrieval conditions, source-document identifiers and masking rules so that the evidence remains reproducible without exposing the whole dataset.
Model and prompt version
Put the model identifier, delivery mode, settings, system prompt, tool descriptions, routing, memory policy, output validation and policy version into one configuration manifest. Avoid mutable labels such as “latest model.” The identifier must allow the tested and deployed configurations to be compared.

Determine retest scope from the actual difference
Full retesting after every edit can paralyze delivery; no retesting hides material risk. Use the change itself as the input to a structured impact decision:
- Identify every changed configuration element.
- Trace the affected business scenarios, tools, permissions and data.
- Identify the related failure modes and existing controls.
- Select the required regression, red-team, business and evidence checks.
- Assign the appropriate evaluator line and reauthorizer.
- Set production monitoring and rollback conditions.
Do not classify change size by lines of code. One new permission can allow a previously impossible send action. A wording correction may remain limited if it cannot affect execution, evidence or a decision. Judge the difference against forbidden boundaries, impact radius, detectability and recoverability.
Ask these questions during classification:
- Can the agent perform a new category of action?
- Has the readable data, writable target or approval path expanded?
- Has the potential effect on people, customers, finance, quality, safety or compliance changed?
- Will existing denial tests, stop procedures and rollback still work?
- Do previous test data and metrics still represent the new operating context?
- Can an incident still be reconstructed from request through tool result?
If the answer is unknown, do not label the change “minor.” Design the evaluation needed to resolve the uncertainty.
Design independent AI assessment as three lines

NIST AI RMF Measure 1.3 involves internal specialists who were not front-line developers and/or independent assessors in regular assessment. NIST Dioptra’s design principles describe first-party testing during development, second-party testing during acquisition or in an evaluation lab, and third-party testing for audit or compliance. Dioptra also emphasizes reproducibility and traceability. The following three-line arrangement translates those ideas into an enterprise workflow; it is a TOMAS TECH operating synthesis, not a NIST-mandated organization chart.
First line: developer and operations self-evaluation
The team that understands the change performs component tests, scenario regression, tool-boundary checks and log validation. Its advantage is speed and technical context. Its limitation is that it can miss its own assumptions and is exposed to schedule and availability pressure. A high-impact change should not enter production solely because the team that built it says it passed.
The first line should submit more than a success chart: the change diff, impact analysis, cases and inputs, expected and actual results, failure logs, exclusions, known limitations, residual risks and rollback procedure should form one evaluation package.
Second line: business, risk and quality challenge
The second line does not reimplement the system. It challenges whether business success was defined correctly, forbidden behavior was tested, thresholds were not moved for convenience, and evidence can reconstruct production behavior. Select participants from business ownership, information security, quality, legal, compliance or internal control according to the scenario.
The goal is not to collect signatures. It is to create independent falsification. Test ambiguous requests, conflicting instructions, stale data, permission boundaries, revoked approval, tool failure, long-running work and language switching that the happy-path development suite may not cover. Track responses to findings and carry unresolved items into the gate decision.
Third line: independent external assessment
Use external assessment where the scenario can cause material external impact, has high autonomy, is regulated or strategically important, follows a serious incident, or requires specialist capability the company does not possess. Hiring a third party does not automatically create independence. The RFP must define who appoints and pays the evaluator, what they can access, where they report, whether results can be suppressed, how conflicts are handled, and who verifies remediation.
On 18 September 2026, Anthropic announced an embedded-evaluation partnership with Accenture, led by Faculty, that includes model evaluation, red teaming, alignment assessment and safeguard testing. Anthropic and Accenture each expect to invest at least USD 1 billion over the next five years to build capacity in this area. Anthropic also says embedded evaluation is new, that standards for evaluator access and reporting do not yet exist, and that there is no settled funding model. This is timely evidence of investment in independent challenge, not a finished universal standard. The developer remains accountable even when an independent evaluator participates.
Allocate the three lines according to change risk
Define the allocation rule before a change request arrives so that a project cannot reduce scrutiny after seeing a difficult result.
| Change or condition | First line | Second line | Third-line consideration |
|---|---|---|---|
| Presentation only; no effect on capability, evidence or logging | Required | Sample review | Normally unnecessary |
| Prompt, retrieval or model-setting change | Required | Recheck scenarios and metrics | Use for high-impact cases |
| New tool, write operation or data population | Required | Independently check permission, business outcome and evidence | Use for material impact |
| Approval bypass, unauthorized action, critical error or failed stop | Root-cause and fix verification | Required | Formally determine need; normally strong case |
| Regulation or contract requires independent assurance | Required | Required | Perform to the defined scope |
This is an operating example, not legal advice or a universal classification. Tailor it to industry, jurisdiction, impact and controls. If third-party evaluation is not used, document why and who approved that decision.
Replace the single “accuracy” score with six dimensions
One aggregate accuracy score can hide the failure that matters. A correct answer is not acceptable if the agent invoked a forbidden tool. A safe stop is insufficient if no investigation trail exists. Keep at least six dimensions separate.
1. Business success rate
Measure whether the scenario’s defined end state was achieved: required inputs checked, valid evidence used, approval obtained and the intended business state reached. Distinguish partial success, handoff to a person, timeout and duplicate execution. Version the scenario set and include normal, exceptional, multilingual, missing-data, conflicting-instruction and long-running cases. Do not let a simple average hide critical scenarios.
2. Unauthorized-action rate
Measure attempts to use a forbidden tool, operation, target, dataset, time window or approval route. Observe both actions that escaped and attempts blocked by policy. A high number of blocked attempts may show that the control works, but it can also reveal a broken prompt or workflow.
Include indirect instructions in retrieved documents, ambiguous natural language, prior conversation, recently changed permissions, expired approval and identifiers from another site or customer. Define thresholds by impact, and never average away a prohibited high-consequence action.
3. Critical-error rate
Separate errors that may materially affect customers, finance, quality, safety, law, privacy or contract from stylistic differences and minor omissions. Define severity, adjudicator, evidence and appeals before the run. A small number of critical errors may still justify HOLD or STOP where detection is weak and recovery is hard. Where a deterministic check or human approval reliably blocks an error, include that control evidence as well.
4. Stop and recovery
Test detection, suspension of new work, handling of in-flight work, credential revocation, queue isolation, manual fallback, state reconciliation and restart authorization. This is broader than terminating a process. Long-running agents may leave partial artifacts, external side effects, retries and duplicate transactions.
OpenAI announced the Agents API on 10 September 2026 as a public beta. It is one example of infrastructure designed for long-running agents that manage context, tools, subagents and persistent work, but its public-beta status is an important maturity limitation. The customer claims on the launch page are marketing examples, not audit benchmarks or general performance guarantees. The relevant lesson is that an agent can be an extended system rather than one model call, so its evaluation unit must include the environment and tools. See our Agents API operations guide for the runtime layer.
5. Evidence completeness
Determine whether each test and production event can correlate the request, inputs, retrieved sources, configuration, policy decision, tool call, approval, outcome, exception, stop and recovery. Evidence completeness does not mean copying every secret or personal record into a log. Define minimum evidence, masking, access, retention, clock synchronization, tamper detection and export.
NIST Dioptra describes reproducibility through resource snapshots and traceability through a history of experiments and inputs. Whether or not Dioptra is selected, those properties can become procurement requirements: reproduce the tested configuration and trace who ran what.
6. Post-change regression
Verify not only that the new feature works, but that previously accepted critical behavior, refusals, stopping and logging remain intact. Use the changed component to select dependent scenarios, and add every material incident and near miss to the permanent regression set. Preserve the difference from the previous version; an unchanged total can conceal improvement in one critical case and deterioration in another.
No universal threshold is prescribed here. Approve thresholds before testing based on scenario impact, risk tolerance, contracts, applicable law and controls. Structural dimensions should not be converted into arbitrary proprietary scores.
Integrate AI red teaming into continuous evaluation
AI red teaming should not be a single pre-launch event. Update the threat hypotheses whenever capability, connections or threats change. Regression asks whether known requirements still hold; a red team searches for an unexpected path across a boundary.
Include indirect prompt injection, ambiguous tool descriptions, permission confusion, memory poisoning, forged provenance, approval fatigue, state changes during long execution, manipulated tool responses, unsafe fallback and responsibility gaps between multiple agents. Evaluate detection, containment, alerting, evidence and recovery—not just attack success.
ETDA’s AI 2026 direction, published on 9 June 2026, highlights AI Governance Testing and Thailand’s first Red Teaming Challenge. ETDA reports 12 governance guideline/toolkit sets ready for use and another two in development during 2026: an AI Ethical Impact Assessment Playbook and AI Value Creation. This shows ecosystem direction in Thailand, not a general legal obligation for every Thai company. The announcement emphasizes finding weaknesses before real-world use; this article extends the discipline into a company’s post-go-live operating process.
Create evidence during execution, not on the eve of an audit
Screenshots assembled just before an audit cannot reconstruct the deployed configuration. Generate a machine-readable manifest and a human-readable decision record with every evaluation.
Configuration manifest
- business-scenario identifier and version;
- model, settings, prompt and policy identifiers;
- tool, API, connector and execution-environment versions;
- permission scope, credential scope and approval path;
- data snapshot, retrieval conditions and scope; and
- test harness, evaluator logic and decision-rule versions.
Run evidence
- case identifier, input, expected outcome and actual outcome;
- source evidence and tool-call chain;
- policy decision, approval, refusal and exception;
- timestamps, correlation identifier and evaluation environment;
- separation of runner, reviewer and approver; and
- failure, remediation, rerun and unresolved finding.
Decision record
- applied thresholds, approval date and threshold owner;
- results across all six dimensions and material differences;
- known limitations, residual risk and compensating controls;
- reason for GO, CONDITIONAL GO, HOLD or STOP; and
- production monitoring, expiry, next evaluation and rollback conditions.
Evidence may contain sensitive information. Where it is provided to an external assessor, define the review location, masking, export, retention, deletion and subcontracting rules in the RFP and contract. Independence requires sufficient access to verify claims, not unrestricted copying of data.
Connect the result to GO, CONDITIONAL GO, HOLD or STOP

An assessment that only produces a report is disconnected from release control. Define four outcomes.
GO
The named configuration meets approved thresholds, has no unresolved material finding, has complete evidence, and has monitoring and recovery ready. GO is not permanent permission. Limit it to the identified version, scenario, permission scope and authorization period.
CONDITIONAL GO
Residual risk is explicit and can be kept within tolerance by time-limited compensating controls, reduced scope, extra monitoring or human approval. Record each condition, owner, deadline and closure test. A condition must not silently become permanent approval.
HOLD
Evidence is missing, the run is not reproducible, a material defect remains, a threshold is missed, or required independent review is incomplete. Stop release and state what must be corrected and rerun before a decision is possible.
STOP
A prohibited boundary was materially breached, residual risk is unacceptable, stop/recovery is ineffective, or evidence integrity has been lost. Stop the affected configuration and execute safe degradation, credential revocation, manual fallback and incident response as applicable.
What an AI audit RFP should require
Objective and evaluation unit
State whether the engagement covers a model or the full business system. For continuous agent evaluation, specify the scenario × tool × permission × data × model/prompt version, change-triggered regression and reauthorization.
Independence and conflict management
Ask whether the evaluator also develops, sells, integrates or operates the system. Define team separation, reporting line, authority over negative findings and remediation verification. “Third party” is a label; governance and incentives determine practical independence.
Access and data protection
Define required access to models, prompts, tools, logs, test environments and production observations. Address confidential and personal data, cross-border transfer, subcontractors, storage, deletion, incident notice and ownership of deliverables.
Method and reproducibility
Require the method for case design, distinction between regression and red teaming, sampling, severity, reruns, evaluator disagreement, tools and version control. If access is black-box only, require a clear statement of what remains unverified.
Six dimensions and release gates
Require separate reporting of business success, unauthorized action, critical error, stop/recovery, evidence completeness and regression. Do not allow one composite score to offset a critical failure. Define decision rights and escalation for GO, CONDITIONAL GO, HOLD and STOP.
Deliverables and remediation verification
Request the manifest, case catalogue, run records, failure evidence, reproduction steps, findings, severity, remediation and rerun results—not only an executive slide deck. Contract for follow-up questions, evidence retention and transfer of new regression cases.
Standards and legal scope
Where an ISO/IEC 42001, NIST AI RMF or legal crosswalk is requested, require the assessor to identify the source text and scope actually reviewed. ISO’s official overview describes ISO/IEC 42001 as a Plan-Do-Check-Act management system for continual improvement and mentions traceability, transparency, risk assessment and audit schemes. We did not review the paid standard text for this article and therefore do not invent clause numbers. For the wider AIMS program, see our ISO 42001 implementation guide.
Reading EU AI Act Article 72 from Thailand
Article 72 in the consolidated English EU AI Act text dated 27 July 2026 requires providers of in-scope high-risk AI systems to establish and document a post-market monitoring system proportionate to the technology and risk. The system actively and systematically collects, documents and analyses relevant performance data throughout the lifetime to evaluate continuous compliance. It is based on a post-market monitoring plan forming part of the technical documentation.
This is a concrete legal example of lifecycle monitoring. It does not mean Article 72 directly applies to every Thai company or every AI agent. A company must examine EU-market activity, its role, high-risk classification and product/service chain and obtain legal advice where appropriate. Even outside scope, the lifecycle, documentation and continuous-compliance concepts can inform internal control.
Operating workflow from change request to reauthorization
- Register the change. Record not only the technical edit but its possible effect on scenarios, capability, data, permissions, users and jurisdiction. Give emergency changes an expiry and post-review requirement.
- Compare with the baseline. Diff the approved manifest against the candidate across model, prompt, tool, permission, data and evaluation logic. Add manual differences for procedures and contracts.
- Assign impact and evaluator lines. Review impact, detection, recovery and controls; decide which lines must participate and record the reason.
- Freeze the test package. Approve cases, inputs, expected results, thresholds, severity and environment before the run. Log any later exclusion.
- Run regression and red teaming. Separate change-focused checks, critical permanent regression, incident-derived cases and exploratory adversarial work.
- Assess the six dimensions. Report critical boundaries, errors, failed stopping and evidence gaps independently from averages.
- Make the gate decision. Separate the approver from the runner and record the authorized configuration, expiry, monitoring and rollback conditions.
- Feed production evidence back. Convert failures, interventions, unauthorized attempts and user feedback into the risk register and future cases.
Common failure modes
- Auditing the model name and missing tool, permission, data and prompt changes.
- Allowing a high success rate to offset a critical unauthorized action.
- Outsourcing the audit before the company defines success and forbidden behavior.
- Creating logs at audit time rather than capturing the deployed decision chain.
- Treating red teaming as an annual event while the attack surface changes weekly.
- Letting a conditional approval live forever without owner, deadline and closure test.
- Leaving assessor access ambiguous, creating either unverifiable results or excessive data exposure.
- Moving thresholds after results are visible, destroying comparability.
Practical checklist
Program design
- [ ] Define the evaluation unit as scenario × tool × permission × data × model/prompt version.
- [ ] Define scheduled and event-driven triggers.
- [ ] Approve three-line allocation and conflict rules.
- [ ] Approve definitions, denominators, severity and thresholds for all six dimensions before testing.
- [ ] Define authority and exceptions for GO, CONDITIONAL GO, HOLD and STOP.
Per-change evaluation
- [ ] Identify the difference from the approved baseline and affected scenarios.
- [ ] Pair permitted success cases with denial cases.
- [ ] Add incidents, near misses and complaints to regression.
- [ ] Update red-team hypotheses for the changed capability.
- [ ] Test stop, recovery, rollback and retry side effects.
- [ ] Reconstruct the chain from input through tool result.
Reauthorization
- [ ] Separate evaluator and approver roles.
- [ ] Record unresolved findings, residual risk and compensating controls.
- [ ] Name the authorized version, scope and expiry.
- [ ] Give conditional approval an owner and closure deadline.
- [ ] Register production monitoring and the next evaluation trigger.
FAQ: AI audit and continuous AI agent evaluation
How often should an AI audit be performed?
No single interval fits every system. Use scheduled review plus triggers for model, prompt, tool, permission, data, business or jurisdiction changes and incidents. Increase monitoring and reassessment where impact and change frequency are high. Document the rationale from risk tolerance, contracts, regulation and operating pace.
How is continuous evaluation different from acceptance testing?
Acceptance testing approves a named pre-release configuration. Continuous evaluation preserves that baseline, uses production observation and change differences to select regression, red-team and evidence checks, then reauthorizes the candidate. They are consecutive controls, not alternatives.
Is an independent third-party AI assessment always required?
Not for every change. Decide from impact, autonomy, regulation, contract, incident history, internal capability and conflicts. If no third party is used, preserve second-line challenge and the reason for the decision. If one is used, define access, reporting, conflict management and remediation verification.
Does AI red teaming complete an AI audit?
No. Red teaming is valuable for discovering unexpected attack paths, but it does not replace business-success testing, critical-error review, stop/recovery, evidence completeness, regression or authorization. Convert useful red-team findings into permanent regression cases.
Can improved accuracy justify reauthorization?
Not by itself. Check unauthorized action, critical error, stop/recovery, evidence and regression separately. A more accurate model can still call a forbidden tool, weaken a refusal boundary or lose audit evidence.
What should an AI audit RFP clarify first?
Clarify the evaluation unit and independence. Define whether the assessor tests a model or the business system, exactly which versions, data and permissions are in scope, and whether the assessor also develops or sells the solution. Then specify method, evidence, data protection, six dimensions, gates and remediation verification.
Does ISO/IEC 42001 certification remove the need for system-level continuous evaluation?
No. A management system supplies a continual-improvement framework, but a particular scenario and changed configuration still need tests and evidence. Also compare the certification scope with the actual AI systems in use.
Summary: connect AI audit to change control and release authority
An AI agent’s audit target changes when its model, prompt, tools, permissions or data change, even after it passed acceptance. Fix the evaluation unit as business scenario × tool × permission × data × model/prompt version and select regression from the actual difference. Allocate development, business/risk and external-assessor lines according to impact. Evaluate business success, unauthorized action, critical error, stop/recovery, evidence completeness and post-change regression separately. Connect the result to GO, CONDITIONAL GO, HOLD or STOP and record the authorized version, expiry, monitoring and rollback. That is what turns an audit report into an operating control.
TOMAS TECH can help structure an AI audit and continuous-evaluation program from the RFP, configuration register and change triggers through test cases, evidence packages and reauthorization. A review can begin while you are still deciding how to extend an existing acceptance-testing or ISO 42001 program.
References
- Anthropic, “Partnering with Accenture on embedded evaluation” (18 September 2026): https://www.anthropic.com/news/accenture-embedded-evaluation
- NIST, AI Risk Management Framework Core — Measure: https://airc.nist.gov/airmf-resources/airmf/5-sec-core/
- NIST, AI RMF Playbook — Measure: https://airc.nist.gov/airmf-resources/playbook/measure/
- NIST, Dioptra Design Principles: https://pages.nist.gov/dioptra/explanation/dioptra-motivation/design-principles.html
- ISO, AI management systems: What businesses need to know: https://www.iso.org/artificial-intelligence/ai-management-systems
- European Union, AI Act consolidated text, Article 72 (27 July 2026): https://eur-lex.europa.eu/legal-content/EN/TXT/PDF/?uri=CELEX%3A02024R1689-20260727
- ETDA, AI 2026 — Driving Trust AI Governance (9 June 2026): https://www.etda.or.th/th/pr-news/aigc_Driving-Trust_AI_Governance.aspx
- OpenAI, “Introducing the Agents API” (10 September 2026): https://openai.com/index/introducing-the-agents-api/
This article is general operational guidance based on public information checked on 19 September 2026. It is not legal advice, certification or a guarantee.