Blog

2026.09.20

AI Agent Observability: Monitoring, SLAs and RFP Guide

AI Agent Observability: Monitoring, SLAs and RFP Guide

AI agent observability is the ability to reconstruct not merely whether an answer appeared, but how a business request moved through models, data, tools and approvals to reach an operational outcome. In manufacturing, a conversation log is not enough. Equipment, product, lot, work order, shift, plant time, tool execution, operator corrections and cost must be connected through a correlation ID and converted into evidence that supports anomaly detection, diagnosis, SLA review and improvement. This guide covers telemetry design, RFP requirements, a proposed 90-day rollout and acceptance for multi-site operations, including factories in Thailand.

The short answer: observe the whole business session, not only the model

An LLM API response time does not reveal the operational quality of an AI agent. An agent combines retrieval, planning, tool selection, ERP or MES queries, approval waits, retries and handoffs. A final answer may look correct even though the agent read prohibited data, queried the wrong asset, executed one request twice or required the operator to make the same correction every time.

The minimum observation unit is therefore not one model call but a session with a business purpose. Place one or more traces below the session, then represent model generation, retrieval, tool calls, approvals, guardrails, external APIs and human intervention as spans. The official OpenAI Agents SDK documentation similarly describes a trace as an end-to-end workflow made up of spans that have start and end times and parent relationships. This is a useful implementation example, not a requirement to copy one vendor’s format as an internal standard.

This article focuses on telemetry and operational measurement. Our guides to AI agent API operations, continuous AI agent evaluation and audit and AI agent incident response cover platform operations, governance/reapproval and post-incident containment respectively. Observability is the cross-cutting layer that supplies time-ordered evidence to all of them.

The basic AI agent tracing model

Session: business purpose and user context

A session is more than time spent in a chat window. It represents a purpose such as “investigate the cause of Line 3 downtime and prepare a maintenance request” or “identify a parts shortage and place a purchase request into approval.” Record a business case ID, site, user role, start/end, outcome, plant business date, data classification and policy version.

If a case continues the next day, connect it with the business ID even if the application creates another conversation. If one chat handles several work orders, give each work order its own trace. Separating UI convenience from the evidence unit is essential for later cost, quality and accountability analysis.

Trace: the execution path to one outcome

A trace is the path from a business request to an outcome. It should identify the session, initiating actor, agent version, model configuration, start/end, status, outcome and related external transaction. When work is delegated to another agent, retain a parent trace or trace link so that the complete path remains searchable.

Do not claim that tracing exposes a model’s complete reasoning. What can be observed is the permitted input/output, retrieval source, tool request, policy decision, approval, external response, timing and usage. Mixing observable facts with assumptions about hidden reasoning weakens an audit.

Span: the operation used to isolate delay, failure and responsibility

A span represents an operation with a start and end: model generation, RAG retrieval, data transformation, tool execution, approval wait or human override. At minimum, retain span ID, parent span ID, operation, timestamps, status, error type, attempt, service, site, tool/model and input/output reference.

More spans are not automatically better. Excessive granularity increases storage and investigation effort; insufficient granularity hides the cause. A practical split occurs when the responsible service changes, the recovery method changes, the charging unit changes, or an approval/external side effect occurs.

AI Agent Observability: Monitoring, SLAs and RFP Guide - figure 1

Connect factory context to evidence

Physical-process context separates factory-agent monitoring from generic chatbot analytics. “Abnormal” means different things for different assets, products, lots, recipes, modes, shifts, sensor freshness and maintenance states. Link traces to:

  • site, area, line and asset IDs;
  • work order, batch, lot and product code;
  • shift, UTC and local plant time with time zone;
  • recipe/version and PLC, MES or ERP transaction ID;
  • sensor snapshot, acquisition time, quality flag and missing-data state;
  • model, prompt, policy, tool, connector and knowledge-base versions;
  • operator or pseudonymous identity, role and approval ID; and
  • final business outcome plus the later quality, downtime or delivery record used to verify it.

Do not copy all sensor data and documents into trace storage. Keep the snapshot in an evidence repository and place its hash, timestamp, schema version and access-controlled reference in the trace. This avoids turning telemetry into another uncontrolled confidential-data lake while preserving reconstruction.

Clock quality also matters. If cloud services, terminals, MES, PLCs and cameras disagree, one event may appear before its cause. Record the time source and synchronization state. Use UTC as the common axis and show ICT or another local time to operators. Explicit uncertainty is stronger evidence than a falsely precise sequence.

Design LLM observability metrics in five layers

1. Use: who uses what, and for which business purpose

Count sessions, active users, business process, site, shift, agent version and completion state. Volume alone is not success. Combine it with quality, elapsed time, human effort and business outcomes. A spike may indicate adoption, but it may also reveal a retry loop or user-interface fault.

2. Reliability: completion, error, retry and duplicates

Separate errors, timeouts, retries, partial completion, duplicate side effects, dead letters, tool denials, fallbacks and abstentions. Safe abstention is not the same as a system error. A rising abstention rate may be an early signal of missing data or out-of-scope requests.

3. Responsiveness: break down the time a person waits

Measure end-to-end latency, time to first response, model, retrieval, tool, approval and queue time. Averages hide peak and long-tail delay, so use percentiles and distributions. Targets must be use-case specific: an advisory quality investigation and an urgent downtime workflow do not need the same response time.

4. Human collaboration: override, approval and handoff

Record approval request/outcome, human handoff, operator edit, override reason, repeated input and unhandled wait. A high override rate may reflect a model issue, but it may also come from poor master data, insufficient permission, a confusing review screen or a mismatch with local terminology. Preserve reason codes and structured diffs, not only free text.

5. Value and cost: connect model charges to business outcomes

Allocate input/output tokens, model calls, retrieval, external APIs, telemetry storage, evaluation and human review to the business case. Proposed internal measures can include cost per completed case, per approved action or per avoided manual minute, in addition to cost per session. Define the baseline and what “avoided” means before measurement, and keep estimates distinct from actuals.

OpenTelemetry’s GenAI semantic conventions include attributes for model, operation, token use, response ID, finish reason, time to first chunk, retrieval and tool calls. They provide an interoperability starting point, but the registry is versioned and attributes can move or become deprecated. The RFP should pin a semantic-convention version and require a mapping to the company’s canonical fields.

Divide the AI agent SLA into three layers

“99.9% availability” does not prove that an agent completes business work. Structure the SLA/SLO in three layers:

  1. Technical service level: API availability, trace ingestion, tool connection, queue processing and monitoring delay.
  2. Agent execution level: session completion, end-to-end latency, errors, abstentions, duplicate side effects and approval wait.
  3. Business outcome level: correct work-order association, operator correction, processing time, deadline completion and quality result.

The following values are examples for design, not universal standards.

MetricExample proposalMeasurement caution
Trace capture for critical work99.5% or higherDefine critical work and whether missing evidence blocks execution
p95 end-to-end latency30 seconds or lessShow values with and without approval wait
Prohibited duplicate write0Reconcile idempotency key with external transaction
Session failure rateBelow 2%Classify abstention, denial and user cancellation separately
Cost anomalyAlert at 30% above baselineFix the baseline period, seasonality and work mix
Override rateReview above 20%Separate useful intervention from poor quality by reason code

Replace the numbers using loss, volume, acceptable delay and supervision. For low-volume work, report counts with rates because one case can dominate a percentage. NIST AI 600-1 warns against over-reliance on quantitative measures without the context and limitations in which they apply. A dashboard is the beginning of a decision, not the decision itself.

Treat tool calls as a ledger of side effects

The most consequential evidence is often not what the agent said, but what it did. Each tool call should record tool name/version, caller, target, operation, argument reference, policy decision, approval, idempotency key, request/response state, external transaction, retry and compensation.

Tool arguments may contain suppliers, prices, assets, personal information or tokens. Do not retain raw arguments indefinitely. Define allow, mask, hash, drop or secure-reference rules by field. Asset ID may remain searchable, free-text comments may require classified storage, secrets must never be logged, and purchase amounts may need restricted viewing.

Make approval a separate span and compare the proposed payload hash with the executed payload hash. If target, quantity or asset state changes after approval, do not reuse the approval. Rejections should retain reason, wait time and alternative treatment because they reveal permission and UI problems.

AI Agent Observability: Monitoring, SLAs and RFP Guide - figure 2

Do not merge error, abstention and override

Operational states should distinguish:

  • system error: service, network, authentication or schema failure;
  • task failure: the business task could not be completed;
  • policy denial: the policy correctly blocked the action;
  • abstention: the agent withheld an answer/action due to weak evidence or scope;
  • human override: a person edited or cancelled the proposal;
  • user abandonment: the user left while waiting; and
  • partial completion: only part of the work finished.

If all states count as failure, safe refusals become a bad KPI and teams are encouraged to reduce them. If all abstentions are considered safe, an agent can hide its inability to work. Define an expected range per use case and combine reason codes with sample review.

Decide redaction, minimization and retention first

The OpenAI Agents SDK documentation explains that generation spans may retain model input/output and function spans may retain tool input/output, and provides settings to disable sensitive-data capture. Microsoft Foundry documentation likewise warns that inputs, outputs and tool arguments/results may be sensitive; it advises keeping secrets out, redacting/minimizing sensitive data before telemetry, and applying production-grade access and retention controls.

Use four control points:

  1. Before collection: do not pass secrets through prompts or tools; inject them at a broker if needed.
  2. At ingestion: detect names, contact details, customers, prices and credential patterns, then drop or tokenize.
  3. After storage: apply encryption, tenant/site separation, role access, audit and expiry.
  4. At display: show summaries to general operations and provide request-based access to confidential detail.

Avoid “keep everything and decide later.” Set retention by investigation need, contract, quality, privacy, storage cost and learning use. For example, 30 days for ordinary spans, 13 months for aggregate data and case-specific preservation for serious events is only a proposal, not a default. Align actual periods with local counsel, customer contracts and company rules.

OpenAI documentation also states that built-in tracing is unavailable to organizations using OpenAI APIs under a Zero Data Retention policy. Verify the actual contract, configuration and exporter. Where built-in tracing cannot be used, provide an internal OpenTelemetry collector or business-event ledger rather than assuming a dashboard equals retained evidence.

Vary sampling by risk and evidence value

One sampling rate allows low-value success logs to consume storage while important failures disappear. A proposed multi-stage policy is:

  • retain prohibited actions, write tools, approvals, overrides, errors, policy denials and critical assets at 100%;
  • temporarily raise new releases, first sites or night shifts to 100%;
  • retain detailed traces for an example 10–20% of routine read-only successes while counting all metrics;
  • use tail-based sampling to keep traces with unusual latency, token use, cost or span count; and
  • preserve representation across process, language, site and shift.

These percentages are example values. Even for sampled-out traces, retain a minimum case ID, result, main metrics and external-transaction reference. Version the sampling rule so that the organization can explain why detail was absent at that time.

A vendor-neutral observability architecture

A practical architecture sends telemetry from the agent runtime to a collector using OpenTelemetry or an equivalent format. The collector normalizes schema, redacts, samples and routes to trace, metric and log/evidence stores. The business database, MES, ERP and quality system are linked through transaction and case IDs. One operational UI may combine them, while original evidence remains separate from aggregate measures.

AWS AgentCore is an implementation example that stores metrics, spans and logs in CloudWatch and offers trace visualization, custom span metrics and error breakdowns. Microsoft Foundry shows span-level review and delivery to Application Insights. Neither example is a product-selection guarantee. Keep the RFP neutral by specifying export format, schema, location/region, encryption, search, cost, outage behavior and portability.

Design collector failure as well. For an important write, decide whether missing evidence blocks execution, goes to a local buffer or permits a degraded mode. Test a full buffer, network loss, replay, duplicate ingestion and clock skew. Use different idempotency keys for business side effects and telemetry delivery.

Questions for an AI agent observability RFP

Data model and correlation

  • How are session, trace, span, business case and external transaction related?
  • How are multi-agent handoff, parallel tools, retry and long-running jobs represented?
  • Which site, line, asset, work order, lot and shift fields are standard?
  • How are schema and semantic-convention versions pinned and migrated?

Collection and confidentiality

  • Are model/tool inputs and outputs retained by default, and can individual fields be excluded?
  • Does redaction occur before transfer or after storage, and what is the fail behavior?
  • How are residency, encryption, customer-managed keys, access, retention and legal hold handled?
  • Who authorizes temporary increased capture for debugging, and how does it expire automatically?

Performance, cost and availability

  • How is telemetry overhead on latency, CPU and network measured?
  • What are peak ingestion, buffer, backpressure and drop policies?
  • Can token, tool, storage, query and egress cost be allocated to a case?
  • Can high-risk writes stop if observability is unavailable?

Operations and evidence

  • How many steps take an operator from alert to trace, tool and external transaction?
  • Are raw export, API, schema, dashboard, runbook and training delivered?
  • Are changes to alerts, redaction, sampling and dashboards auditable?
  • Can evidence and settings be exported in a standard format at contract end?

Evaluate with screens, export samples, failure tests, search time and responsibility—not a “supported” checkbox. Integrate telemetry-loss, correlation, redaction, sampling and cost-reproduction tests into the AI agent acceptance-testing gates. For OT write access, separately satisfy the privilege, execution-broker and safety-layer requirements in our industrial AI agent security guide.

A proposed 90-day implementation roadmap

This is a proposed schedule, not a universal rule.

Days 0–15: define the business ID and operational questions

Limit the first scope to one process and one site. Write the questions operators must answer: Why was it slow? Which tool failed? Who corrected it? What did one case cost? Did the ERP transaction complete? Inventory existing logs, data classification, clocks, transaction IDs and ownership; then define the minimum session/trace/span schema.

Do not begin by polishing dashboards. If the case ID cannot connect the UI, agent, tool and external system, a beautiful graph will not support an investigation.

Days 16–30: implement instrumentation, redaction and time

Add spans for model, retrieval, tool, approval and handoff. Normalize to the internal schema in the collector. Implement field policies for secrets, personal data, customers and prices. Confirm UTC/local time and the clock source, and separate development, validation and production environments.

Days 31–45: connect metrics and alerts to outcomes

Aggregate latency, error, retry, abstention, override, approval, tokens and tool cost, then reconcile with ERP/MES completion. Begin with explicitly proposed thresholds and tune them against a two-week baseline. Give every alert an owner, runbook, severity, acknowledgement expectation and closure condition.

Days 46–60: inject failures and reconstruct evidence

Inject tool timeout, expired authentication, schema mismatch, collector outage, full buffer, expired approval, duplicate retry and clock skew. Confirm that an arbitrary business case can be reconstructed through trace, span, external transaction and operator action. Use negative tests to confirm confidential content does not leak into search or export.

Days 61–75: tune SLO, sampling and cost

Propose use-case-specific p95, error budget, override review and cost anomalies, then approve them using measurements and loss impact. Test full capture for high risk and sampling for low-risk success. Balance storage/query cost, network use and alert workload against evidence value.

Days 76–90: accept and hand over

Complete FAT/SAT-equivalent evidence using day/night shifts, Thai/English/Japanese input, local network conditions and role handover. Confirm that operations can search, form a cause hypothesis, verify redaction, close alerts and export evidence without the supplier. Assign a conditional acceptance with owner, deadline, workaround and retest, or stop production when critical evidence is missing.

AI Agent Observability: Monitoring, SLAs and RFP Guide - figure 3

Make acceptance evidence reproducible

Acceptance should include at least these cases:

  1. Trace one business request from session to external transaction.
  2. Preserve relationships for parallel tools, handoffs and retries.
  3. Identify model/tool version, policy and knowledge snapshot.
  4. Split latency into queue, model, retrieval, tool and approval.
  5. Classify error, abstention, denial, override and abandonment.
  6. Reconcile write argument, approval, idempotency and result.
  7. Keep secrets and designated fields out of trace, log, dashboard and export.
  8. Detect loss/duplicates across collector outage and reconnection.
  9. Preserve a minimum case/result reference for sampled-out detail.
  10. Recalculate case cost with supporting data.
  11. Correctly link UTC/plant time, asset, work order and lot.
  12. Let an operator use the runbook and retain a closure reason.

Link requirement, test, trace, expected/actual result, evidence, reviewer, defect and retest IDs. Retain machine-readable exports and queries, not screenshots alone. The same cases should run after a dashboard or platform change.

Operating differences to verify at a Thailand site

Thailand’s ETDA Generative AI Governance Guideline discusses limitations around domain context, freshness, explainability and inconsistent outputs, and recommends human participation and continued technology updates. It also distinguishes adopter, customizer and maker roles. A plant using a commercial SaaS, one building its own RAG/tools and one developing models do not own the same telemetry scope.

At a Thai site, include local asset names and abbreviations, shifts, local approvers, ICT timestamps and network conditions—not only a Japanese-language head-office policy. Language-level abstention and override rates cannot be compared without considering volume and task difficulty. Human-review representative cases and classify whether translation, master data, UI or permission caused the difference.

Access should also differ. Head office may use aggregates and cross-site comparison, site owners may see local detail, security may request confidential traces, and suppliers may receive a pseudonymized subset. This is a design example; contracts, privacy and labor requirements decide the final model.

Common failures

Saving complete conversations and calling it monitoring

A conversation cannot show tool side effects, approval or external outcomes, while it can greatly increase confidential data. Center the design on structured IDs/events and minimize content.

Watching only average model API latency

Users wait for queue, retrieval, tool and approval as well. An average hides peak and long-tail delays. Show end-to-end and span percentiles.

A dashboard without a defined response

A red chart does not improve a process. Give each metric an owner, threshold, runbook, deadline and closure condition. Removing noisy alerts is also part of operations.

Choosing either full retention or one uniform sample

Treating risky writes and routine reads equally loses either evidence or budget. Combine risk-, tail- and release-based sampling.

Depending on a vendor UI as the only evidence

Evidence may disappear after termination, outage or product change. Contract for raw export, schema, queries, settings, retention and migration.

FAQ: AI agent monitoring and SLAs

How is AI agent observability different from ordinary monitoring?

It adds session, generation, retrieval, tool, approval, human correction, business outcome and cost to CPU, memory and API availability. It extends APM rather than replacing it.

Does tracing reveal the agent’s chain of thought?

Not necessarily. It shows the inputs, outputs, events, retrieval, tools, approvals, policies, timing and versions deliberately collected. Keep observed facts separate from causal hypotheses.

What should be recorded at 100%?

There is no universal answer. High-risk writes, approvals, prohibited actions, errors, overrides and critical assets are candidates for complete structured evidence. That does not mean retaining all raw text.

How should numeric AI agent SLA targets be set?

Use business loss, volume, acceptable delay, human fallback and data quality. Take a baseline and define targets across technical, execution and business-outcome layers rather than copying an external benchmark.

What should be reduced first when logging cost is high?

Remove duplicate content, low-value success detail and unnecessarily long retention. Keep important metrics and correlation IDs; use risk-based sampling, compression and tiered storage. Pre-collection minimization reduces both confidentiality risk and cost.

Does observability make an agent safe?

No. It improves detection, reconstruction and learning. Permissions, deterministic constraints, approvals, safety PLCs, incident response and continuous evaluation remain separate controls. Observability is not control.

Conclusion: follow one business case through evidence and cost

AI agent observability is not a project to produce more logs. It connects session, trace and span to the business case, asset, work order, lot, tool, approval, human intervention, external transaction and cost so that operations can reconstruct a case. Separate errors, abstentions and overrides; break down latency; minimize sensitive content before collection; and sample according to risk. Require schema, export, redaction, retention, outage behavior and portability in the RFP, then validate correlation, fault injection, SLOs and handover within the proposed 90-day plan.

TOMAS TECH can help structure existing plant IDs, permissions and logs into an AI agent observability RFP, 90-day plan and acceptance cases before a product is selected. If you first need to determine which process, evidence and operating owner belong in scope, contact TOMAS TECH.

References