Blog

2026.09.05

AI Agent Workflow Automation: 7 Failure Modes We Hit In-House

AI Agent Workflow Automation: 7 Failure Modes We Hit In-House

AI agent workflow automation is easy to demonstrate. What is hard is keeping it running unattended at a fixed time every day, and staying able to notice when it breaks. At TOMAS TECH we run several AI agents daily across our own Thailand operation, covering email, calendars, meeting recordings, quotations, follow-up tasks and internal notifications. This article documents the seven failure modes we actually hit, and the design rules we adopted because of them, in a form other companies can reuse.

Conclusion: automation succeeds on “being able to tell it broke,” not on model intelligence

Automation does not stop only because a model is inaccurate. In our operation, the events that actually destroyed business value were mundane. A job finished successfully and reported zero records while its connection had already been lost. The same item was filed as a brand-new task every morning. An aggregate reversed its conclusion because internal memos were mixed into the denominator. A configuration existed in two places, and the copy we fixed was not the copy being read.

None of these were “the AI got it wrong.” They came from the area around the agent: input health, record uniqueness, monitoring, and the boundary of human approval. So the design centre for in-house automation is these five points.

  1. Can you distinguish “the input could not be fetched” from “the input was empty”?
  2. Can you define the identifier that prevents writing the same event twice, independently of the implementation?
  3. Does every output carry its denominator and its period?
  4. Is there a line between operations a human approves and operations the agent may perform on its own?
  5. When it breaks, who is told, when, and by what channel?

The sections below explain why each of these became necessary, in the order we hit the failures. If you want to organise the broader adoption path first, see the AI adoption roadmap for companies in Thailand; for the procurement view including outsourcing, see how to select a Bangkok AI development company.

What we automated: nine daily back-office pipelines

To set the context, here is the scope. All of it is internal work, not customer systems.

PipelineInputAgent outputJudgement left to people
Email correspondence logCorporate mail clientAccumulates sent and received mail into a database (internal broadcasts and marketing excluded)Whether a reply is required
Calendar intakeBusiness calendarCreates stub records for today’s and tomorrow’s meetings in the meeting databaseCreating or changing the appointment itself
Meeting record reflectionVoice recorder transcriptsMatches summaries to the meeting record; non-Japanese transcripts are translated and kept alongside the originalChecking factual accuracy
Quotation syncCore system (PostgreSQL)Incremental sync of quotations and daily update of days-openDeciding price and terms
Follow-up extractionAll of the aboveFiles the day’s required actions into the task databasePriority and execution
Reply draftingEmail correspondence logCreates an email draft (never sends)Finalising and sending
Appointment gap checkEmail plus calendarDetects appointments confirmed by email but missing from the calendar, and notifiesAdding it to the calendar
Deadline alertsHR roster dataCalculates who is approaching a deadline and by when notice must be given, then files itEvaluation and personnel decisions
Morning notificationTask databaseSends each owner a numbered list of open tasks over business chatReporting completion
AI Agent Workflow Automation: 7 Failure Modes We Hit In-House - figure 1

The shape is the same throughout: collect → normalise → record → notify. Generative AI mainly does the normalising (classifying, summarising, translating, matching). We deliberately record into the places people already look — the task and meeting databases, business chat, the mail drafts folder — and built no dedicated agent screen. Automation that requires staff to learn a new screen stops being looked at within the first week of operation.

Note also that we have not handed the agent any operation that sends, registers or decides. Email stops at the draft. The gap check stops at detection and notification. The deadline alert stops at calculation and filing. The reasoning behind that line is set out later.

Failure mode 1: a failure disguised as “zero records”

This is the failure we went longest without noticing. The job that fetches and records email kept finishing successfully and reporting “zero records in scope” even though its connection prerequisite had been lost. Every log said success; the notification said “nothing for today.” In reality, nothing had been fetched for fifteen days.

The cause was that the implementation expressed “the result was empty” and “the fetch could not happen” with the same empty array. The agent received an empty array and faithfully reported zero. Nobody lied, and the process was dead for fifteen days.

We now put three things into every job.

  • Record connection health as a signal separate from the record count. Whether the source was reachable, whether authentication succeeded, and whether the period query returned, are all stored independently of a count of zero.
  • Treat a run of zeros as an anomaly. Business email at a site of our size does not fall to zero for two consecutive working days. Keep a consecutive-zero counter and, past a threshold, notify as a warning rather than a success.
  • Always state source status in the completion report. Not “0 records” but “connection OK, window 24 hours, 0 matching.” A human sees the difference immediately.

Batch monitoring is usually designed to fire when something fails. AI agents have the opposite property: they round failures up into successes. Build in dead-man monitoring that treats silence as a fault from day one.

Failure mode 2: when the idempotency key drifts, everything looks new every morning

If the same check runs daily, “do not file an item that has already been filed” is mandatory. We decide this with a hash of a string that uniquely represents the item — an idempotency key.

We got this wrong twice. The separator character and the concatenation order were rewritten each time the key was reimplemented. Change the separator and the hash changes completely. Every item filed on previous days was judged new, and duplicate tasks piled up. The second time we hit it, we froze the key formula as a written specification and forbade implementations from reassembling it themselves.

In reusable form:

  • Write the idempotency key definition in prose, in exactly one place, and have code refer to it. For example: “concatenate the date and the sorted occurrence list with |, then take the first 12 characters of the SHA-1 digest” — the separator is part of the specification.
  • Never mix run time or run environment into the key material. Current timestamp, host name and file path all turn the same event into a different one on every run.
  • Make the deduplication outcome visible. Print “n new, m updated, k unchanged” every run. When the ratio jumps, you see it. If everything is new, you see it that morning.

Failure mode 3: the search you assumed exists does not work in the real environment

A query you assume is trivially available at design time may be unusable in production. In our case it was full-text search over email bodies. The standard filter syntax destabilised the client, the full-text operator was disabled at the environment level, and the asynchronous search API never signalled completion — three constraints at once.

What mattered was accepting the constraint as a premise rather than an exception, and changing the design. We gave up body search, narrowed candidates by subject, sender and recipient addresses, and period, and moved body-level judgement into the agent after retrieval. It processes more data, but it works, which is worth more.

When writing requirements for AI workflow automation, verify the following on the real system. Judging from the catalogue and API reference alone leads to a rebuild late in implementation.

  • Whether the search or filter syntax completes at real data volumes (test with production-scale counts, not a 100-record sample).
  • Whether an asynchronous API signals completion. If polling is required, what the termination condition is.
  • Whether attachments, bodies or custom fields carry permission and licensing constraints.
  • What happens at the rate limit. Confirm retry intervals and ceilings in the provider’s own documentation — Microsoft Graph and the Notion API both publish their limits and expected client behaviour.

Failure mode 4: without a denominator, an aggregate reverses its conclusion

A different problem appeared when we began aggregating the accumulated data to find trends. Trying to count contact frequency per customer across the meeting database, we found the customer-name field empty in nearly half of the records, plus several hundred internal memos that were not customer meetings at all. Summaries existed only for records created after a certain date.

Aggregate that as-is and you cannot tell whether a “rarely contacted customer” is genuinely rare or merely unnamed. An aggregate with no stated denominator is more dangerous than a silent error, precisely because it produces a plausible conclusion.

AI Agent Workflow Automation: 7 Failure Modes We Hit In-House - figure 2

The rules we settled on:

  • Every aggregate states its period, population, exclusions and missing-data rate.
  • Fields whose missing rate exceeds a threshold need a completion mechanism before they are used in aggregation. In our case we auto-complete by normalised matching against the customer master, and never overwrite an existing value.
  • Publish both “the ratio among records that have the field” and “the ratio across all records.” Looking at one alone guarantees misreading.
  • Mark the point at which the meaning of the data changed (a field added, a process changed) as a boundary, and annotate any aggregate that spans it.

This is not unique to generative AI, but agents execute the aggregation they were asked for without doubting it and return the result as assertive prose, so a missing denominator couples easily to an overconfident conclusion.

Failure mode 5: dates and times drift silently between environments

For an operation spanning a Thai site and a Japanese head office, time handling is a permanent source of incidents. We hit three.

First, a shell environment silently ignored the timezone setting and ran in Coordinated Universal Time. There is no error; all that remains is a result dated one day off. Second, mail from counterparts at Japanese sites is written in Japan time while the calendar is kept in Thai time, and people were mentally correcting the seven-hour difference. Third, we back-calculated a past working date from the current run time and disagreed with the actual creation date.

The countermeasures are simple but need to be enforced.

  • Use timezone-aware datetimes throughout the agent’s internal representation. Do not pass naive date strings around.
  • Decide, per destination, the rule that display is local site time and storage is absolute time.
  • Do not infer past dates. Read them from real data — modification timestamps, mail headers.
  • Log the runtime timezone setting at start-up so it can be checked on every run.

Failure mode 6: caches and sync sit between your write and your read

When automation writes into an internal system or your own website, you meet the phenomenon of content you wrote disappearing. We confirmed two variants.

The first is content delivery network edge caching. Read immediately after writing and you get stale content; build the next update on top of it and you overwrite and destroy the change you just made. The fix has two parts: always cache-bust on read, and read back after every write to confirm what is actually stored.

The second occurs when deliverables are built inside a cloud-synced folder. The file is swapped mid-sync, and the finished artefact silently lacks the edits. We changed the practice: run generation in a local working area, verify the artefact, and only then place it into the synced folder.

Both are the kind of event that, without knowing the cause, gets misread as “the agent did not make the change I asked for.” Make read-back-and-compare part of the same verification step as the write.

Failure mode 7: two sources of truth, and the one you fixed is not the one being read

As an operation grows, the same processing gets defined in more than one place. In our case the definition read by the scheduled run and the definition read by manual invocation had been duplicated into separate directories. We fixed one, and only noticed the next morning when the same defect reappeared.

This is not an AI-specific problem, but it hurts more under agent operation. An agent reads the location it was pointed at and executes exactly what is written there, so it will confidently run to completion on a stale definition. A human would stop and think “but I fixed that.” An agent does not stop.

  • Choose one source of truth and make the other reference-only. If a copy is required, treat it as a generated artefact and never hand-edit it.
  • Log which path the definition was read from.
  • After changing a definition, treat it as done only once you have checked the next automated run.

Where to keep human approval: set autonomy in tiers

Given the failure modes above, how far to trust the agent should be decided by the reversibility of the effect, not by intuition. We split operations into tiers and open them up from the bottom.

AI Agent Workflow Automation: 7 Failure Modes We Hit In-House - figure 3
TierExample operationReversibilityCurrent handling
L1 ReadFetching mail, calendar, core system dataNo effectAutomated
L2 RecordAccumulating into a database, adding summariesFixable by editingAutomated
L3 FileCreating tasks, calculating deadlinesReversible by deletionAutomated (idempotency key mandatory)
L4 DraftCreating a reply draftNo effect unless sentAutomated
L5 NotifyNotifications and alerts to ownersIrreversible but low harmAutomated (restricted recipients)
L6 Send / registerSending mail, creating appointments, publishing externallyIrreversibleHuman performs
L7 DecidePrice, personnel terms, contract conditionsIrreversibleHuman performs

The point is the line between L5 and L6. Automate up to notification; keep anything with external effect with a person is a simple enough rule that nobody on the floor has to deliberate. This line also hedges against the fact that generative AI errors appear not as visible misreadings but as plausible assertions. The harder an error is to detect, the earlier the approval gate has to sit.

Notification automation has its own failure mode. In our case the notification bot’s usage permission was limited to its creator, so notifications we believed we had sent to other owners reached nobody. The send API returns success. For anything in the notification path, verify once on real devices that the message appeared on the recipient’s screen — not merely that it was sent.

The minimum you need to run this in production

Here is the same material as a minimum requirement list for organisations starting out. Whether these seven items exist mattered far more to continuity than any advanced feature.

ElementMinimum actionConsequence of skipping it
Input healthRecord connectivity, authentication and period-query outcomes separately from countsFailures keep passing as zero records
IdempotencyFreeze the event’s unique key as a specificationDuplicate filing, or everything treated as new
Run recordDefinition path read, period covered, counts of new / updated / unchangedNobody can explain what changed
MonitoringDetect consecutive zeros, consecutive failures, and runs that never happenedIt stops in silence
Re-runnabilityThe same command resumes from where it stoppedEvery failure needs manual recovery
Approval boundaryKeep externally visible operations with a personYou learn of a wrong send after it happened
Output caveatsAlways state denominator, period and exclusionsAggregates lead to wrong conclusions

None of these seven require a large platform or a dedicated observability stack. In our own setup the destinations are existing business tools, execution is a daily scheduler, and monitoring is a message into the same chat. What was needed first was not more tooling but a decision about how failure should look.

How to start: build small, and test the way it breaks first

If you are considering in-house automation, we recommend this order. For evaluation and verification planning, see AI PoC cost items and success criteria; for staff capability, see practical testing for generative AI employee training.

  1. Pick exactly one pipeline. It should occur daily, cause no catastrophe if delayed, and allow the correct answer to be verified after the fact. Ours was email accumulation.
  2. Record into an existing tool. Build no new screen. Put the output where people already look.
  3. Stop at L4 (draft). Do not let it send at the start. For one or two weeks, have a person grade the drafts.
  4. Test how it breaks, first. Cut the connection, remove the permission, empty the input. Look at what the notification says in each case; if you cannot tell them apart, fix the design.
  5. Add the consecutive-zero alert before you move to unattended operation.
  6. Align the second pipeline’s key format and run-record format with the first. If each pipeline invents its own conventions, neither monitoring nor handover is possible.
  7. Audit output denominators quarterly. Operational changes quietly change what the data means.

FAQ

How large a team does AI agent workflow automation require?

Ours was assembled by people who understand the business, on top of existing business tools and a scheduler. There is no dedicated AI platform team. That works because this is internal work with no external blast radius. Automation touching customer systems or equipment needs its own requirements, testing and operations structure.

How accurate does generative AI need to be before it is usable in production?

No single accuracy figure decides it. What matters is whether a person can notice an error, and whether the operation can be undone. Reversible operations — recording, drafting — tolerate some error. Irreversible operations — sending, registering, deciding — realistically keep a human approval step regardless of measured accuracy.

What should we automate first?

Something that occurs daily, whose correctness can be verified after the fact, and where delay is not catastrophic. The opposite profile — monthly, unverifiable, immediately costly when late — is a poor first target.

How do we notice that an automated process has stopped?

“Notify on failure” is not enough. Monitor for consecutive zero-record runs and for runs that did not happen at all (dead-man monitoring). Every long-lived outage we experienced presented as “succeeded, zero records.”

Is it acceptable to send internal data to generative AI?

That depends on your information classification and your contracts. At minimum, confirm from the terms of service and the product settings which data is transmitted, where it is stored, and whether it is used for training. If personal data of people in Thailand is involved, confirm PDPA scope with a qualified adviser and design to pass only the minimum data required.

Should we let an agent send email?

We do not. We automate up to the draft and a person sends. A wrong send cannot be recalled and affects the relationship directly. Even when opening this up later, restricted recipients, a pre-send confirmation and an auditable send log should be in place first.

Summary: automation is judged by how it fails quietly

What we learned running AI agent workflow automation is that the hard part is not the capability of the generative model but everything around it. A lost connection reported as success. An identifier change turning the same item into a different one. An aggregate with no denominator producing a confident conclusion. Caches and sync swallowing a change. A duplicated definition keeping the stale copy alive. None of these are dramatic outages; all of them quietly erode business value.

So the order of design is to decide how it fails, and to build what failure looks like, before adding features. Record input health separately from counts, freeze the idempotency key as a specification, attach a denominator to every output, and keep externally visible operations with a person. With those four in place from the start, our operation did not break down as pipelines were added.

If you are designing back-office automation for a Thai site, or in-house automation that integrates with a core system, contact us through the TOMAS TECH enquiry page. Once the target work and the existing systems are known, we can make concrete where to start and where approval should stay.

References