Blog

2026.09.06

Enterprise Chatbot Case Studies: A 90-Day PoC for Japanese Companies in Thailand

Enterprise Chatbot Case Studies: A 90-Day PoC for Japanese Companies in Thailand

Enterprise chatbot case studies can be impressive yet difficult to translate into a practical first step. A Japanese company in Thailand may have HR questions in Thai, policies approved in Japanese, service tickets in English, and data split across Microsoft 365, an ERP and local files. The useful lesson is therefore not which model a famous company selected. It is how the company defined the job, governed knowledge, limited permissions, measured true resolution and handed difficult cases to people.

This article examines five published cases—Ada, Klarna, Cemex, Orion Health and DoorDash—through that operating lens. It then turns the lessons into a 90-day proof-of-concept plan for HR inquiry AI and internal help desk AI in Thailand. Every company metric below is a result publicly reported by the named company or its technology provider. None is presented as a Thai deployment, a TOMAS TECH customer result, or a guaranteed outcome for your organization.

Before comparing enterprise chatbot case studies, define “resolved”

A bot response, a conversation that never reached a human, and a genuinely solved request are three different events. A dashboard that mixes them can make a poor user experience look efficient.

The published Ada story makes this distinction explicit. It calls the share of conversations kept away from a human the containment rate, while the resolution rate asks whether the user received a relevant, accurate and safe outcome without human intervention. Ada reports that an earlier product had 70% containment but only 30% resolution. Customers moved to the newer system typically saw resolution rates up to 60%, while the highest-performing customers were above 80%.

Those are Ada/OpenAI figures, not targets for Thailand. The transferable practice is to count failure honestly. If an internal help desk AI returns VPN instructions but the employee still cannot connect, that request is not resolved. It becomes resolved only when connectivity is restored or when a ticket with the right device, operating-system and error details reaches the right support queue.

KPIWhat it measuresCommon mistake
Automated response rateA response was producedSays nothing about correctness
Containment rateNo human transfer occurredCan reward a dead end
Resolution rateThe intended outcome was safely completedNeeds a clear outcome rule
Repeat-contact rateThe user returned with the same issueMust link channels where possible
Time to resolutionTime from request to completed outcomeDifferent from first-response speed

Five enterprise chatbot case studies and what a Thai operation can reproduce

Case 1: Ada—optimize resolution quality, not confinement

Ada rebuilt its evaluation around whether conversations were truly resolved. Its published framework assesses relevance, accuracy and safety and, according to the case, reached 80–90% agreement with a human reviewer in testing. The headline movement from about 30% resolution to as much as 60%, and above 80% for leading customers, sits on top of that evaluation work.

A Thai subsidiary should not copy those percentages into a business case. It can copy the measurement architecture. For each selected inquiry, define the observable completion event, the acceptable sources, the conditions for refusal, and the evidence a reviewer needs. Review a sample of both “successful” and escalated conversations. A low handoff rate is valuable only when resolution and trust remain high.

Case 2: Klarna—measure repeat contacts and elapsed time as well as volume

OpenAI reports that during its first live month Klarna’s assistant handled 2.3 million conversations, equal to two-thirds of Klarna customer-service chats. It performed work described as equivalent to 700 full-time agents, reduced repeat inquiries by 25%, and helped customers resolve matters in less than two minutes compared with eleven minutes previously. The story also reports 24/7 operation across 23 markets and more than 35 languages.

These are Klarna’s published global results. “Equivalent to 700 agents” does not mean another company can remove 700 positions, and it should not be treated as a Thai staffing promise. What can be reproduced is the balanced scorecard: scale, repeat contact, resolution time and user satisfaction together.

For HR inquiry AI, a quick answer to “How many leave days do I have?” is not sufficient if the employee then emails HR to verify it. Link chat, email, phone and ticket events to the same case where privacy and systems permit. If cross-channel identity matching is not appropriate, use sampled follow-up surveys and ticket-reopen rates. The goal is to detect displaced work rather than merely moving it out of the chatbot dashboard.

Case 3: Cemex—place HR inquiry AI in the employee’s existing channel

Microsoft’s Cemex customer story describes Consult HR, built with Copilot Studio and integrated with Microsoft Teams, ServiceNow and SAP SuccessFactors. The proof of concept was completed in November, followed by a controlled 25-user pilot. The system moved to full production on 26 January 2026 and served approximately 800 employees in Central Mexico.

This is a Central Mexico deployment, not a project in Thailand. Its reusable pattern is the channel and workflow design. Employees ask questions in Teams, retrieve information from approved SharePoint sources, and can create and track a ServiceNow request. The chatbot is not a separate destination that employees must remember; it is an entry point to the existing service process.

For a Japanese company in Thailand, governance around local applicability matters. A Japanese headquarters policy translated into Thai is not automatically the rule for a Thai employee. Record an owner, entity/site, eligible population, effective date, expiry date, original language and approved translation for each source. Benefits, tax, disciplinary and employment questions may need local review even when the wording looks similar.

Case 4: Orion Health—internal help desk AI starts with knowledge governance

AWS reports that Orion Health built Oribot to search across fragmented internal knowledge. The published case says employees can retrieve answers from more than 500,000 records in under one minute, a working prototype launched within two months, and the support team is expected to reclaim around 50 staff hours per day from reduced search effort.

Orion Health is a New Zealand-headquartered healthcare software company. The 50-hours-per-day number is an expectation in the AWS case, not a guaranteed saving and not a Thailand result. More important is the architecture behind it: retrieval-augmented generation across approved sources, session orchestration, a vector store and hosting inside an environment where access and model configuration can be controlled.

FAQ automation will not repair weak source material. Conflicting versions, unknown owners, excessive permissions and outdated procedures will surface as inconsistent answers. Before improving a prompt, determine whether the failure came from intent classification, document retrieval, access rights, source quality, tool execution or presentation. A model upgrade cannot make an obsolete policy current.

Case 5: DoorDash—voice support needs a latency budget and careful attribution

AWS states that DoorDash and the AWS Generative AI Innovation Center produced a reference architecture suitable for production A/B testing over eight weeks, or roughly two months. Testing capacity increased 50 times, and the published response latency with Claude 3 Haiku was 2.5 seconds or less.

The same AWS page also reports that the pre-existing Connect Customer and Amazon Lex IVR had reduced agent transfers by 49%, improved first-contact resolution by 12%, and produced USD 3 million in year-over-year operational savings. Those three figures belong to the existing IVR layer; they should not be attributed solely to the later generative-AI work. Maintaining that causal boundary is essential when building an investment case.

Voice requires separate tests for Thai, Japanese and English. Measure speech-recognition errors, response-start latency, interruptions, confirmation loops and human transfer for each language. One average multilingual score can conceal a serious failure in the language used by frontline staff.

Enterprise Chatbot Case Studies: A 90-Day PoC for Japanese Companies in Thailand - figure 1

What the five chatbot implementations have in common

CaseBounded jobPublished measureReusable design principle
AdaResolve customer issuesResolution about 30% to up to 60%; leading customers above 80%Measure completed outcomes, not just containment
KlarnaCustomer-service tasks2.3M first-month conversations, two-thirds of chats, 25% fewer repeats, under 2 minutesPair scale with quality and elapsed time
CemexHR information and service entry25-user pilot, production for about 800 staff in Central MexicoUse the employee’s existing channel and ticket system
Orion HealthInternal knowledge retrieval500,000+ records, under one minute, two-month prototype, expected 50 hours/dayGovern knowledge and access before scaling
DoorDashVoice self-serviceEight weeks, 50x testing capacity, 2.5 seconds or lessTest under production-like conditions and keep causal attribution clean

None began as “an AI that answers everything for everyone.” Each case had a bounded job, source or system access, an operating channel and a metric. “Automate HR FAQs” is too broad. “Help permanent employees at the Bangkok office find the current approved leave procedure and escalate an exception to HR with the required context” is testable.

Select the first HR or internal help desk job in Thailand

Good starting points for HR inquiry AI

Suitable first jobs are frequent, backed by an approved source, low in consequence when handled conservatively, and easy to escalate. Examples include where to start a leave request, where to view a payslip, how to file a standard medical-benefit claim, the training calendar, or which onboarding forms are normally required.

Discipline, harassment, medical detail, individualized tax judgments and termination questions should not be autonomously resolved in an initial PoC. The agent can acknowledge, protect the user’s privacy and route the request to the correct person. A correct refusal and rapid human handoff are successful outcomes.

Good starting points for internal help desk AI

Potential jobs include official password-reset guidance, initial Wi-Fi or VPN troubleshooting, an approved-software request, device-replacement intake and ticket-status lookup. The AI can add value by collecting the site, device, operating system, error and urgency before drafting or creating a ticket.

Keep privileged access, security-alert overrides and factory-equipment configuration behind human approval. Start with retrieval, then drafts, then approved low-risk actions. For a deeper treatment of knowledge controls, see Internal Policy Search AI for Japanese Companies in Thailand.

A 90-day chatbot PoC plan

Days 1–15: choose one job and baseline today’s work

Classify eight to twelve weeks of tickets, email, chat and phone notes. Record volume, handling time, waiting time, repeat contact, routing and reopen rates. Score candidate jobs on frequency, availability of an approved answer, identifiable owner, limited downside and observable completion.

Agree the scope with Japanese and Thai management, the operational owner, IT/security and representative users. Define acceptance criteria before building. Avoid a vague target such as “automate 60%.” For example, the company running the PoC could require at least 90% accuracy on reviewed answers, zero critical incorrect answers, a 20% reduction in median time to resolution, no deterioration in repeat-contact rate versus its own baseline, and at least 95% correct routing of high-risk questions. These are illustrative thresholds for the participating company to adapt to its own risk and baseline; they are not universal pass marks or results reported by the five case-study companies.

Days 16–30: curate knowledge and write escalation rules

Attach an owner, approver, site/entity, audience, language, approved translation, effective date, review date and confidentiality level to every source. Exclude obsolete and ownerless documents. Define when the agent answers, asks a clarifying question, drafts a ticket, requests approval, refuses or transfers immediately.

Build separate evaluation sets for Japanese, Thai and English. Include spelling mistakes, abbreviations, mixed languages, romanized Thai, vague Japanese, references to old policy, attempts to access another employee’s data, prompt injection and ambiguous requests. Report quality by language and failure type, not only as an overall average.

Days 31–60: build read-only, integrate the channel and run a controlled pilot

Keep business systems read-only during the first pilot. Show the source document, effective date and relevant passage with each answer where possible. When evidence is weak, transfer the case instead of inventing an answer. Integrate the channel employees already use and ensure ticketing, authentication, logging and notification work together.

A cohort of roughly 20–30 users is often manageable, but the right number depends on inquiry frequency and risk. Include different departments, seniority levels, languages, sites and digital skill levels. Explain the collection purpose, log scope, retention and support route before the pilot. If LINE is in scope, LINE Chatbots for Business in Thailand in 2026 covers channel-specific design considerations.

Enterprise Chatbot Case Studies: A 90-Day PoC for Japanese Companies in Thailand - figure 2

Days 61–75: run production-like A/B tests and classify failures

Compare the current flow with the AI-assisted flow using the same definitions. Examine resolution, repeat contact, human handling time, user waiting time, reopen rate, critical error severity and cost. Include knowledge preparation, integration, review, security and maintenance costs, not only model API charges.

Classify failures: wrong intent, wrong retrieval, stale source, access-control failure, unsupported calculation, incorrect tool action, delayed escalation, or confusing interface. This prevents a model change from becoming the default response to every problem.

Days 76–90: decide expand, rework or stop

Review business value, answer quality, safety and operability. Higher automated volume with more repeat contact is a rework signal. Good answers with no sustainable knowledge owner are an operational failure. A small PoC that supports a clear “not yet” decision has prevented a larger mistake.

If the result supports expansion, release capability in stages: answers and links, form or ticket drafts, submission with human approval, and only then low-risk reversible actions. Verify approvals, audit trails, rollback and change ownership at every stage.

Governance for enterprise chatbots: scope, permission, evaluation and people

OpenAI introduced Presence on 22 July 2026 as an approach in which each deployment starts with a specific job and receives only the knowledge and system access needed for that job. The company defines policies, approved actions and when a person takes over. Pre-launch simulations cover routine, edge and high-risk scenarios, while production sessions and escalations become inputs to controlled improvement.

OpenAI reports that its own English-language phone-support channel resolves 75% of inbound issues without human assistance, and that its improvement loop reduced human handoffs by 15 percentage points in ten days. These are specifically OpenAI’s English phone-support results—not an average across customers, a Thai-language result or a guaranteed outcome. The transferable lesson is the continuous evaluation and controlled-change loop.

Enterprise Chatbot Case Studies: A 90-Day PoC for Japanese Companies in Thailand - figure 3

Review personal-data handling flow by flow in Thailand

Employee and customer inquiries may include names, contact details, employee IDs, purchase history, support history and sometimes sensitive HR or health information. Map where input is sent, why it is used, who can view it, where it is stored, when it is deleted and whether it is reused for evaluation.

The Thai PDPC/GPPC privacy notice is a useful primary reference for thinking about visible purposes for collection, use and disclosure. It does not mean that publishing one notice makes every chatbot workflow compliant. Review purpose, lawful basis, retention, processors, cross-border transfers and employee/customer communications for the specific data flow with Thai legal and privacy specialists. This article is not legal advice.

Ten questions for chatbot vendor selection

  1. Can the vendor describe the job and its observable completion event?
  2. Can answers show the approved source and effective date?
  3. Are Japanese, Thai and English evaluated separately?
  4. Does retrieval respect the user’s existing access rights?
  5. Can high-risk, low-confidence and system-error cases reach a person reliably?
  6. Can it integrate with the existing chat, ticket, identity and notification stack?
  7. Are source, response, tool action, approval and configuration changes auditable?
  8. Can the vendor explain storage, retention, cross-border flow and model-training use?
  9. Are production failures converted into regression tests before a controlled change?
  10. Does the annual cost include knowledge ownership, quality review, integration and support?

Do not reduce the decision to “which model is strongest.” The more useful question is whether the system can explain, contain, reverse and improve a failure. Training staff to use generative AI and implementing a production chatbot are different workstreams; how to select a generative-AI training provider in Thailand addresses the former.

Build the financial case around time reallocated, not positions removed

It is risky to copy an FTE-equivalent figure from a case study into a headcount-reduction plan. Inquiry work includes more than the visible conversation: searching for information, verifying identity, making a judgment, recording the result, obtaining approval and coordinating with another team. A chatbot may shorten some of those steps without replacing accountable judgment. A more defensible model starts with annual inquiry volume multiplied by the time spent on the in-scope steps and the improvement observed in the PoC. It then states where the recovered capacity will be reallocated.

HR might spend more time on complex employee cases, manager support and workforce-experience improvements. IT might redirect capacity toward incident prevention, security hardening and business-process improvement. Framing the benefit as reduced waiting, rework and search time makes it easier for local operations and management to share the same definition of value.

Include integration, source curation, approved translation, access design, evaluation-set creation, security review, production-log review, document updates and user support in the cost model. Usage-based model, speech and storage charges should include both a normal month and a peak month. Overlong answers, overbroad retrieval and unnecessary history retention can increase cost as well as risk.

Use optimistic, central and conservative scenarios. Vary repeat-contact rate, human-review rate, data-preparation effort and monitoring effort separately instead of hiding every uncertainty inside one automation-rate assumption. The purpose of the PoC is not to prove the optimistic scenario. It is to replace unknown assumptions with measurements from the company’s own work.

Agree responsibilities between Japanese headquarters and the Thai entity before the PoC

A multi-site chatbot needs a decision map before it needs more technical capability. Japanese headquarters may own the common platform, security standard and vendor agreement, while the Thai entity owns local policy, approved Thai wording, employee communication and escalation. Write these responsibilities down. Any area in which each side assumes the other is checking becomes a likely source of stale knowledge after launch.

At minimum, name a business-process owner, data owner, system owner, quality evaluator, security decision maker, privacy reviewer and continuity owner. Decide who can suspend the agent when answers are repeatedly wrong, who communicates the alternative service channel, and how the team works when the model or ticket integration is unavailable. This prevents an AI component incident from becoming an outage of the whole help desk.

Set a knowledge-change service level as well. When an organization chart, Thai holiday, benefit, application screen or security procedure changes, define the permitted time from source approval to chatbot availability. If an urgent change cannot be reflected safely, temporarily disable automated answers for that topic and route it to a person.

User communication is also an operating control. Explain the agent’s scope, the possibility of error, when to check the original source and how to reach a person. Keep the “report a problem” action close to the answer. Low-friction reporting turns small signs of confusion into evaluation data before they become serious incidents.

Frequently asked questions

Can enterprise chatbot case-study numbers be used directly in our ROI model?

Use them as external reference points, not direct inputs. Your inquiry mix, volumes, handling time, language, system landscape, knowledge quality and labor costs differ. Establish your own baseline and compare the PoC using identical definitions.

Where should help desk automation begin?

Start with one frequent, low-risk job whose approved answer exists and whose completion can be observed. Begin read-only, cite the source, and transfer unresolved cases with context rather than forcing containment.

Can HR inquiry AI support Thai and Japanese at the same time?

Yes technically, but translation alone is insufficient. Keep approved knowledge and applicability clear by entity, site and employee group. Test correctness, refusal and escalation separately in each language.

What is the value of generative AI in FAQ automation?

It can handle paraphrases, multi-condition questions and follow-ups better than keyword matching. It must still be grounded in approved sources with citations, boundaries, refusal and human escalation.

How much system access should internal help desk AI receive?

Initially, approved-knowledge retrieval and ticket drafting are safer. Add ticket creation or status lookup with approval after quality is demonstrated. Do not grant privileged changes, equipment control or mass updates without explicit approval, rollback and audit.

Is 90 days enough to decide on production?

It is enough to make an expand, rework or stop decision for a tightly bounded job if the baseline, acceptance rules, knowledge owners and evaluators are clear. It is not proof of the impact of an enterprise-wide rollout. Sustainable monitoring and knowledge updates remain part of production readiness.

Conclusion: reproduce the design, not the headline

The most valuable parts of the five cases are not their largest numbers. Ada shows how to define resolution. Klarna links scale to repeat contact and elapsed time. Cemex integrates HR self-service with the employee’s existing channel and a controlled pilot. Orion Health emphasizes governed retrieval. DoorDash demonstrates production-like testing, a latency budget and disciplined attribution.

For a Japanese company in Thailand, a sound 90-day PoC narrows the job, records today’s baseline, curates approved multilingual knowledge, begins read-only, evaluates each language separately and treats the human handoff as a designed success path. Once true resolution is visible in a small scope, the organization can decide where to expand and where people should remain in control.

If you are considering HR inquiry AI or an internal help desk AI and want to define the first job, baseline KPIs or Japanese/Thai evaluation plan before selecting a product, contact TOMAS TECH. We can help structure what the 90-day PoC must prove, even at the early exploration stage.

Sources

*The company figures in this article are drawn from primary materials published by the named companies or technology providers. They do not describe implementations in Thailand and do not guarantee comparable outcomes. Published information and product availability may change.*