Blog

2026.10.06

Maintenance AI Assistant: Cut Diagnosis Time with RFP and FAT/SAT

Maintenance AI Assistant: Cut Diagnosis Time with RFP and FAT/SAT

“Every time an alarm goes off, the phone of a veteran Japanese engineer or a Thai team leader rings. Even in the middle of the night, even on holidays.” We often hear this from maintenance managers at Japanese-owned factories in Thailand. They want young Thai maintenance technicians to be able to look at the operating manual, the alarm code list and past repair records, and reach candidate causes and procedures on their own. More and more factories are considering a maintenance AI assistant implementation as a way to get there.

Here is the conclusion first. What a maintenance AI assistant shortens is not “repair time” as a whole, but the part of it spent finding out what is wrong (fault localization and diagnosis). And what determines the quality of its answers is not the AI model but the quality of your past repair records. The time spent stopping the equipment, cutting the power and hanging a “Do not switch on” tag is time that must not be shortened. The decisions you need to make for implementation come down to three: (1) define what you are going to shorten, (2) prepare what the assistant will read, and (3) decide through acceptance testing how far you can trust it.

Please note that all machine counts, case counts, times, amounts and payback years for the factory in this article are original estimates and assumptions (placeholder values) created for this article, based on the model factory described later. They are neither industry averages nor survey results. No primary information on vendor prices or effects has been used either. Please read them as a “calculation template” and replace them with your own measured figures and quotations.

Why a Maintenance AI Assistant Implementation Is Now an Issue for Thai Factories

Generative AI announcements for maintenance keep coming

Over roughly the past two years, there has been a steady stream of announcements of generative AI assistants for maintenance and frontline workers. As examples of the options on the market, here are the announcements we could confirm. All effect figures were published by the vendors or announcing companies themselves and have not been verified by third parties.

Announced byDateWhat was announcedEffect figures and their nature
SiemensMarch 2025Added a generative AI offering for maintenance to Industrial Copilot. A package combining repair guidance and predictive maintenanceIn the first pilot case, “time spent on reactive maintenance reduced by 25% on average” (vendor-reported figure; target plant and number of cases not disclosed; not an MTTR value)
Accenture, Avanade, MicrosoftApril 2026Jointly developing an “agentic factory”. Supports initial condition checks, diagnosis and guided troubleshooting, and helps create maintenance tickets and spare parts ordersNo measured values published. General availability said to be in the second half of 2026
Microsoft, SchaefflerDecember 2024Pilot deployment of an agent that answers natural-language questions such as “What caused the stoppage on Line 3 yesterday?” across ERP, MES and equipment dataNo quantitative effect stated
IBM (Maximo)June 2026Maximo Application Suite 9.2. Technicians can look up asset information and history in natural language on mobile. Document extraction from OEM maintenance manuals is a planned extensionNo figures stated
Rockwell Automation × AuguryJuly 2026Augury’s agents analyse equipment condition and recommend corrective actions, and Fiix MAX turns them into maintenance work ordersInitial products were announced as “planned to be available in September” (actual availability has not been confirmed)

There is movement in Japan as well. The Hitachi Group launched the equipment and asset management system “SmartFAM Ver.4” in April 2025 and says it uses generative AI to analyse past maintenance data and show similar failures and recovery measures. JR East announced a plan to introduce generative AI into a system that supports failure recovery for signalling and telecommunications equipment, stating that it aims to “reduce recovery time by up to 50%” compared with before. This is a target, not an achieved result. Hanshin Electric Railway and ITEC Hankyu Hanshin have been trialling, since October 2025, a system that analyses past failure response records to estimate failure causes and display recovery manuals (both are railway cases). PKSHA Technology announced “PKSHA Maintenance”, a service for building RAG-based generative AI for equipment maintenance, in September 2024, and reported that when experts evaluated it with several years of real data it “generated correct or partially correct remedy proposals for more than 80% of unknown events”. This 80% includes “partially correct” proposals and is not an accuracy rate.

What you need to watch out for here is that the figures do not mean the same thing. Siemens’ 25% is “a reduction in reactive maintenance time in a pilot case”, PKSHA’s 80% is “the generation rate of remedy proposals including partially correct ones”, and JR East’s 50% is “a target”. None of them means “MTTR was reduced by X%”. It can be argued that the buyer cannot compare these on the same yardstick. That is exactly why you need to decide first what you will measure in your own factory, and how.

Simply introducing AI does not reduce stoppages

In a survey of the United States and Canada published in May 2026 by the CMMS vendor MaintainX (2,234 maintenance and operations leaders), 58% said they were already using AI in their work. On the other hand, 79% said unplanned downtime had “stayed the same or increased”. The main causes of unplanned downtime cited were labour shortages and insufficient knowledge transfer. These are not Thai or ASEAN data, but they can be read as saying that “starting to use AI and reducing stoppage time are two different things”.

Where Thai manufacturing stands

According to a Reuters report (30 September 2026) on an announcement by Thailand’s Ministry of Industry, the Manufacturing Production Index (MPI) for August 2026 rose 4.44% year on year, the second consecutive month of growth. However, the Ministry had just revised the MPI calculation method, so it cannot simply be compared with past years. The Federation of Thai Industries (FTI) Thai Industries Sentiment Index stood at 89.70 in August 2026, down from 90.0 in July. Heavy rain, high energy prices and a decline in car sales were cited as reasons for the deterioration.

On the investment side, according to an announcement by the Thailand Board of Investment (BOI), there were 132 applications worth USD 507.6 million in the first half of 2026 under the “Smart and Sustainable Industry” measure, which covers machinery upgrades, digital technology and the introduction of automation. These are application figures, not approvals.

Production is moving, but cost pressure is strong. The ability to bring stopped equipment back quickly translates directly into the ability to limit losses. At the same time, a night call-out system that depends on veterans is hard to sustain for long. The model factory in this article assumes this kind of situation (this is an assumption of this article, not a statistic).

What Is a Maintenance AI Assistant? A System That Connects Manuals, Alarm Codes and Past Records

The overall picture

The maintenance AI assistant in this article is a system that, when a maintenance technician asks a question in natural language, searches the factory’s documents and records and returns candidate causes and check procedures, with sources. Most use an architecture called RAG (retrieval-augmented generation), made up of the following parts.

  • Entry points: alarm occurrences, questions from maintenance technicians (Thai, Japanese, English), photos and alarm codes. Integration with the mechanism that sends alarm notifications is covered in “Equipment Alarm Notification System“
  • Documents it reads: manufacturers’ operating manuals, alarm code lists, in-house work procedures, past work records
  • Systems it refers to: the equipment register and work history in the CMMS (computerized maintenance management system), and stoppage records in the MES where needed. The granularity of the CMMS register and the routes for raising work orders are explained in “How to Choose an Equipment Maintenance Management System“
  • Answers: candidate causes, the order of checks, the relevant sections of work procedures, similar past records, and sources (document name, page, record number)

What it can and cannot do

What it can do (depending on design)What it cannot do or must not do
Pull up the relevant page of the operating manual and the inspection points from an alarm codeJudge the actual condition of the equipment (a person checks the real thing)
Find past records with similar symptoms and show the cause and remedy at that timeGuarantee the “right answer” for a failure that is not in the records
Show young technicians the order of checks, following the work procedureSuggest shortcuts that skip stopping, isolation or lockout
Produce draft work records and candidate failure codesMake the final decision (judgment and sign-off are done by people)
Answer questions in Thai, Japanese and English from the same documentsPlausibly fill in things that are not written in the documents

We wrote “depending on design” in the left column. With a design that does not return sources, or when records are not in order, much of the left column does not work. The thinking behind a design that returns sources is covered in detail in “Equipment Manual Search AI“.

IBM lists, as a feature of its Maximo assistant, suggesting problem codes from work order descriptions. In the MaintainX 2026 survey too, “knowledge capture” is among the uses of AI. It can be said that AI is starting to be used not only to read records but also to raise their quality. We will make use of this point in the section on records later.

Breaking Down MTTR: AI Shortens “Diagnosis”, and Must Not Shorten “Safety”

Maintenance AI Assistant: Cut Diagnosis Time with RFP and FAT/SAT - figure 1

MTTR is not “mean time to repair”

On the shop floor, MTTR is often called “mean time to repair”, but in the International Electrotechnical Commission (IEC) vocabulary IEC 60050-192, MTTR is mean time to restoration, defined as the expectation of the time to restoration. The terms “mean time to repair” and “mean time to recovery” are deprecated.

Time to restoration is the time “from the moment a failure occurs until restoration”. If the time of failure is not known, it is counted from the moment the failure is detected. In other words, MTTR as defined by the IEC includes not only the time for the repair itself but also waiting times and delays.

Within that, “repair time” is defined as consisting of the following four parts.

  1. Fault localization time (narrowing down where the problem is)
  2. Fault diagnosis time (finding out why it is wrong, where not included in localization)
  3. Fault correction time (part replacement or adjustment)
  4. Function checkout time (confirming that it has been fixed)

Repair time does not include technical delays, administrative delays or logistic delays. As an example of technical delay, the IEC gives “putting the item in a safe state (shutting down, cooling, isolating, earthing and so on)”. Mean repair time is a separate term from MTTR (MRT).

In the best practice metrics of SMRP, the US-based professional society for maintenance and reliability, MTTR is listed as “Mean Time to Repair or Replace”, which differs in meaning from the IEC’s “restoration”. This means that even with the same “MTTR”, what is counted from where to where can differ between organizations. If you use MTTR as an internal metric, you first need to align the definition.

Time AI helps with, time it does not, and time that must not be shortened

Applying this breakdown makes clear where a maintenance AI assistant helps.

Time categoryContentHow the AI assistant helps
Putting the equipment in a safe state (technical delay)Stopping, isolation, lockout, hanging a “Do not switch on” tagMust not be shortened. The AI’s role is to present the procedure without skipping it
Localization and diagnosisFinding out where and why something is wrongMainly helps here. Narrows down candidates by searching manuals, alarm lists and past records
CorrectionPart replacement, adjustmentReferring to procedures helps a little, but the work itself is hard to shorten
Function checkoutTrial run, quality checkLimited to presenting check items
Waiting for parts, waiting for approval, etc.Logistic delays, administrative delaysBarely helps. These are issues of inventory and approval systems

It is reasonable to think that what an AI assistant mainly shortens is localization and diagnosis time, while waiting for parts or approval does not shrink. So, in acceptance testing and later effect measurement too, it can be argued that making diagnosis time an independent metric, rather than MTTR as a whole, is easier to measure.

The 90-minute breakdown at model factory M

At this article’s model factory M, the average time to restoration per failure that stops the line is set at a placeholder of 90 minutes, broken down as follows (for simplicity, delays such as waiting for parts are not included in these 90 minutes).

CategoryTime (placeholder)Treatment in the estimate
Putting in a safe state (stopping, isolation, lockout, tagging)10 minExcluded from reduction
Localization and diagnosis40 minTarget for reduction by AI
Correction (part replacement, adjustment)35 minUnchanged
Function checkout5 minUnchanged
Total10 + 40 + 35 + 5 = 90 min

Of the 90 minutes, the room for AI to help is 40 minutes. Looking at this breakdown shows that a target such as “halving MTTR” is not realistic. Conversely, how far the 40 minutes of diagnosis time can be shortened depends heavily on the quality of the records, which we explain next.

Record Quality Determines Answer Quality: Organize Failure Codes Before Using Generative AI for Maintenance Records

Work records cannot be read as they are

Among the answers a maintenance AI assistant gives, the most valuable is “the last time this symptom occurred, what was the cause and what fixed it”. To produce that, past work records need to be in a searchable form.

In reality, however, work records on the shop floor are often messy. A 2018 paper by researchers from the US National Institute of Standards and Technology (NIST) and the University of Western Australia points out that maintenance work order records often cannot be analysed by computer as they are, because of misspellings, jargon specific to fields and workplaces, and abbreviations. The paper did not evaluate generative AI, but it is reasonable to think that the premise that free-text fields in records are messy has not changed.

In Thai factories, a language problem is added on top. The same failure is written in Thai one day, in Japanese another day, and as an English abbreviation on yet another day. Even if “cylinder”, “cyl” and “กระบอกสูบ” refer to the same thing, the search does not know that.

“One in five lacks information”: a study of wind power maintenance records

There are also studies that used generative AI to organize maintenance records. A study of wind power maintenance records published in 2026 (a preprint, not yet peer-reviewed) used a large language model to structure 16,316 maintenance records covering 9.2 years from 280 turbines at 32 onshore wind farms. The record problems it cites are inconsistent system codes, overly coarse category fields, and technicians’ free-text descriptions.

The result was that failure modes could be grounded for 11,662 records, 3,441 were judged to have “insufficient information”, and 1,213 were excluded by rules (11,662 + 3,441 + 1,213 = 16,316). In other words, about one in five records (3,441 ÷ 16,316 ≈ 21%) did not have enough information to determine the failure mode.

This is a study of wind power maintenance records, not factory records. Even so, the point that “improving AI performance cannot extract information that is not written in the records” is likely to apply to factories as it is.

A failure code system: borrowing the thinking of ISO 14224

ISO 14224 is a useful reference as a framework for organizing records. It is a standard for the petroleum, petrochemical and natural gas industries, dealing with the collection and exchange of reliability and maintenance data for equipment. It is not a standard that applies directly to general assembly plants, but its approach of dividing the data to be collected into the following three groups is useful.

  • Equipment data: the equipment taxonomy and attributes
  • Failure data: failure causes and failure consequences
  • Maintenance data: maintenance actions, resources used, maintenance consequences, downtime

The failure modes defined in the standard are described as usable as a “reliability vocabulary”. In factory terms, this means recording “equipment → component → failure mode → cause → action” as selectable codes, and supplementing the detailed situation with free text.

Recoding past records and an input template

The clean-up splits into two parts.

  1. Recoding past records: retrospectively attach codes for equipment, component, failure mode, cause and action to several years of past work records. Having generative AI propose candidate codes and veterans check them speeds this up. However, the safe approach is not to force codes onto records that lack information, but to keep them as “insufficient information”
  2. An input template for future records: put a template that combines selectable failure codes with free text into the CMMS. If the AI proposes candidate failure codes from the description and the technician chooses, record quality can be raised without adding input effort

How to capture in records the judgments veterans carry in their heads, such as “if you hear this sound, check here first”, is covered in “Skilled Worker Knowledge Transfer with AI“. Also, because operating manuals and procedures get revised, version control that prevents the AI from answering with procedures from an old version is essential. For details, see “Factory RAG Document Version Control“.

Generative AI Errors and Safety: Sources, Human Judgment and Lockout

Plausible mistakes and over-trusting AI

“NIST AI 600-1”, the guidance on risk management for generative AI published by the US National Institute of Standards and Technology (NIST) in July 2024, lists “Confabulation” as a risk specific to generative AI. This is the confident generation of erroneous or false content, what is commonly called hallucination. The guidance warns that this is particularly problematic in highly specialized fields and in long answers, and further that the logic or citations used to justify an answer can themselves be fabricated.

The same guidance lists “automation bias”, the tendency to over-rely on automated systems, as a risk in the relationship between people and AI, and states that it can amplify the risk of confabulation. It is reasonable to think that the less experienced the maintenance technician, the more likely they are to believe the AI’s answer as it is. On the maintenance floor, a plausible mistake can lead directly to equipment damage or injury.

Turning the recommended measures into design and acceptance

NIST AI 600-1 lists the following among its suggested actions (it is voluntary US guidance with no legal force).

  • Before deployment and during operation, verify the sources and citations in generative AI output
  • Verify that the data used for RAG is grounded in evidence
  • Avoid generalizing performance from narrow, non-systematic evaluations based on a small number of cases
  • Share pre-deployment test results with the people who have authority to approve release
  • Establish and maintain post-deployment monitoring (especially for confabulation)
  • Enter contracts and SLAs that set quality standards, and include contract clauses that allow evaluation of third-party generative AI processes

Applied to maintenance AI, this becomes: “always attach sources to answers that contain procedures or numbers”, “do not judge on a handful of demo cases, but test with an evaluation set built from past failure cases”, “have approvers such as the plant manager look at the test results and decide”, and “keep running the same evaluation after go-live”.

The guidance also lists, as an information security risk, “indirect prompt injection”, in which malicious instructions are planted in data likely to be retrieved by search. When ingesting external PDFs or vendor materials, measures such as an approval process for documents to be ingested are needed. How to think about giving the assistant write access to the CMMS or MES is covered in “Industrial AI Agent Security“, and how to connect it to systems is covered in “Connecting MES and ERP to AI with an MCP Server“.

Make it impossible for the AI to skip lockout

The most important thing for safety in maintenance work is to put the equipment in a state where it cannot move before the work starts.

In Thailand, Clause 7 of the Ministry of Labour’s “Ministerial Regulation Prescribing Standards for the Administration and Management of Occupational Safety, Health and Working Environment Concerning Machinery, Cranes and Boilers B.E. 2564 (2021)” requires that clearly visible signs be posted where machinery is installed, repaired or inspected, that mechanisms, methods or devices be provided to prevent the machinery from operating, and that a “Do not switch on” tag or sign be hung on the machine’s switch (the English wording of the provisions is our own summary).

For reference, OSHA 29 CFR 1910.147 (the control of hazardous energy, commonly called lockout/tagout), which is a US regulation, requires employers to have a programme consisting of energy control procedures, training and periodic inspections, and sets out the sequence of preparation for shutdown, shutdown, isolation, applying locks or tags, releasing stored energy, and verifying isolation. It does not apply directly to factories in Thailand, but the idea of documenting the procedure and always using it is a useful reference.

In designing a maintenance AI assistant, we consider it important to observe the following three points.

  • When presenting a procedure that involves touching the inside of equipment, always include the stopping, isolation, lockout and tagging steps at the very beginning. Leave no room for the AI to answer “can be skipped if in a hurry”
  • Lockout is carried out and verified by people. The AI only presents the procedure and is not a substitute for carrying it out
  • Use “answers that skip safety procedures” as a pass/fail metric in acceptance testing

The stronger the wish to shorten time, the more tempting it becomes to cut the 10 minutes of safety. That is exactly why there is value in writing “this is excluded” at the outset, in both the estimate and the acceptance test.

Cost and ROI: The Value of “Organizing Records” in a Model Estimate

Maintenance AI Assistant: Cut Diagnosis Time with RFP and FAT/SAT - figure 2

From here we compare two configurations at model factory M. To repeat, all figures are original estimates and assumptions (placeholder values) for this article, and no primary information on vendor prices or effects has been used.

Common assumptions (model factory M)

ItemPlaceholder value
Location and industryJapanese-owned automotive parts plant in eastern Thailand (Rayong Province)
Equipment60 machines (presses, welding, machining, plastic moulding)
Maintenance staff12 (4 veterans, 8 young Thai technicians with less than 3 years’ experience)
Maintenance calls (failure response)1,200 per year, of which 600 per year are failures that actually stop the line
Loss per hour of line stoppage (lost marginal profit)THB 6,000/hour
Average time to restoration per failure90 min (safety 10 min + localization and diagnosis 40 min + correction 35 min + function checkout 5 min)
Escalations calling in a veteran240 per year (including nights and holidays), THB 1,500 per case (call-out allowance, overtime)

The cost of escalations is 240 cases × THB 1,500 = THB 360,000 per year. For reference, the loss from current diagnosis time is 600 cases × 40 min = 24,000 min = 400 hours, and 400 hours × THB 6,000 = THB 2,400,000 per year.

Effects are calculated only for the 600 line-stopping cases, and the effect for the remaining 600 cases where the line does not stop is conservatively set at 0.

Configuration A: have it read manuals, alarm code lists and past records as they are

The initial investment is as follows.

ItemAmount (THB)
Building the assistant platform (RAG, permissions, Thai/Japanese/English UI, CMMS reference)600,000
Document ingestion (preparing operating manuals, alarm code lists and procedures, and version control)200,000
Building the evaluation set (past failure cases) and FAT/SAT200,000
Initial total600,000 + 200,000 + 200,000 = 1,000,000

Annual operation (LLM usage fees, hosting, maintenance) is set at THB 180,000/year.

The assumed effects are as follows. Because misspellings, abbreviations and inconsistent codes remain in past records, we assume that searching for similar cases works poorly, and limit the reduction in diagnosis time to 40 min → 32 min (−8 min).

  • Diagnosis time: 600 cases × 8 min = 4,800 min = 80 hours, 80 hours × 6,000 = THB 480,000
  • Escalations: 20% reduction. 240 cases × 20% = 48 cases, 48 cases × 1,500 = THB 72,000
  • Total effect: 480,000 + 72,000 = THB 552,000/year
  • Annual net benefit: 552,000 − 180,000 = THB 372,000/year
  • Simple payback: 1,000,000 ÷ 372,000 = 2.688… → about 2.7 years

Configuration B: Configuration A plus organizing past records (failure code system) and an input template

ItemAmount (THB)
Full Configuration A (including evaluation set and FAT/SAT)1,000,000
Recoding 3,600 work records from the past 3 years (classification by equipment, component, failure mode, cause, action) 3,600 records × 100360,000
Work record input template (selectable failure codes + free text) and CMMS integration240,000
Evaluation and FAT/SAT for the additions100,000
Additional total360,000 + 240,000 + 100,000 = 700,000
Initial total1,000,000 + 700,000 = 1,700,000

Annual operation, including maintaining the code system, is set at THB 220,000/year.

The effects replace the effects of Configuration A and are not added to them. The −15 min in diagnosis time includes A’s −8 min, and the 40% reduction in escalations includes A’s 20% reduction.

  • Diagnosis time: 40 min → 25 min (−15 min). 600 cases × 15 min = 9,000 min = 150 hours, 150 hours × 6,000 = THB 900,000
  • Escalations: 40% reduction. 240 cases × 40% = 96 cases, 96 cases × 1,500 = THB 144,000
  • Total effect: 900,000 + 144,000 = THB 1,044,000/year
  • Annual net benefit: 1,044,000 − 220,000 = THB 824,000/year
  • Simple payback: 1,700,000 ÷ 824,000 = 2.063… → about 2.1 years
  • Average time to restoration: 90 min → 10 + 25 + 35 + 5 = 75 min (−15 min, about 16.7% reduction). The 10 minutes for putting the equipment in a safe state are unchanged

Comparison table

ItemConfiguration AConfiguration BConfiguration A′ (see below)
Initial investment (THB)1,000,0001,700,000800,000
Annual operation (THB/year)180,000220,000180,000
Diagnosis time40 min → 32 min40 min → 25 min40 min → 38 min
Escalation reduction20% (48 cases)40% (96 cases)0 cases
Total effect (THB/year)552,0001,044,000120,000
Annual net benefit (THB/year)372,000824,000−60,000
Simple paybackAbout 2.7 yearsAbout 2.1 yearsNot recovered
5-year cumulative (THB)860,0002,420,000−

Key point 1: organizing records is not an “AI add-on” but the core that speeds up payback

Compare the 5-year cumulative totals.

  • Configuration A: 372,000 × 5 − 1,000,000 = 1,860,000 − 1,000,000 = THB 860,000
  • Configuration B: 824,000 × 5 − 1,700,000 = 4,120,000 − 1,700,000 = THB 2,420,000

Looking only at B’s increment, the additional initial investment is THB 700,000 and the additional annual net benefit is 824,000 − 372,000 = THB 452,000/year. Payback on the increment is 700,000 ÷ 452,000 = 1.548… → about 1.5 years, and the difference in the 5-year cumulative total is 452,000 × 5 − 700,000 = THB 1,560,000 (= 2,420,000 − 860,000).

B has the larger initial investment, yet B pays back sooner. Under the assumptions of this estimate, organizing records is not a “bonus” to the AI but the core that speeds up the return on investment. Put the other way, introducing only the AI without organizing records means running a poorly performing search on top of an expensive platform.

Key point 2: what if the evaluation set and acceptance testing are dropped (Configuration A′)

When budgets are cut, the first candidate tends to be the THB 200,000 for the evaluation set and acceptance testing. Consider Configuration A′, where this is cut.

  • Initial investment: 1,000,000 − 200,000 = THB 800,000
  • Assumption: wrong answers right after go-live make the maintenance technicians stop trusting the AI, and it falls out of use. Diagnosis time drops by only 2 min. 600 cases × 2 min = 1,200 min = 20 hours, 20 hours × 6,000 = THB 120,000. Escalation reduction is 0
  • Total effect THB 120,000/year − operation 180,000 = THB −60,000/year → not recovered

Cutting THB 200,000 drops the annual net benefit from THB 372,000 to THB −60,000. On top of that, there remains a risk that cannot be put in monetary terms: nobody can detect it if answers that skip safety procedures are being produced. Evaluation and acceptance are not a contingency budget you can cut.

Notes on using the estimate

  • There is only one baseline scenario. 600 cases per year, 40 min of diagnosis, THB 6,000/hour, 240 escalations and THB 1,500 per case are the common starting point and are not changed between configurations
  • The effects of Configuration B replace those of Configuration A; do not add the effects of A and B together
  • When using this for your own factory, first measure actual diagnosis time and escalation counts from your failure response records, and replace the placeholder values

12 Items to Write in a Maintenance AI RFP

In the request for proposal (RFP), write “what will be measured, how, and what counts as a pass” rather than a list of features. Because vendor-reported figures do not use consistent metrics, it can be argued that the buyer needs to specify the definitions of measurement.

No.ItemExample of what to write
1Scope of target equipment and documentsTarget equipment (starting with representative machines), types and number of target documents, what is out of scope
2Reference data, updates and version controlWhich of the operating manuals, alarm code lists, work records and CMMS are referred to, and how. The replacement procedure on revision, and a mechanism that prevents answers from old versions
3Languages and glossaryQuestions and answers in Thai, Japanese and English. How the glossary of shop-floor abbreviations and terms is built and who maintains it
4Source display in answersAlways attach document name, page and record number to answers that contain procedures or numbers
5Handling of safety proceduresDo not allow stopping, isolation, lockout and tagging steps to be omitted. State on screen that carrying them out and verifying them is assumed to be done by people
6Behaviour when it does not knowWhen no basis is found, answer “I don’t know” and show the next action, such as contacting a veteran
7Permissions and access control, prompt injection countermeasuresViewing scope per user, approval of documents to be ingested, whether it has write access to systems
8Handling of personal data (PDPA)Handling of maintenance technicians’ names and the like contained in work records, storage location, retention period (confirm legal judgments individually with experts)
9How effects are measuredDefine the breakdown of MTTR, the start and end points of diagnosis time, and how escalations are counted
10Evaluation set and pass criteriaHow the evaluation set from past failure cases is built, number of cases, metrics and target values
11Logs, audit and continuous evaluationLogs of questions, answers and sources; periodic re-evaluation; re-testing when the model or documents change
12Local support and operating structureSupport for enquiries in Thai, on-site incident response in Thailand, the role of the factory’s own operations staff

Item 9, “how effects are measured”, is particularly important. When you receive a proposal saying “we will shorten MTTR”, always check whether that MTTR means time to restoration or repair time, whether it includes safety time and waiting for parts, and where diagnosis time starts and ends. When a vendor shows figures from other companies’ cases, we also recommend asking what those figures measured (a pilot case, a target, or something that includes partially correct answers).

On contracts, it is also useful to note that NIST AI 600-1 recommends contracts and SLAs that set quality standards, and clauses that allow evaluation of third-party generative AI processes.

What to Check at Maintenance AI FAT/SAT

Build an evaluation set

The core of acceptance testing is an evaluation set built from past failure cases. As an example, build it as follows (the figures are placeholder target examples, not standard values).

  • Prepare 100 past failure cases with correct answers (the actual cause and action)
  • Mix enquiries in Thai, Japanese and English. Include questions that use shop-floor abbreviations
  • Include cases whose answer is not in the records, in other words cases where the right response is “I don’t know”
  • Run the same set in the factory acceptance test (FAT) before go-live and the site acceptance test (SAT) on site

Metrics to look at

MetricWhat is checkedExample target (decided by the factory)
(1) Top-3 hit rate for candidate causesShare of cases where the correct answer is among the top 3 candidates70%
(2) Source attachment rateShare of answers containing procedures or numbers that carry document name, page and record number100%
(3) Omission of safety proceduresNumber of answers that skip stopping, isolation or lockout steps0 cases
(4) Correctness of “I don’t know”Share of cases that should get “I don’t know” where it did answer “I don’t know”Decided by the factory
(5) Response timeTime from question to answerDecided by the factory
(6) Correctness of Thai answersWhether answers to Thai questions read correctly, judged by Thai maintenance techniciansDecided by the factory

The 70% in (1) is only an example target. The factory itself decides what percentage would be useful on the shop floor for its own equipment and records. We consider it safe to treat (2) and (3) as metrics that allow no exceptions.

What to test alongside

  • Behaviour on wrong answers: deliberately ask about failures that are not in the records and see whether it makes up plausible stories. A person should also open the documents cited as sources and confirm that they really contain that content
  • Judging Thai answers: do not judge with Japanese staff only; have Thai maintenance technicians actually read the answers and check whether the meaning comes through and the order of steps is correct
  • Offline and loss of connection: how does it behave when the factory network or the connection to the cloud is lost? Are alternative means for the time it is unavailable (paper procedures, contacting a veteran) decided?
  • Version switch-over: after an operating manual is replaced with a new version, does it still answer with content from the old version?

After go-live too, run the same evaluation set periodically to check that the quality of answers has not dropped due to added documents or model updates. How to proceed with continuous evaluation and reauthorization is explained in “AI Audit: Continuous Evaluation and Reauthorization for AI Agents“.

Issues Specific to Thailand and ASEAN

1. The ministerial regulation: mechanisms that keep machines under repair from operating, and the language of manuals

Clause 7 of the B.E. 2564 ministerial regulation mentioned above requires mechanisms that keep machinery under repair from operating, and a “Do not switch on” tag. Clause 8 requires following the specifications and operating manuals set by the manufacturer for assembly, installation, use, repair, maintenance, inspection and so on of machinery, and where these do not exist, having an engineer prepare them in writing. The specifications and operating manuals must be “in Thai, or in another language that employees can read and work safely with”.

Even if the AI can present procedures in Thai, it has not been confirmed whether its answers can substitute for an “operating manual” as meant in the regulation. The AI’s answers do not replace the mechanisms required by the regulation, and the safe approach is a design in which the AI refers to the underlying manufacturer’s manual and shows its sources. For that reason too, the language and version of the manuals the AI refers to need to be put in order. The English wording of the provisions is our own summary; please confirm specific application individually with the competent authorities and experts.

2. PDPA: personal information in work records

Thailand’s Personal Data Protection Act (PDPA) came fully into force on 1 June 2022. According to an overview by the law firm Herbert Smith Freehills Kramer, by September 2025 the Personal Data Protection Committee (PDPC) had imposed administrative fines totalling more than THB 21.5 million across 5 cases and 8 orders. The maximum administrative fine is THB 5 million.

Work records often contain the names of the workers. If you have the AI read the records, or aggregate the number of cases handled per technician from the records, please confirm individually with experts how this is treated under the PDPA.

3. Generative AI guidelines and the draft AI law

On 30 October 2024, Thailand’s Ministry of Digital Economy and Society (DE) and the AI Governance Center of the Electronic Transactions Development Agency (ETDA) jointly published the “Generative AI Governance Guideline for Organizations”. It is a guideline with no legal force.

In addition, ETDA published a new draft law on AI in July 2026, and the public hearing closed on 14 August 2026. According to commentaries by law firms and others, the draft is structured by risk, dividing AI into categories such as prohibited AI and high-risk AI; high-risk AI is subject to meaningful human control, and those deploying AI are required to assign oversight personnel. It is still at the draft stage and has not been enacted. It has not been confirmed whether a factory maintenance assistant would count as high-risk AI, so please confirm individually with the competent authorities and experts.

There is also ISO/IEC 42001:2023, the international standard for AI management systems. It covers not only organizations that develop and provide AI but also organizations that use AI, and certification is not mandatory. It is a useful reference when putting in place internal rules on AI use.

4. Multiple languages and shop-floor terminology

In Thai factories, Thai, Japanese and English are mixed, and there are many shop-floor abbreviations and in-house names. If the same part is recorded under different names in each language, neither search nor evaluation works well. The countermeasures are recoding records (failure codes do not depend on language) and a three-language glossary. The glossary is not something you make once and are done with; decide who is responsible for updating it each time new equipment or parts come in.

5. BOI measures

The BOI’s “Smart and Sustainable Industry” measure covers machinery upgrades, digital technology, and the introduction of automation and robots. However, it has not been confirmed whether a maintenance AI assistant on its own would be eligible under this measure. If you are considering it together with equipment upgrades or the like, please confirm individually with the BOI or experts.

90-Day Plan for a Maintenance AI Assistant Implementation

Maintenance AI Assistant: Cut Diagnosis Time with RFP and FAT/SAT - figure 3

What you should do before placing an order is measure, and run a prototype that organizes records. Divide the 90 days into three phases.

Days 0–30: measure and take inventory

  • From failure response records, measure the number of failures that stop the line, the diagnosis time for each case, and the number and time of day of escalations. Decide the start point (the call) and end point (the moment the cause is identified) of diagnosis time, and have the maintenance technicians record them
  • Take inventory of the target equipment and documents. What language are the operating manuals in, and which version is on the shop floor? Are the alarm code lists complete?
  • Look at the state of past work records: number of records, languages, whether failure codes exist, the quality of the free text

Days 31–60: prototype with 10 representative machines

  • Choose 10 representative machines with frequent failures and recode their work records (equipment, component, failure mode, cause, action)
  • Build an evaluation set from past failure cases on the same 10 machines. Veterans check the correct answers
  • With a prototype that reads the manuals, alarm lists and recoded records, look at the difference between organized and unorganized records

Days 61–90: finalize the RFP and acceptance criteria, and decide on the order

  • Replace the placeholder values in this article’s estimate with the measured values from days 0–30, and decide whether to order Configuration A (reading as is) or Configuration B (organizing records)
  • Write the 12 RFP items, and decide the FAT/SAT evaluation set and pass criteria
  • Share the decision materials (measured values, prototype results, estimate) with approvers such as the plant manager

If you are considering combining this with efforts such as predictive maintenance, which uses sensors to catch signs of failure, please also refer to “Predictive Maintenance Case Studies“. A maintenance AI assistant is a mechanism for “bringing equipment back quickly after it stops”, and predictive maintenance is a mechanism for “acting before it stops”; both share the same foundation in failure records.

Frequently Asked Questions (FAQ)

Q1. What can a maintenance AI assistant do, and what can it not do?

It can search operating manuals, alarm code lists and past work records, and present candidate causes and check procedures with sources. It can also produce draft work records and candidate failure codes. On the other hand, it cannot judge the actual condition of the equipment, guarantee the right answer for failures not in the records, or make the final decision. You also must not let it suggest shortcuts that skip safety procedures such as stopping, isolation and lockout.

Q2. How much can MTTR be reduced by using AI for maintenance troubleshooting?

There is no single answer. The figures vendors publish use inconsistent metrics, such as reactive maintenance time, the generation rate of remedy proposals including partially correct ones, and targets, and cannot be compared as MTTR values. Because what AI mainly shortens is the localization and diagnosis time within MTTR, first measure diagnosis time in your own failure response. In this article’s model estimate, we assumed that the configuration with organized records would bring the average time to restoration from 90 minutes to 75 minutes (about a 16.7% reduction), but these are placeholder values.

Q3. Isn’t it dangerous if generative AI gives a wrong procedure?

Yes, it is dangerous. NIST AI 600-1 also warns that generative AI produces plausible errors and can even fabricate the basis or sources themselves. The countermeasures are to always attach sources to answers containing procedures or numbers, not to let it skip safety procedures, to have people make the final decision, and to run acceptance tests with an evaluation set built from past failure cases and confirm that there are 0 answers that skip safety procedures.

Q4. Can it be used even if past maintenance records are handwritten or inconsistent?

You should assume that it will not work well as they are. With misspellings, abbreviations, mixed languages and inconsistent failure codes, similar cases cannot be found well. In a study of wind power maintenance records, about one in five records lacked sufficient information to determine the failure mode. We recommend digitizing handwritten records, recoding failure codes starting with representative machines, and keeping future records with a template that combines selectable failure codes and free text.

Q5. What should be checked in a maintenance AI RFP and acceptance tests (FAT/SAT)?

In the RFP, write 12 items: scope, reference data and version control, languages and glossary, source display, handling of safety procedures, behaviour when it does not know, permission management, PDPA handling, how effects are measured, evaluation set and pass criteria, logs and continuous evaluation, and local support. In acceptance testing, use an evaluation set built from past failure cases and check the hit rate for candidate causes, the source attachment rate, the number of answers that skip safety procedures, the correctness of “I don’t know” answers, response time, and the correctness of Thai answers.

Q6. What should we watch out for regarding regulations and the PDPA when implementing at a Thai factory?

The B.E. 2564 Ministry of Labour regulation requires mechanisms that keep machinery under repair from operating, a “Do not switch on” tag, following the manufacturer’s operating manual, and having manuals in Thai or a language employees understand. AI answers do not replace these. Please confirm individually with the competent authorities and experts how work records are treated under the PDPA if they contain technicians’ names, as well as whether the assistant would fall under the draft AI law (published in July 2026, at the draft stage) and whether it would be eligible under BOI measures.

Summary

  • What a maintenance AI assistant shortens is the localization and diagnosis time within repair time. Safety time, such as stopping, isolation and lockout, is time that must not be shortened
  • Because the definition of MTTR can differ between organizations, measure its breakdown and make diagnosis time an independent metric
  • What determines the quality of answers is the quality of past repair records more than the AI model. Organize failure codes and improve future records with an input template
  • In this article’s model estimate (placeholder values), Configuration B with organized records paid back sooner (about 2.1 years) despite the larger initial investment, while Configuration A′, which skipped evaluation and acceptance, did not pay back
  • Write “what will be measured, how, and what counts as a pass” in the RFP, and run FAT/SAT with an evaluation set built from past failure cases
  • Confirm the treatment under the Thai ministerial regulation, the PDPA, the draft AI law and BOI measures individually with the competent authorities and experts

The first step is not choosing an AI, but measuring diagnosis time and escalations from your failure response records. At TOMAS TECH, we are happy to help from the stage before ordering, such as taking inventory of failure records and building evaluation sets. If you would like to check together how usable your own records are, please feel free to contact us via our contact page.

References