Near-miss report AI analysis, using generative AI to translate, summarize and classify reports, is the option that comes up when safety officers (จป.), general affairs / OSH section managers, and plant managers at Japanese-owned factories in Thailand tell us: “Near-miss reports come in as Thai handwriting on paper or as LINE messages. Just reading them, putting them into Japanese, classifying them, entering them into the register and compiling the monthly report takes all our time. We do count them, but we cannot follow up on whether the countermeasures were actually closed.”
Here is the conclusion up front. Whether near-miss report AI analysis delivers results is not decided by “the AI’s classification accuracy.” It is decided by “whether reporting becomes easier so that the number of reports rises, and whether the time until a countermeasure is assigned and closed gets shorter.” No matter how much faster classification and counting become, accidents will not go down unless owners, deadlines and completion checks for countermeasures actually work.
There are five things to decide in an implementation: (1) the reporting channel, (2) the scope handed to AI, (3) how to create the correct classification (ground truth), (4) countermeasure assignment and completion checks, and (5) how personal data and statutory reporting are handled. Deciding them in this order keeps product selection, the RFP and the acceptance tests from drifting.
Note that all report counts, processing times, labor costs, costs, accident losses and payback periods for the factory in this article are this article’s own estimates and assumed (placeholder) values based on the model factory described below. They are neither industry averages nor survey figures. No vendor prices are used in the calculation either. Read it as a “calculation template” to be replaced with your own actual figures and quotes.
Why Near-Miss Report AI Analysis Is on the Table for Thai Factories Now
Four walls: paper, LINE, multiple languages and the monthly report
At Japanese-owned factories in Thailand, most near-miss reports arrive in Thai. Some are handwritten on paper report forms, some are sent by team leaders over LINE as a photo with a short sentence, and some are told verbally and written up by the safety officer on the worker’s behalf. Formats and levels of detail vary widely. Depending on the factory, reports in other languages, such as the native languages of workers from neighboring countries, are also mixed in.
To turn these into something Japanese managers can read, someone has to read each report one by one, grasp its meaning, put it into Japanese, classify it by accident type and location, and enter it into the register. At the end of the month, they count the reports, make charts and compile the monthly report. Much of the safety officer’s time disappears into this “read, translate, sort, count” work, leaving little time for what really matters: “thinking about where the danger is, taking countermeasures, and confirming they were closed.”
Even more troublesome, the more reports come in, the harder it is for processing to keep up. In principle, the more near-misses are collected, the easier it should be to see where the hazards are; but from the processing side, more reports also mean more workload. As a result, an unspoken feeling that “fewer reports make life easier” takes hold, and the reporting channel gradually narrows. This is where the value of using AI to reduce the effort of translation, summarization and classification lies.
Why collect near-miss reports at all
The Japanese Ministry of Health, Labour and Welfare (MHLW) “Workplace Safety Site” describes hiyari-hatto (near-miss) activities as taking up experiences that gave someone a fright or a start and linking them to accident prevention, and says they are effective as one way of identifying hazards. At the same time, it notes that reporting a near-miss is “not something to be very proud of” for the person involved, so unless an agreement not to blame workers is made and carried out, the system will not last. It adds that even when the work was not done according to the procedure manual, this should be seen as an opportunity to review the manual, and that sharing improvement cases widely, for example in company newsletters, helps spread improvements to other lines and sites and raises awareness.
In the context of near-misses, Heinrich’s law is often cited. The glossary on the same site introduces it as a law published by Heinrich, a safety engineer at an American property and casualty insurance company: of 330 accidents caused by the same person, 1 was a serious accident, 29 were minor injuries and 300 were accidents without injury. However, the same explanation also notes that Heinrich himself said this ratio differs between, for example, steel erection and clerical work, and concludes that what matters is not the numbers in the ratio but the fact that there are many hazards.
The ratio has also been criticized. Safety+Health, the magazine of the U.S. National Safety Council (NSC), presents Fred A. Manuele’s criticism: the original files no longer exist and cannot be verified by third parties; the ratio was revised from edition to edition without explanation; and starting from the premise that “most accidents are caused by unsafe acts” points countermeasures at workers rather than at the work system. It is reasonable to treat the ratio as something introduced as a rule of thumb. It cannot be used as a tool for calculating that reducing near-misses by so many cases will reduce serious accidents by so many.
In this article, too, the goal of near-miss AI analysis is not placed on “reducing the number of reports.” The goal is to learn quickly where the hazards are, take countermeasures, and confirm that they have been closed.
Developments in 2025 to 2026: AI that creates reports from video, and AI that reads every report
In the near-miss x AI field, two broad lines of development can be seen (this two-line framing is this article’s own).
The first is the line that finds “near-misses nobody reported” from video and sensors. On March 25, 2026, NEC and Sumitomo Heavy Industries announced joint development of a system that extracts near-misses from the video of cameras mounted on hydraulic excavators and from sensor information, and automatically creates reports in which generative AI summarizes how the event developed. They completed technical verification in September 2025, and aim for technology development and business verification in fiscal 2026 and practical use in fiscal 2027. This is the plan as of the announcement and is not yet a commercialized product. Its target is also construction machinery video, not a system for classifying the text of factory reports. At a construction and surveying exhibition in June 2026 it was introduced as the “near-miss automatic detection report creation function HAW (tentative name),” and press coverage cited, as background, the problem that near-miss reports are hard to get from construction sites. We covered the idea of detecting dangerous situations with factory cameras in AI safety cameras.
The second is the line that reads every human-written text report and classifies and summarizes it. Nomura Research Institute (NRI) has published a solution that uses natural language processing to extract and structure near-miss-related elements from work reports, and states that there is a case in which analysis that used to be done on samples was expanded to full analysis of approximately 24,000 reports per year, with sharing on dashboards also automated, cutting about 600 hours per year. This is the figure from a single implementation case published by the vendor, and it does not mean the same effect will appear at every factory. Overseas, Benchmark Gensuite promotes AI functions such as automatically identifying potentially serious events, in addition to hazard reporting via mobile forms. This, too, is the vendor’s claim: no accuracy figures are published, and we could not confirm support for Thai or Vietnamese.
This article mainly deals with the second line: analyzing text reports with AI.
What research has shown so far
On the research side, in 2025 Ando and colleagues at the University of Occupational and Environmental Health, Japan, published results of automatically classifying occupational accident reports with large language models (LLMs). From 2021 data on occupational accidents requiring 4 or more days of lost work, they took 2,619 deaths and injuries caused by “falls on the same level” in the healthcare and retail industries, classified them with 4 LLMs, and compared the results with experts’ manual classification. Newer-generation models generally exceeded a kappa coefficient (agreement after subtracting chance agreement) of 0.7 in many categories other than “agent of injury,” and accuracy for location (indoor or outdoor) exceeded 91%. In the hardest category, agent of injury, o4-mini’s kappa coefficient was 0.662. Processing all cases took about 90 minutes and cost about 11 dollars, according to the study.
However, this study covered reports of occupational accidents that actually happened, not near-misses. The reports were in Japanese and limited to falls in 2 industries. You cannot generalize this to “LLM classification accuracy is 90%.” How closely AI agrees on your own reports, using your own classification scheme, can only be measured in acceptance testing.
In maritime transport, a study by the University of Strathclyde in the UK and the ferry operator CalMac Ferries analyzed 4,360 near-miss reports covering 5 years with topic modeling, used GPT to assist in labeling the topics, and then verified the output through a structured review by multiple experts. They extracted more than 75 incident topics and 87 preventive-measure topics, and showed seasonal variation, such as passenger-related risks being more common in summer and equipment failures in winter, as well as links between incidents and countermeasures. This study is about shipping and does not report accuracy figures, but it is a useful reference for a design in which “labels assigned by AI are verified by people.”
In addition, a preprint study (ABEX-RAT) on construction accident reports in the United States names “severe class imbalance,” where rare but serious accident types become minority classes, as a challenge. The more serious and less frequent an accident type is, the harder it tends to be for AI to classify. This is why, in the acceptance tests described later, “missed serious events” are given priority over overall agreement.
Context from Japanese statistics
For reference, let us look at Japanese figures. All of the following are Japanese statistics, not Thai figures. According to the occupational accident statistics for Reiwa 7 (2025) published by the MHLW on May 27, 2026 (excluding COVID-19 infections), there were 700 deaths, the lowest on record, and 135,333 deaths and injuries requiring 4 or more days of lost work, down 0.3% from the previous year. By industry, manufacturing had the most deaths and injuries at 26,371. By accident type, “falls on the same level” were the most common cause of deaths and injuries across all industries at 37,195, while among deaths, “caught in or between” accounted for 117, up 6.4% from the previous year.
Japan’s 14th Occupational Accident Prevention Plan sets a target of reducing deaths and injuries from “caught in or between” accidents involving machinery in manufacturing by 5% or more by Reiwa 9 (2027) compared with Reiwa 4 (2022). There are types with many cases, such as falls, and types with fewer cases but serious outcomes, such as caught in or between. Near-miss analysis needs a design that overlooks neither.
What Is Near-Miss AI Analysis: What AI Handles and What People Decide

A 4-stage flow from report to countermeasure
Near-miss AI analysis is easier to organize when thought of in the following 4 stages.
- Reporting: A worker reports a near-miss event via smartphone or on paper
- AI preparation: Translation, summarization, suggested classification tags, merging duplicates and estimating the location
- Counting and visualization: Counting by accident type, location and time of day, and finding the places where reports concentrate (hotspots)
- Countermeasures and completion checks: Deciding an owner and deadline, taking countermeasures, confirming their effect and closing them
The figure above illustrates this 4-stage flow. AI works mainly in stages 2 and 3; stages 1 and 4 are determined by how people and systems are designed. Implementation discussions tend to focus on AI performance because stage 2 is the most visible, but what actually determines results is stages 1 and 4.
The scope you can hand to AI
The tasks that are easy to hand to AI are “preparation” tasks such as the following.
- Translation: Reports in Thai, Vietnamese, English and other languages into Japanese, or Japanese instructions into local languages
- Summarization: Turning long LINE messages or handwritten reports into the form “when, where, what, and what almost happened”
- Suggested classification: Attaching candidate tags for accident type, agent of injury, location and expected severity
- Merging duplicates: Proposing candidates to merge when several people report the same event
- Estimating the location: Estimating the factory area from shop-floor expressions such as “behind the press on Line 2” or “by the pillar in the shipping area”
- Prompting for missing details: Asking back at the input stage when a report lacks the location or time
All of these are tasks that “propose candidates on the premise that a person makes the final check.” AI does not finalize the classification; the safety officer reviews each report and approves or corrects it.
The scope people decide
On the other hand, the following judgments stay with people.
- Final severity judgment: Whether the event could easily have led to a serious accident
- Deciding the countermeasure: Whether to stop it with equipment, change the procedure, or supplement with signage and training
- Assigning owner and deadline: Who does what by when
- Confirming completion and effect: Whether the countermeasure was implemented and whether the same report has come up again
- Statutory reporting: When an accident actually occurs, notifying the competent authority and submitting the written report
The reason severity judgment is not handed to AI is, as mentioned above, that AI tends to struggle more with serious but infrequent types. The AI’s severity estimate is used to “sort the reports that people should look at first,” while the judgment itself is made by a person. As for countermeasure approaches, we cover making mistakes impossible on the equipment side in poka-yoke, and using AI to support writing corrective actions in corrective action report AI.
Designing Multilingual Near-Miss Reporting: A 30-Second Channel and a No-Blame Agreement
If the channel is narrow, AI finds nothing
AI analysis only works on the reports that come in. If the reporting channel is hard to use, the number of reports will not grow, and analysis will not reveal where the hazards are. This is why the channel is the first thing to decide in an implementation.
As a guideline, it helps to use whether a worker can submit a report from a smartphone in 30 seconds as the standard. Specifically, a design like the following.
- Workers can report directly from the apps they already use (such as LINE or Teams)
- They can send it just by taking 1 photo and adding a few words. The text can be short
- They can write in their own language, such as Thai, Vietnamese or Japanese. Speaking and converting voice to text can also be an option
- The location can be set by scanning an area QR code or picking it from a map
- The reporter’s name is not mandatory. An anonymous option remains
Paper report forms can stay as well. Some areas do not allow smartphones, and some people are more used to writing. Paper reports can be photographed and read by AI, and once a person checks the result, they can go into the same register. We also covered the idea of digitizing paper daily reports and inspection records in factory daily report AI and inspection checklist digitalization.
Put the “no-blame agreement” in writing in the local language
As the MHLW explanation says, without an agreement not to blame reporters, a near-miss system will not last. The same holds in Thailand. If anything, because of the language barrier, it is harder for the shop floor to know how Japanese managers receive reports, and the worry that “I might get scolded if I write this” tends to linger.
So put the following in writing in Thai and other local languages, and repeat it in morning meetings and safety training.
- Nobody will be rated lower or reprimanded for reporting a near-miss
- Even if it turns out the work was not done according to the procedure, the procedure and equipment are reviewed first
- Every report gets a reply about what was done (“received,” “countermeasure decided,” “countermeasure completed”)
- Good reports, and improvements that came from reports, are shared and recognized within the company
An ISO 45001 commentary site explains that clause 5.4 (consultation and participation of workers) requires identifying and removing obstacles to participation, citing as examples failing to respond to workers’ suggestions, language or literacy issues, and policies that discourage participation. A multilingual channel and replies to reports are exactly this kind of “removing obstacles.” AI translation can be positioned as a means of lowering the language barrier.
Designing the Scheme for Automatic Near-Miss Report Classification: Accident Type, Agent, Location and Severity
Decide the 4 axes first
Before having AI classify reports, you need to decide on a classification scheme. If you simply ask AI to “classify these” without a scheme, each report gets tagged with different words and nothing can be counted.
The basics are the following 4 axes.
| Axis | Content | Examples |
|---|---|---|
| Accident type | What almost happened | Fall on the same level, caught in or between, fall from height, struck by, cut or abrasion, flying or falling object |
| Agent of injury | What almost caused it | Forklift, press machine, conveyor, hand cart, stairs, oil on the floor |
| Location | Where in the factory | Plant 1 press area, shipping area, warehouse aisle, outdoor loading entrance |
| Expected severity | How bad the damage would have been | Could lead to a lost-time accident, minor injury level, property damage level |
Aligning accident type names with the categories of Japan’s occupational accident statistics makes it easier to report to the Japanese head office and to compare with other sites. However, making them too fine-grained causes both AI and people to disagree more often. It is easier to operate by starting with around 10 main categories and adding categories when more cases start falling into “other.”
Build a mapping table between shop-floor words and management words
The biggest stumbling block in classification is the gap between shop-floor language and management language. A worker writes, “The lift backed up and nearly hit me.” In management terms, that is “struck by, agent of injury: forklift.” The place the shop floor calls “behind No. 2” might be “Plant 1 press area east aisle” on the drawings.
So create a mapping table that links expressions commonly used on the shop floor with the categories of the classification scheme: Thai expressions, Japanese expressions, and area nicknames alongside their names on the drawings. This table can be used as-is in instructions to the AI (prompts) and as reference examples for classification. How to check how accurately generative AI handles Thai text is covered in detail in evaluating Thai-language generative AI.
The ground truth is created by safety officers
To judge whether the AI’s classification is correct, you need a “correct answer.” The correct answer is the classification assigned by safety officers. As described in detail in the acceptance testing section, select several hundred past reports, have 2 safety officers classify them independently, and resolve disagreements through discussion. At this point, also record the rate at which the 2 disagreed. For categories where even people disagree, there is no point in demanding that AI agree more than people do. The agreement between people serves as a guide to the upper limit of what to require from AI.
From Near-Miss Data Analysis to Countermeasures: Finding Hotspots, Owners, Deadlines and Completion Checks

Look at hotspots as “location x accident type x time of day”
Once classified reports accumulate in the register, counting becomes possible. The most useful count is by location. The figure above illustrates a heat map overlaying the concentration of reports on the factory floor plan in color intensity, showing places where reports tend to concentrate, such as where forklift aisles and pedestrian walkways cross.
Once you find a hotspot, overlay accident type and time of day. Even at the same intersection, the right countermeasure differs depending on whether “near-contact” reports concentrate around shift changes or on the night shift. AI is well suited to quickly producing these combined counts and extracting common points from free-text descriptions. However, people decide what to read from the results and where to act.
If you only look at the types with the most cases, you will miss the serious ones. In Japanese statistics too, the most common cause of deaths and injuries is falls, yet the top causes of death include “caught in or between,” which has far fewer deaths and injuries than falls. Design the counting screen so that reports with high expected severity always appear at the top, separately from the ranking by number of cases.
Put “owner, deadline and completion check” behind classification
This is the part most easily overlooked in near-miss AI analysis. No matter how fast classification and counting become, if no owner and deadline are set for countermeasures and completion is not confirmed, the hazard remains as it is.
For each report, or for each hotspot, keep the following 4 items in the register.
- Owner: Who takes the countermeasure (a role rather than a personal name is fine)
- Deadline: By when
- Countermeasure: What will be done, and whether through equipment, procedures, signage or training
- Effect check: How reports of the same type at the same location changed after the countermeasure
An ISO 45001 commentary site explains that clause 10.2 requires responding to incidents and nonconformities in a timely manner, investigating them to determine the causes, evaluating the need for corrective action and implementing it, reviewing its effectiveness and keeping records, and that it also requires workers to participate in this process. In addition, in the same site’s paraphrase of the definition, incidents in ISO 45001 also include near-misses. Following this approach, it is natural to manage the near-miss register as one continuous flow through to implementing countermeasures and checking their effect.
AI can help in this flow too: notifying owners of countermeasures whose deadlines are approaching, putting overdue items at the top of the monthly report, and alerting when a similar report comes from the same place after a countermeasure.
The metric to watch is not “the number of reports”
After introducing AI, the metrics to track are things like the following.
- Time from a report until an owner is assigned
- On-time completion rate of countermeasures
- Rate at which reports of the same type at the same location come up again after a countermeasure was closed
- Share of high expected-severity reports whose countermeasures have been closed
- Time until the reporter receives a reply
The number of reports itself is not a metric to reduce. If the channel is made easy and the no-blame agreement gets across, it is natural for the number to rise. It is closer to reality to see an increase as a sign that hazards have become more visible. For an overall picture of how to measure the impact of AI adoption, see also measuring the effect of AI adoption.
Cost and ROI of Near-Miss AI Analysis: A Model Calculation
From here on is a calculation for a model factory. All figures below are this article’s own estimates and assumed (placeholder) values, and are neither industry averages nor survey figures. No vendor prices are used.
Common assumptions (Model Factory M)
Assume a Japanese-owned parts factory in central Thailand.
| Item | Assumed value | Calculation |
|---|---|---|
| Size | About 600 people including employees and on-site contractors, 3 shifts | – |
| Current number of reports | 1,800 per year (150 per month). Paper and LINE, mainly in Thai | – |
| Current processing per report | 10 minutes for reading, putting into Japanese, classifying and entering into the register | 1,800 x 10 / 60 = 300 hours/year |
| Current monthly report preparation | 16 hours per month | 16 x 12 = 192 hours/year |
| Current total | – | 300 + 192 = 492 hours/year |
| Processing per report after AI | 3 minutes of human checking | 1,800 x 3 / 60 = 90 hours/year |
| Monthly report preparation after AI | 4 hours per month | 4 x 12 = 48 hours/year |
| Total after AI | – | 90 + 48 = 138 hours/year |
| Safety officer labor cost (assumed, including social insurance etc.) | THB 350/hour | – |
| Initial cost | THB 600,000 (multilingual report form, LINE/Teams channel, classification scheme design, dashboard, register integration) | – |
| Annual cost | THB 240,000 (licenses, API, maintenance) | – |
| Loss per lost-time accident | THB 400,000 (placeholder total of treatment, replacement staff, production impact and investigation) | – |
The labor-saving effect is as follows.
- Hours saved: 492 – 138 = 354 hours/year
- Amount: 354 x 350 = THB 123,900/year
Comparing 3 cases
Annual net is “annual benefit – annual cost,” and payback period is “initial cost / annual net.” All amounts are in THB.
| Case | Annual benefit | Annual cost | Annual net | Payback period |
|---|---|---|---|---|
| A Labor saving in counting only | 123,900 | 240,000 | -116,100 | Does not pay back |
| B Labor saving + 1 lost-time accident prevented per year | 123,900 + 400,000 = 523,900 | 240,000 | 283,900 | 600,000 / 283,900 = about 2.1 years |
| C Labor saving + 1 lost-time accident prevented every 2 years (0.5 per year) | 123,900 + 200,000 = 323,900 | 240,000 | 83,900 | 600,000 / 83,900 = about 7.2 years |
Case A counts only the labor saving from translation, classification and the monthly report as the benefit. 123,900 – 240,000 = -116,100, meaning a net outflow of THB 116,100 every year. Labor saving in counting alone does not pay back.
Case B is where, in addition to the labor saving, 1 lost-time accident per year was prevented as a result of countermeasures taken on trends found by AI. Annual benefit is 123,900 + 400,000 = 523,900, annual net is 523,900 – 240,000 = 283,900, and payback is 600,000 / 283,900 = 2.11…, or about 2.1 years.
Case C is where the lost-time accidents prevented were 1 every 2 years (0.5 per year). The accident prevention benefit is 400,000 x 0.5 = 200,000, annual benefit is 123,900 + 200,000 = 323,900, annual net is 323,900 – 240,000 = 83,900, and payback is 600,000 / 83,900 = 7.15…, or about 7.2 years.
If the number of reports rises
Suppose that, as a result of making the channel easier, reports double to 3,600 per year. Checking time with AI is then 3,600 x 3 / 60 = 180 hours, plus 48 hours for the monthly report, for 228 hours/year. This is below the current 492 hours, so even with twice the reports, as far as the time spent checking reports and preparing the monthly report is concerned, the calculation shows a lighter workload for the safety officer than today.
However, these 228 hours are not added to the benefit amount. The increase in reports is itself a goal, and the benefit is counted through the labor saving and accident prevention in Cases A to C. If you added to the savings the time it would hypothetically take to process the extra reports manually as today (3,600 reports x 10 minutes, for example), you would be counting as “saved” time that is not actually being spent today, overstating the benefit.
Key takeaway: what decides payback is “accidents prevented”
What this calculation shows is that what decides payback is neither classification accuracy nor labor saving in counting, but “whether countermeasures taken on trends found by AI really prevented lost-time accidents.” The labor saving in Case A alone does not cover the annual cost; what makes the difference is the accident prevention benefit.
And the number of accidents prevented differs completely from factory to factory. The 1 per year and 0.5 per year used here are placeholders for the calculation, not evidence-based figures. Also, simply introducing AI does not reduce accidents. The accident prevention benefit arises only as the result of reports coming in, hotspots being found, owners and deadlines being set for countermeasures, and countermeasures actually being closed. If the mechanism behind classification does not work, this benefit is 0.
If, at the PoC (proof of concept) stage, you decide on full deployment based only on classification accuracy without measuring this “until countermeasures are closed,” you are likely to end up in the situation of Case A. For how to conclude a PoC and the criteria for deciding on full deployment, see also AI PoC exit criteria.
Notes to avoid double counting
- Labor saving is the safety officer’s time; accident prevention is the loss from accidents. They are different things, so it is fine to add them together
- Accident prevention is counted “only when accidents were prevented as a result of countermeasures taken from AI analysis.” Reductions unrelated to the AI implementation are not included
- The checking time when reports increase (228 hours) is not added to the benefit amount
- Translation savings are included in the 10 minutes of processing per report. They are not added separately
When calculating for your own factory, replace at least these 5 items with actual figures: “annual number of reports,” “processing time per report,” “monthly report preparation time,” “safety officer labor cost” and “loss per lost-time accident.” The number of accidents prevented cannot be known in advance. It is something to confirm after implementation by matching records of countermeasures against records of accidents.
12 Items to Include in an RFP for Near-Miss AI
When getting proposals from multiple vendors, write the following 12 items into the RFP to align the premises of the proposals.
| No. | Item | What to write |
|---|---|---|
| 1 | Reporting channel and languages | Types of channel such as LINE, Teams, web forms and photos of paper, supported languages (Thai, Japanese, English, Vietnamese, etc.), whether voice input is available, whether anonymous reporting is possible |
| 2 | Target data and migration of past reports | How many years of past reports to migrate, formats such as paper, Excel and LINE history, division of migration work |
| 3 | Classification scheme and the role of AI | Categories for accident type, agent of injury, location and severity, and a flow in which AI assigns candidates and a person finalizes |
| 4 | Handling of low-confidence cases | Whether AI can answer “I don’t know,” and the mechanism and threshold for routing low-confidence cases to people |
| 5 | Criteria for missed serious events | Evaluating misses of serious events such as caught in or between, work at height and forklift contact separately from overall agreement |
| 6 | Quality criteria for translation and summarization | Evaluating mix-ups of negation, quantities and locations, and a screen that shows original and translation side by side |
| 7 | Notification and owner assignment | Immediate notification of reports with high expected severity, whether automatic owner assignment is possible, recipients and means of notification |
| 8 | Countermeasure progress management | Whether owner, deadline, countermeasure and effect check can be managed in the register, and notification of overdue items |
| 9 | Personal data and PDPA | Masking of reporter names, names of those involved and injury information, access rights, retention period, processing before sending to AI, data storage location |
| 10 | Audit trail and change history | Who corrected the AI’s classification and when, and whether both the original report and the corrected classification can be kept |
| 11 | Output for statutory reports and safety committee materials | Output format for monthly safety committee materials and for information used as material for statutory reports (stating that it does not replace the statutory report itself) |
| 12 | Local support and Thai-language screens | Thai support for screens, forms and notifications, and inquiry support within Thailand with support hours |
Items 4, 5 and 6 in particular are often missing from proposals. A system in which AI silently attaches a plausible classification even to reports it is not confident about is more dangerous than one that honestly answers “I don’t know.” Even if overall agreement is high, missing a single serious event means the analysis cannot be trusted as a safety management tool. If translation turns “did not” into “did,” or “2 meters” into “20 meters,” even the direction of the countermeasure changes.
Item 9 is especially important when the design sends report text to an external generative AI service. Reports sometimes contain details of injuries or the names of those involved. For an overall picture of how to send information to generative AI and how data is handled, see also preventing generative AI data leaks.
What to Check in FAT/SAT for Near-Miss AI

Build the evaluation set first
Before acceptance testing, build an evaluation set. For example, extract 300 past reports. Make sure they mix Thai, Japanese, English and Vietnamese, and include photos of handwritten reports and short LINE messages. Have 2 safety officers assign classifications (accident type, agent of injury, location and expected severity), and treat these as the “correct answers.” Cases where the 2 disagree are decided through discussion, and the disagreement rate is also recorded. This is because the agreement between people serves as a guide to the upper limit of the agreement to require from AI.
The figure above illustrates the acceptance test: human classifications (left) and AI classifications (right) are placed side by side and matched, and cases that do not match get a warning for a person to check.
In FAT (testing in the vendor’s environment), provide the evaluation set to the vendor and test it in their environment. In SAT (testing with your own real data through the actual operational channel), run the same procedure through the real reporting channel. By running both with the same evaluation set and the same procedure, you can see in numbers any gap such as “it worked in the vendor’s environment but not through the shop-floor channel.” General thinking on acceptance testing for systems that incorporate AI is also covered in AI agent acceptance testing.
The target values below are placeholder examples. They are not standard values. Each factory decides its own pass criteria.
Metrics to check
- Agreement by category: Calculate the agreement rate with human classification and the kappa coefficient for each category. An example target is “close to the agreement between people for the main categories.” Do not look only at the overall average; look individually at categories that tend to split, such as agent of injury
- Number of missed serious events: Count the serious events, such as caught in or between, work at height and contact with forklifts, that AI missed. This metric takes priority over overall agreement, and an example target is 0
- Whether it can answer “I don’t know”: Whether there is a mechanism for AI to route low-confidence cases to people. Check that it is not producing plausible classifications despite low confidence
- Whether translation and summarization change the meaning: Check against the original for mix-ups in negation (did / did not), quantities (distance, weight, number of people) and location (which line, which aisle)
- Time from report to notification and owner assignment: Time from when a report with high expected severity comes in until the owner is notified and assigned
- Handling of personal names and injury information: Whether masking works, whether what is visible changes according to access rights, and whether data is deleted after the retention period
- Audit trail: Whether a record of who corrected the AI’s classification is kept along with the classification before correction
After go-live, people re-classify once a month to check
Passing acceptance is not the end. The way reports are written changes with the seasons and staff turnover, and the generative AI model itself may be updated. Even after go-live, once a month, randomly select 20 reports, have a person re-classify them, and check the agreement with AI. If agreement drops, review the mapping table and the reference examples for classification. Attaching this monthly check result to the safety committee materials lets the committee share how far the AI analysis can be trusted.
Issues Specific to Thailand and ASEAN
1. AI counting cannot replace statutory accident reporting
Section 34 of Thailand’s Occupational Safety, Health and Environment Act B.E.2554 (2011) sets out the employer’s duties when, among other things, a serious accident occurs at a workplace. According to an unofficial English translation, when an employee dies, the employer must notify the safety inspector immediately upon learning of it by telephone, fax or similar means, and report the details and cause in writing within 7 days of the date of death. When a fire, explosion, leak or other serious accident causes damage to the workplace, a production stoppage or injuries, the employer must notify immediately and submit a written report stating the cause and measures to prevent recurrence within 7 days of the accident. For injuries and illnesses covered by workers’ compensation, after reporting to the Social Security Office, a copy must be submitted to the safety inspector within 7 days. Violations carry a fine of up to 50,000 baht.
Near-miss AI counting cannot replace these statutory reports. Note that this section does not impose a duty to file near-misses themselves. The use of the AI register is, when an accident actually occurs, to quickly pull up related past reports and countermeasure history and gather the material for the statutory report sooner. The party responsible for reporting and the deadlines remain as set out in the law, and introducing AI does not change them. Please confirm the latest practice and how it applies to your company individually with the competent authority or experts.
2. How to feed AI counts into the safety committee and จป. statistical analysis
A 2022 ministerial regulation (B.E.2565) sets out the categories and assignment of safety officers (จป.). Under Section 25, workplaces with 50 or more employees must establish a safety, occupational health and environment committee within 30 days of the date they reach that number, and under Section 33 the committee meets at least once a month. The duties of จป. at the technical level include collecting statistics and submitting reports and recommendations to the employer, and at the senior technical and professional levels they also include analyzing the data.
AI counting can be positioned as a tool that speeds up this “collect, report and analyze statistics” work. For example, the monthly safety committee could be given a list of hotspots, the on-time completion rate of countermeasures, overdue countermeasures, reports with high expected severity and their response status, and the results of the monthly check of AI agreement. Reading the analysis and putting together recommendations remains the job of the จป. and the committee. Assignment categories vary by industry, so please confirm with the competent authority or experts, including which category your company falls into.
3. PDPA: injury and health information is sensitive data
Thailand’s Personal Data Protection Act (PDPA, B.E.2562) came into full force on June 1, 2022. According to an unofficial English translation, Section 26 in principle prohibits collecting sensitive personal data such as health data without the data subject’s explicit consent. Exceptions listed include preventing danger to life, body or health (when the data subject cannot give consent), when necessary for legal claims, and when necessary to comply with the law (such as preventive or occupational medicine, assessment of employees’ working capacity, employment protection and social security, all subject to appropriate safeguards). According to a commentary article, depending on the type of violation, administrative fines of up to 5 million baht or criminal penalties of imprisonment for up to 1 year are possible.
Near-miss reports sometimes contain injury details such as “cut my finger slightly” or “hurt my back,” or chronic conditions. This is information that may qualify as health data. Things to decide in the design are as follows.
- How reporter names and names of those involved are handled (not mandatory, with an anonymous option)
- Masking before sending to generative AI (names, employee numbers, injury details)
- Access rights (separating those who can see the full text from those who can see only the counts)
- Retention period and deletion procedure
Leave the judgment on which exception applies to legal professionals. The relationship between factory data and the PDPA is also covered in factory IoT and PDPA.
4. Multilingual: getting reports to Japanese managers without changing their meaning
When delivering Thai reports to Japanese managers, the scariest thing is the meaning changing in translation: a negation dropped, a quantity off by a digit, “right” and “left” swapped, a line number shifted. People who read only the translation cannot notice these mix-ups.
There are 2 countermeasures. One is to test translation mix-ups specifically in acceptance testing. The other is to always show the original and translation side by side on the register screen, so that staff who can read Thai can go back to the original at any time. At factories where reports written in other languages are mixed in, include reports in those languages in the evaluation set as well.
5. A no-blame agreement only gets across when written in the local language
An operating rule not to blame reporters will not reach the shop floor if it is only written in Japanese regulations. It needs to be put in writing in Thai and other local languages so that team leaders and supervisors can explain it in their own words. Replies to reports are also given in the reporter’s language. AI translation can also be used to lighten this work of “replying in the local language.”
6. Vietnam sites: Personal Data Protection Law, Law 91/2025/QH15
If you have a site in Vietnam, check the Personal Data Protection Law (Law No. 91/2025/QH15). It was passed on June 26, 2025 and took effect on January 1, 2026. According to an English translation in a legal database, the maximum fines are 5% of the previous year’s revenue for violations involving cross-border transfer, 10 times the illegal revenue for buying and selling data, and 3 billion dong for other violations. If you design the system to process near-miss reports from a Vietnamese factory in Thailand, Japan or an overseas cloud, check the procedures for cross-border transfer. The scope of sensitive data depends on a list set by the government, so it is safest to confirm with experts, including how health-related information is handled.
A 90-Day Plan for Near-Miss AI Analysis Implementation
Proceed in the following 3 phases, with 90 days as a guideline.
Days 0-30: Inventory and building the ground truth
- Collect the past 1 year of near-miss reports and count them by format, such as paper, LINE and Excel
- Decide the classification scheme for accident type, agent of injury, location and severity, and create the mapping table to shop-floor language
- Select 300 past reports and have 2 safety officers classify them to build the ground truth. Record the disagreement rate as well
- Put the no-blame agreement for reporters in writing in the local language
Days 31-60: Trial on 1 line and matching
- Trial the reporting channel (reporting from smartphones) on 1 line or in 1 area
- Match AI classifications against human classifications, and count agreement by category and missed serious events
- Check for translation mix-ups (negation, quantities, locations)
- Measure the number of reports and the time reporters spend on input
Days 61-90: Measure whether countermeasures close, and decide on full deployment
- For hotspots in the trial area, assign owners and deadlines, and measure the on-time completion rate of countermeasures
- Measure the time from report to owner assignment
- Decide the scope of full deployment, the RFP (the 12 items above) and the FAT/SAT pass criteria, and judge whether to place an order
Even if classification agreement is sufficient at day 60, if not a single countermeasure has been closed by day 90, reviewing the mechanism behind classification comes before full deployment. Conversely, even if agreement is only moderate, if the number of reports is rising and countermeasures have started getting closed, AI classification is worth continuing to use on the premise of human checking.
Frequently Asked Questions (FAQ)
How accurate can automatic near-miss report classification with AI be?
In research, when Japanese occupational accident reports (2,619 fall accidents in 2 industries) were classified with LLMs, newer-generation models generally exceeded a kappa coefficient of 0.7 in many categories other than agent of injury, and accuracy for indoor versus outdoor exceeded 91%. For agent of injury, the hardest category, the kappa coefficient was 0.662. However, this covered reports of occupational accidents that actually happened, not near-misses. Results change if the industries, accident types or classification scheme differ. Please measure agreement during acceptance testing, using your own reports and your own classification scheme.
When AI handles near-miss data analysis, is the number of reports a metric to reduce?
No, it is not a metric to reduce. If the channel is made easy and the no-blame agreement for reporters gets across, it is natural for the number to rise, and an increase can be a sign that hazards have become more visible. What to track is the time from report to owner assignment, the on-time completion rate of countermeasures, and the rate at which reports of the same type at the same location come up again after a countermeasure.
Can AI also analyze multilingual near-miss reporting, including Thai?
Using generative AI translation and summarization, reports in Thai, Vietnamese, English and other languages can be brought into a form that Japanese managers can read. However, if negation, quantities or locations get mixed up in translation, even the direction of the countermeasure changes. Test mix-ups specifically in acceptance testing, and design the register to show the original and translation side by side.
What are typical implementation costs and payback periods for near-miss AI analysis?
In this article’s model calculation (all placeholder values), against an initial cost of THB 600,000 and an annual cost of THB 240,000, labor saving in counting alone yields only THB 123,900 per year in benefit and does not pay back. If countermeasures taken from AI analysis prevent 1 lost-time accident per year, the calculation shows payback in about 2.1 years; if 1 every 2 years, about 7.2 years. The number of accidents prevented differs completely from factory to factory, and simply introducing AI does not reduce accidents. Please recalculate with your own actual figures and quotes.
Can AI counting substitute for Thailand’s statutory accident reporting?
No, it cannot. Section 34 of Thailand’s Occupational Safety, Health and Environment Act B.E.2554 requires employers, among other things, to notify immediately and report in writing within 7 days for fatal and serious accidents. The AI register can be used to gather material for statutory reports sooner, but it cannot replace the reports themselves. The same applies to holding safety committee meetings and the duties of the จป. Please confirm the latest practice and how it applies to your company individually with the competent authority or experts.
What should an RFP and acceptance tests (FAT/SAT) for near-miss AI check?
In the RFP, write 12 items including the reporting channel and languages, the classification scheme and the role of AI, handling of low-confidence cases, criteria for missed serious events, quality criteria for translation and summarization, countermeasure progress management, PDPA and audit trails. In acceptance testing, use as the ground truth the classifications that 2 safety officers assigned to 300 past reports, and check agreement by category, the number of missed serious events, whether the system can answer “I don’t know,” mix-ups of meaning in translation, the time to notification and owner assignment, handling of personal data, and the audit trail.
Summary
- Whether near-miss AI analysis delivers results is decided not by classification accuracy but by whether reporting becomes easier so that the number of reports rises, and whether the time until countermeasures are assigned and closed gets shorter
- AI handles translation, summarization, suggested classification, merging duplicates and estimating the location. People decide the final severity judgment, countermeasures, owners and deadlines, and completion checks
- The channel should be “30 seconds from a smartphone, in your own language, without being blamed.” Put the no-blame agreement in writing in the local language
- Safety officers create the ground truth, and the agreement between people serves as a guide to the level to require from AI
- In the model calculation (placeholder values), labor saving in counting alone does not pay back. What decides payback is whether countermeasures taken from AI analysis really prevented accidents
- In acceptance testing, give priority to missed serious events and translation mix-ups over overall agreement
- Thailand’s statutory accident reporting, safety committee and PDPA obligations do not change when AI is introduced. Confirm how the law applies with the competent authority or experts
You are welcome to start from taking stock of past reports, building the classification scheme, or creating ground-truth data together with your safety officers. If you are considering near-miss report AI analysis at a factory in Thailand, we can help you sort things out from the very first stage of your evaluation, so please feel free to reach out via our contact page.
References
- Heinrich’s law (MHLW Workplace Safety Site): https://anzeninfo.mhlw.go.jp/yougo/yougo24_1.html
- Hiyari-hatto / near-miss (MHLW Workplace Safety Site): https://anzeninfo.mhlw.go.jp/yougo/yougo26_1.html
- Examining the foundation (Safety+Health magazine, October 1, 2011): https://www.safetyandhealthmagazine.com/articles/6368-examining-the-foundation
- Joint development of a near-miss automatic extraction and report generation system by NEC and Sumitomo Heavy Industries (NEC, March 25, 2026): https://jpn.nec.com/press/202603/20260325_02.html
- Introduction of the near-miss automatic detection report creation function HAW (tentative name) (Norimono News, June 25, 2026): https://trafficnews.jp/post/681242/2
- Ando et al., automatic classification of occupational accident reports with LLMs (medRxiv abstract, October 2025): https://api.biorxiv.org/details/medrxiv/10.1101/2025.10.02.25337141
- Uncovering patterns in ferry near-misses (Ocean Engineering, University of Strathclyde): https://strathprints.strath.ac.uk/96699/1/Black-etal-OE-2026-Uncovering-patterns-in-ferry-near-misses-a-topic.pdf
- ABEX-RAT (arXiv, September 2025): https://arxiv.org/abs/2509.02072
- Near-miss analysis (Nomura Research Institute): https://www.nri.com/jp/service/solution/incident_analysis_ai.html
- Concern Reporting Software (Benchmark Gensuite): https://benchmarkgensuite.com/app/concern-reporting-software
- Occupational accident statistics for Reiwa 7 (MHLW, published May 27, 2026): https://www.mhlw.go.jp/stf/newpage_73382.html
- Occupational accident statistics for Reiwa 6 (MHLW, published May 30, 2025): https://www.mhlw.go.jp/stf/newpage_58198.html
- Thailand Occupational Safety, Health and Environment Act B.E.2554, unofficial English translation (Open Development Mekong): https://data.opendevelopmentmekong.net/dataset/a3d165ff-700e-4c32-a4eb-dd6b393e866b/resource/b171c6cf-58a1-46d7-a30e-fb4734dd2d24/download/occupational-safety-health-and-environment-act-b.e.-2554-a.d.-2011.pdf
- Occupational health and safety regulations in Thailand (Gratanet, November 14, 2025): https://gratanet.com/publications/occupational-health-and-safety-regulations-in-thailand
- Ministerial regulation on safety officers and committees B.E.2565 (Royal Gazette PDF, reposted by SHECU, Chulalongkorn University): https://www.shecu.chula.ac.th/data/boards/748/กฎกระทรวง%20การจัดให้มีเจ้าหน้าที่ความปลอดภัยในการทำงาน%202565.PDF
- Thailand Personal Data Protection Act B.E.2562, unofficial English translation (Mahasarakham University): https://pdpa.msu.ac.th/wp-content/uploads/2022/05/EN-PDPA-2019.pdf
- Thailand Personal Data Protection Act (Securiti): https://securiti.ai/thailand-personal-data-protection-act-pdpa
- Vietnam Personal Data Protection Law, Law No. 91/2025/QH15 (LuatVietnam): https://english.luatvietnam.vn/dan-su/law-on-personal-data-protection-law-no-91-2025-qh15-405135-d1.html
- Commentary on ISO 45001 clause 10.2 (Presencis): https://cdn.presencis.com/regulations/iso-45001/article-10.2/
- Commentary on the ISO 45001 definition of incident (Presencis): https://cdn.presencis.com/regulations/iso-45001/definitions/incident/
- Commentary on ISO 45001 clause 5.4 (ISO 9001 Checklist): https://www.iso-9001-checklist.co.uk/iso-45001/5.4-consultation-and-participation-of-workers.htm