Solution 61

AI-OCR (Forms & Invoices)

Invoices, delivery notes, purchase orders, acceptance certificates, receipts, and handwritten slips filled in on the shop floor. An almost countless number of forms and documents change hands every day in manufacturing and logistics. Someone has to work through those sheets of paper and PDFs line by line and key the information into the accounting system or the core business system — a scene that is still part of daily life at many Japanese-affiliated companies. Yet this task of “reading and typing” is far more demanding than it looks. It is error-prone, and it steadily consumes people who could be doing something else.

AI-OCR is the technology that is now bringing a major turning point to this long-standing challenge of digitizing forms. Conventional OCR (optical character recognition) stopped at “reading characters.” The latest AI-OCR, which combines generative AI and large language models (LLM), goes all the way through in a single flow: it reads the characters, understands their meaning, assigns each value to the correct field, and passes the result to your systems. TOMAS TECH provides this kind of form-digitization automation — invoice transcription being the most common starting point — to Japanese-affiliated manufacturers and logistics companies operating in Thailand.

This article explains, as concretely as possible and from a shop-floor perspective, what AI-OCR is, how it fundamentally differs from conventional OCR, which document types it covers, how the process works from reading through verification to data output, how it handles multiple languages and formats, how it connects to accounting, ERP and RPA, and what to expect in terms of implementation process and security. Our aim is to give anyone thinking “surely AI can do something about this form processing” a single article that supports a real decision.

What AI-OCR is — the shift from “reading” to “understanding”

AI-OCR is the general term for technology that combines conventional OCR with artificial intelligence (AI) — in particular machine learning, generative AI and large language models (LLM) — to read the characters written on a form or document and then structure that content as meaningful data. Its defining feature is that it does more than convert print into text: it also assigns meaning, recognizing “this is the invoice amount,” “this is the payment due date,” “this is the supplier name,” and shapes the result into a form your systems can import directly.

The limits of conventional OCR

Conventional OCR recognizes the shapes of characters within an image or scanned file and converts them into text. That is genuinely useful in its own right, and for standardized forms whose positions and formats are fixed in advance — an internal document that always arrives in exactly the same layout, for example — it can deliver high accuracy. In real-world form processing, however, there were several walls that conventional OCR alone could not get past.

  • Weak when the layout is not fixed: Every supplier uses a different invoice layout. The position of the amount, the wording of the field labels, the order of the line items — all of it varies. Conventional OCR tends to be designed on the premise of “read the field at these coordinates,” so the settings had to be rebuilt every time a layout changed.
  • It can read characters but not meaning: Even if it reads the figure “1,200,000,” conventional OCR cannot judge whether that is the amount before or after tax, or whether it is a subtotal or a grand total. In the end, a person still had to check it and file it in the right place.
  • Weak with handwriting and degraded text: With slips scribbled on the shop floor, or older documents where the print has faded, recognition accuracy tended to drop sharply.
  • Hard to handle wording variations: “Amount billed,” “invoice amount,” “total amount” — the same meaning can be expressed in countless ways. Hard-coded rules inevitably miss some of them.

Field extraction with generative AI and LLM

Adding generative AI and large language models to the picture changed the situation dramatically. An LLM has learned from an enormous volume of text and is good at inferring meaning from context. Just as a person looking at an invoice naturally judges that “this figure is at the bottom right, so it must be the total” or “it says ‘Grand Total,’ so this is the overall amount,” AI uses the overall layout of the form and the surrounding text as clues to infer which piece of information belongs to which field.

As a result, without specifying coordinates in advance to say “read here,” the system can extract fields such as supplier name, invoice number, issue date, payment due date, subtotal, tax amount, total and line items from an invoice layout it has never seen before — and do so with an understanding of what those values mean. This is the heart of AI-OCR, and the single biggest reason that automatic processing of non-standardized forms has become a practical reality. The invoice transcription AI system provided by TOMAS TECH is likewise designed around this flow of “read, understand the meaning, turn it into data.”

The real-world challenges of form processing

Why do so many companies struggle with digitizing their forms? To understand the value of AI-OCR, it helps to first lay out what is actually happening on the ground.

Enormous hours lost to manual entry and re-keying

Every month, or every day, staff in accounting, purchasing and order management open the forms that have arrived one by one and type the supplier name, date, amount, item, quantity and other details into accounting software or the core business system. Even at a few minutes per document, once that adds up to several hundred or several thousand documents a month it becomes a workload that cannot be ignored. During the month-end and month-start closing period in particular, all of this data entry hits at once and places a heavy burden on the people responsible.

Transcription errors and the double burden of checking

Manual entry always comes with mistakes. A digit typed wrong, a date mixed up, an incorrect supplier code. Errors involving amounts can lead to serious consequences: late payments, overpayments, distorted closing figures. That is why many organizations have a second person cross-check what was entered — a double check. It does reduce errors, but it also consumes more people. Both the person entering the data and the person checking it are spending their time on it.

Non-standardized forms and dependence on individuals

The more suppliers you have, the more varied the form layouts you have to handle. “On Company A’s invoice the total is written here”; “Company B puts the line items in two columns” — the person who remembers all these quirks is usually a long-serving, experienced member of staff. The moment that person takes leave or transfers to another role, processing grinds to a halt. Form processing is a textbook example of work that becomes dependent on specific individuals.

The complexity that multiple languages bring

For Japanese-affiliated companies operating in Thailand, this challenge becomes one level more complex. Communication with the head office in Japan is in Japanese, dealings with suppliers inside Thailand are in Thai, and exchanges with overseas suppliers are in English. The forms that arrive are a mixture of Japanese, Thai and English, and the currencies mix baht and yen. It is not unusual for multilingual handling itself to become a burden — splitting responsibilities by language, or translating documents while processing them.

Documents that AI-OCR can handle

AI-OCR is not limited to invoices. All kinds of paper, PDF and image documents generated in daily operations are candidates. Here are the most typical ones.

  • Invoices: Supplier name, invoice number, issue date, payment due date, line items, subtotal, tax amount, total and more. This is the document type in highest demand.
  • Delivery notes: Item, quantity, unit price, delivery date and so on. Used for matching against orders and goods acceptance.
  • Purchase orders: Ordered items, quantities, requested delivery dates, amounts. Useful on both the receiving and the ordering side of the process.
  • Acceptance and receipt certificates: Accepted items, accepted quantities, acceptance dates. These form one leg of three-way matching (order, delivery, invoice).
  • Receipts: Store name, date, amount, consumption tax and so on. Directly linked to automating expense reimbursement.
  • Handwritten slips and daily work reports: Shipping slips, inbound and outbound records, work reports and similar documents filled in on site. Traditionally the hardest area of all to digitize.
  • Application forms and ledgers of all kinds: Internal standard forms, inventory sheets, inspection checklists — any document with a fixed format.

Which documents to target depends on where the heaviest load sits in your particular operation. TOMAS TECH recommends starting digitization with “the most painful form,” confirming the results, and then widening the scope from there.

How AI-OCR works — from reading to data output

Let us walk stage by stage through how AI-OCR turns a form into data. Having the overall picture in mind also makes the later discussion of accuracy and verification flows easier to follow.

Step 1: Capturing the form and pre-processing

First, the form is brought into the system as an image or PDF. The capture method can be chosen to suit your operation: scanning on a multifunction printer, photographing with a smartphone or camera, or automatically retrieving PDFs attached to emails. The captured data then goes through pre-processing — skew correction, brightness and contrast adjustment, noise removal — to put it in the best condition for the recognition stage that follows. Thanks to this pre-processing, recognition accuracy is easier to maintain even for forms photographed at an angle or in slightly dim light.

Step 2: Character recognition (reading)

Next, the system detects where text appears on the form, reads it, and converts it into text data. Up to this point the process is shared with conventional OCR, but an AI-based recognition engine copes more readily not only with printed characters but also with handwriting and with strings that mix several languages. The recognized text is retained together with its positional information — roughly where on the form it appeared — which becomes a clue for the field extraction that follows.

Step 3: Field extraction (assigning meaning)

This is the decisive difference from conventional OCR, and the stage where generative AI and LLM show their strength. Looking across all of the recognized text, the system judges from layout and context which value is the supplier name, which is the invoiced amount, which is the payment due date, and extracts them accordingly. Treating various expressions such as “合計” (total in Japanese), “Total” and “รวม” (total in Thai) as carrying the same meaning, or breaking each line item down into item, quantity, unit price and amount — the judgments that people used to make in their heads are taken over by AI.

Step 4: Assigning confidence scores and verification

AI can attach a confidence score to each extracted field, indicating how certain it is. A clearly printed total amount, for example, will carry a high confidence score, while a faint handwritten figure will carry a low one. Using these scores, you can build an efficient verification flow: accept high-confidence fields as they are, and have a person check only the low-confidence ones. In parallel, the system runs arithmetic consistency checks such as “does subtotal plus tax equal the total?” and cross-references against your existing supplier master, automatically detecting obvious anomalies.

Step 5: Human review (human in the loop)

AI-OCR is not about replacing people entirely; it is about dramatically reducing their workload. Staff check and correct on screen only the low-confidence fields and the documents flagged by the consistency checks. Because the original form image and the extracted results can be displayed side by side, even this review work is far faster than the traditional read-and-type approach. The corrections people make can also be fed back into improving the system. This division of roles — AI handles almost everything, people look only at the exceptions — is the key to achieving accuracy and efficiency at the same time.

Step 6: Data output and system integration

Finally, the confirmed data is passed to your accounting system or core business system via CSV, Excel or an API. This completes the flow of “a form arrives, AI reads it, a person checks only the exceptions, and the system registers it automatically” — and the manual entry step all but disappears from the operation.

Support for multiple languages and formats

For Japanese-affiliated companies doing business in Thailand, multilingual and multi-format support is where the true value of AI-OCR shows most clearly. It is also a point that TOMAS TECH, as an IT integrator with deep knowledge of operations in Thailand, treats as a priority.

Handling a mix of Japanese, Thai and English

On the ground in Thailand, it is not unusual for several languages to appear on a single form. Documents where the company name is in English, the item description is in Thai and the remarks are in Japanese circulate as a matter of course. AI-OCR built on generative AI and LLM can extract fields from such multilingual documents without having to be told where one language stops and another begins. It can go further, too — for example translating extracted Thai or English fields into Japanese and showing both side by side — which smooths the path all the way through to preparing reports for head office.

Absorbing layout differences between suppliers

As noted earlier, invoice and delivery note layouts differ from supplier to supplier. The position of field labels, the order of line items, whether or not a table is used — with conventional OCR you had to build settings for each layout, but AI-OCR, which extracts by understanding meaning, absorbs those layout differences automatically to a considerable degree. Even when a new supplier is added, processing can usually begin without any special configuration, which greatly reduces the operational effort. Of course, for unusual layouts or proprietary forms, we can receive samples in advance and optimize the setup to achieve more stable accuracy.

Attention to differences in currency, date and number notation

In cross-border trade there are many small differences in notation: currency symbols (฿, ¥, USD), date formats (day/month/year versus year/month/day), and decimal and thousands separators (comma versus period). Misreading these differences is a direct cause of incorrect amounts and dates. AI-OCR infers these notation rules from the context of the form and can convert everything into a unified format, making it far easier to manage data centrally across regions and suppliers.

Integration with accounting, ERP and RPA

To get the most out of AI-OCR, it is essential to connect the extracted data reliably to the business systems downstream. If the process stops at digitization, you risk the self-defeating outcome of someone re-keying that data by hand further along the line. Drawing on our experience with core systems for factories and logistics, TOMAS TECH includes this “connecting up” work in what we propose.

Integration with accounting systems

Data extracted from invoices and receipts is passed to accounting software as journal entries or payment data. By linking it to the supplier master and the chart of accounts, accounting staff can concentrate on checking and approving rather than typing. This shortens the time required for the monthly close and contributes to faster financial reporting.

Integration with ERP and core business systems

Feeding purchase order, delivery note and acceptance certificate data into ERP or purchasing and inventory management systems connects the whole chain — ordering, receiving, acceptance and payment — with data. Three-way matching of order, delivery and invoice is particularly important for catching discrepancies in amounts and quantities early, and digitizing each of those documents with AI-OCR allows that matching to be semi-automated. This is a high-impact integration that goes straight to the core operations of manufacturing and logistics.

Combining with RPA

For existing systems where API integration is difficult, another option is to combine AI-OCR with RPA (robotic process automation) so that the data produced by AI-OCR is entered automatically through screen operations. In many cases this “AI-OCR reads, RPA types” arrangement makes automation achievable while keeping modifications to existing systems to a minimum.

CSV, API and a range of output formats

As for the specific means of integration, you can choose flexibly according to your environment: widely compatible CSV or Excel output, handover via an API suited to real-time integration, or placing files in an existing shared folder. Being able to adopt AI-OCR without major changes to existing systems is an important factor in lowering the practical barrier to implementation.

The benefits of implementation

What changes on the ground when AI-OCR is introduced? Here we summarize the most typical benefits. Note that the scale of the effect varies with the type and volume of forms and with how you operate today, so we calculate concrete figures based on each customer’s individual situation.

A large reduction in data entry hours

The most direct benefit is cutting the time previously spent on manual entry. Because AI reads most of the fields automatically and people only review the exceptions, the processing time per document falls substantially. In general, the greater the volume of forms and the more routine the processing, the larger the reduction tends to be. It also helps ease the overtime that used to concentrate around the closing period.

Stable quality and fewer errors

When AI takes on data entry work that used to depend on human concentration, the variability caused by fatigue and inattention decreases and work quality stabilizes. Because automatic checks based on confidence scores and arithmetic consistency verification are built in, obvious anomalies are detected at an early stage. The reliability of information such as amounts and dates — where mistakes have the biggest impact — improves accordingly.

Faster processing

Shortening the lead time from a form arriving to it becoming data also speeds up the downstream steps, such as payment processing and inventory updates. Having information reflected in systems in a timely manner contributes to the speed of management decision-making as well.

Breaking free from dependence on individuals

Escaping the situation where “only that person can process this” is another benefit not to be overlooked. Because AI absorbs the quirks of each layout, form processing can be run at a consistent quality by anyone, without relying on one particular member of staff. That makes for a stable operating structure that holds up against staff transfers and time off.

Redirecting people to higher-value work

Staff freed from simple data entry can devote their time to higher value-added work such as analysis, improvement activities and communication with suppliers. At a time when labor shortages are a real problem, being able to apply limited human resources to judgment rather than typing is of enormous value to a company.

The implementation process — from sample collection to production and improvement

The key to a successful AI-OCR implementation is to proceed steadily in stages rather than rolling it out company-wide from the outset. TOMAS TECH listens carefully to your operational challenges first, and then supports the implementation along the following path.

Step 1: Interviews and sample collection

First we ask in detail about your current operation: which forms consume how many hours, and what layouts are involved. At the same time we receive samples of the invoices, delivery notes and other documents you actually process. Understanding how much variation exists in your real-world forms is the starting point for an accurate design.

Step 2: Design, learning and configuration

Based on the samples we receive, we design the fields to extract, the confidence thresholds, the verification rules, the output format and the systems to integrate with. For unusual layouts and proprietary forms, we tune and optimize the setup so that AI can read them reliably. At this stage we also firm up the operating rules — which fields are entrusted to AI, and from where a person takes over the checking.

Step 3: Validation (PoC / trial)

Before going into production, we run a trial with your actual forms to verify accuracy, processing speed and whether the operational flow fits your business. Any issues that surface here — a particular layout that is hard to read, a field that needs checking too often — are identified and the configuration adjusted. Because you confirm the results with real data before moving to full deployment, you can make the decision with confidence.

Step 4: Production rollout and user training

Once the validation gives you confidence, we deploy to the production environment and provide operating instructions to the people who will use the system. We can deliver these explanations in both Japanese and Thai, so local staff can start using the system without difficulty. We provide careful support so that the new workflow takes root on the ground.

Step 5: Operation and continuous improvement

Implementation is not the end. We keep raising accuracy and efficiency while the system is in use. By reviewing the patterns in the data people check and correct and revising the configuration, and by adding support for newly introduced document types, continuous improvement makes AI-OCR fit your operation better the more you use it. You can also expand the range of target documents gradually, growing the scope of automation over time.

How to think about accuracy — setting realiztic expectations

Accuracy is what most people want to know about when considering AI-OCR. This deserves an honest answer, so let us explain it in a little more detail.

The reading accuracy of AI-OCR varies considerably with conditions: the condition of the form, its layout, the language, and whether the text is handwritten or printed. For a cleanly printed, standardized invoice you can expect very high accuracy, but with a faded handwritten slip or an extremely irregular layout, accuracy falls correspondingly. Claiming a figure such as “always 100%” is therefore not realiztic. What matters is not raw reading accuracy in itself, but the mechanizm by which AI honestly signals, through confidence scores, the places where it is unsure, and a person checks those places — which is what guarantees the accuracy of the final data.

In other words, the essential value of AI-OCR lies in reducing manual effort while preserving the accuracy you need. Rather than aiming for full automation, we design the optimal division of roles between AI and people so that reduced workload and assured quality go together. That is TOMAS TECH’s approach. The most reliable way to find out what accuracy and efficiency you can actually expect is to confirm it concretely through a validation exercise using your own forms.

Security, data storage and handling of personal information

Invoices and other business forms contain highly confidential information: supplier names, amounts, bank account details, the names of individuals. When AI-OCR is used in real operations, security and data handling can never be treated lightly.

Data handling and storage

Where and how form data and extracted results are stored is designed in line with your security policy. We build the configuration on the premise of preventing information from leaking outside — encrypted communication, access rights management, careful selection of storage locations. We also believe it is advisable to agree in advance on retention periods and deletion rules for processed data.

Protecting personal and confidential information

For personal and confidential information contained in forms, we keep the scope of handling to the minimum necessary and control permissions so that no one outside the relevant parties can access it. Operational design must take account of the laws and guidelines of the region where you do business, including Thailand’s Personal Data Protection Act (PDPA). TOMAS TECH works through these points together with you to build a setup you can use with confidence.

Auditing and traceability

When, by whom, which form was processed, and which fields were corrected — keeping a record (log) of these actions means the processing history can be traced later. In work that involves monetary amounts, this traceability is also important from the standpoint of internal control and audit response.

Frequently asked questions (FAQ)

Q1. Can it handle invoices whose layouts differ from supplier to supplier?

Yes, it can. Unlike conventional OCR, AI-OCR based on generative AI and LLM does not hard-code the layout by coordinates; it extracts fields by understanding their meaning. That is why a single mechanizm can handle non-standardized forms in a wide variety of layouts. For unusual layouts, we can receive samples in advance and optimize the setup to achieve more stable accuracy.

Q2. Can it read handwritten slips?

Handwriting is supported. Compared with printed text, however, accuracy varies with the writer’s habits and how irregular the characters are. For handwritten forms we therefore recommend combining the process with a confidence-based human review flow to guarantee the accuracy of the final result. A validation exercise using your actual slips will show how much of the work can be automated.

Q3. Is it all right if a form mixes Japanese, Thai and English?

That is no problem. Even when several languages appear on a single form, AI-OCR can extract the fields without needing to be told where one language changes to another. It can also translate the extracted content into Japanese and show both versions together, producing data that is easy to use both at the site in Thailand and at the head office in Japan.

Q4. Can it connect to our current accounting or core business system?

In most cases, yes. You can choose the method that fits your existing systems: CSV or Excel output, API integration, or placing files in a shared folder. For systems where API integration is difficult, there is also the option of combining with RPA to automate screen entry. We ask about your current environment and then propose the most suitable integration method.

Q5. Will accuracy reach 100%?

To be honest, we cannot guarantee 100% at all times for every kind of form. Accuracy varies with the condition and layout of the document. The value of AI-OCR lies less in raw reading accuracy than in the mechanizm whereby AI indicates the places it is unsure about and a person checks only those places — protecting the accuracy of the final data while cutting workload substantially. Please think of it as delivering results through the optimal division of roles between AI and people, rather than through full automation.

Q6. How long does implementation take?

It depends on the types and volume of forms to be covered and on the state of the systems to integrate with. The usual approach is to start small with a trial covering a subset of forms, confirm the results, and then move on to production and expansion. Proceeding in stages lets you contain risk while steadily widening automation. We present a concrete schedule after the initial interviews.

Q7. Which form should we start with?

We recommend starting with the form that consumes the most hours, or the one where mistakes have the biggest impact. In many cases, high-volume invoices or receipts are the first target. Beginning digitization with the most painful task and expanding the scope once people can feel the benefit lets you advance automation while building internal buy-in.

Q8. We are concerned about security. Will confidential information be protected?

Because forms contain highly confidential information, we treat security as a fundamental premise of the design. We propose an operating model that reflects your policies and the applicable laws, including Thailand’s PDPA — encrypted communication, access rights management, agreed storage locations and data retention rules, and traceability through operation logs. We will build a setup you can use with confidence together with you, so please do not hesitate to share your concerns.

Conclusion — freeing your team from “read it, then type it”

Digitizing invoices, delivery notes and other business forms is work that has long depended on human hands in many organizations. The hours spent on manual entry, transcription errors, the burden of double checking, dependence on individuals caused by non-standardized forms, and the complexity of multilingual handling — every one of these problems stems from the fact that a person reads and a person types.

AI-OCR combined with generative AI and LLM changes that structure at its root. It does not merely read characters; it understands meaning, extracts the required fields from non-standardized, handwritten and multilingual documents, keeps human checking to a minimum on the basis of confidence scores, and passes the data on to accounting, ERP and RPA. As a result, removing the “read it, then type it” step from your operations altogether has become a realiztic option.

What matters is not holding up full automation as an ideal, but building a realiztic mechanizm in which AI and people divide the roles optimally so that reduced workload and accuracy go hand in hand. As an IT integrator that knows the manufacturing and logistics workplaces of Thailand, TOMAS TECH supports AI-OCR implementation tailored to your forms and your operations — end to end, from initial interviews through validation and production to continuous improvement.

If you have ever thought “surely AI can do something about this form processing,” please feel free to get in touch. We will listen to your current forms and pain points and propose an approach likely to deliver results for your company. You can reach us through our contact form.