A meeting is called to settle “cloud or on-premise,” no conclusion is reached, and the item is carried over to the next meeting. We have watched this scene repeat at Japanese-owned factories in Thailand. The reason nothing gets decided is not a lack of information. It is that the question of moving to a cloud production management system is being answered once, for the system as a whole. The unit of decision is the function, and the deciding axis is not cost but how many minutes that function is allowed to stay down. And after the migration, the first thing to stop is neither the server nor the application. It is authentication.
The premise for moving production management to the cloud changed in 2026
Until now, whenever a site in Thailand looked at cloud, the discussion stalled at the same place: where will the data physically sit? Singapore or Tokyo — either way, outside the country. The IT department at the Japanese head office refers the question of cross-border transfer to legal, legal asks for material on which to base a decision, and the project sits still for several months. That has been the standard picture for the last few years.
In 2026, that premise changed, because keeping data inside Thailand became a standard option for the first time.
What it means that Bangkok now has cloud regions
Let us set out the facts. AWS (Amazon Web Services, Amazon.com, Inc.) made the Asia Pacific (Thailand) Region (ap-southeast-7, 3 availability zones) generally available on 8 January 2025. Google Cloud (Google LLC) opened its Bangkok region (3 availability zones) on 21 January 2026 and has announced investment on the scale of USD 1 billion. The Thailand region of Microsoft Azure (Microsoft Corporation), Thailand South, is at present still at the stage of an announced intention to build, and in October 2025 it was disclosed that True IDC (True Corporation) will be used as one of the availability zones.
The difference between these 3 situations translates directly into a difference in how you can run the evaluation. A platform that is already generally available and a platform whose opening has been announced are not the same thing when it comes to what you can write in an internal approval request. If your head office has made “data must stay inside Thailand” a condition, only the former can satisfy that condition today. If you build your plan on the latter, you are writing a migration plan on a premise whose opening date is not fixed, and that is worth flagging explicitly as a risk in the plan.
One more point matters in practice: there are 3 availability zones. In a single-site data centre with only 1 zone, a power or air-conditioning failure becomes a service outage directly. With 3 zones, you can spread a database across multiple zones and fail over automatically when one side goes down. This is a different order of thing from putting 2 servers in your own factory and calling it redundancy. Redundancy inside the factory cannot beat a building-wide power cut or a flood.
That said, this is not the point to jump to a conclusion. “A region exists in Thailand, therefore everything can go to the cloud” does not follow. What changed is only that the data-residency constraint has been lifted. None of the other constraints has moved.
Why “everything to the cloud” still is not the answer
What tends to happen when one constraint is lifted is that people assume the remaining constraints disappeared with it. In reality, 3 problems survive even when the region is inside the country.
The first is the line between the factory and everything beyond it. Even with the data centre in Bangkok, the path from the factory to it runs over a carrier’s line. Construction work, power cuts, equipment failures, and physical faults during the rainy season. That segment is your own responsibility no matter where the region sits. Even if the cloud side quotes a service availability of 99.9%, if the line that reaches it goes down, availability as seen from the factory floor is zero.
The second is that the cost structure becomes harder to read. In the JUAS “Corporate IT Trends Survey 2026”, the most frequently cited reason for IT budget increases was “renewal, modification and expansion of existing systems” at 66.3%, followed by “the impact of the weak yen, rising labour costs and vendor price increases” at 46.6%, “growth in cloud services” at 45.0%, and “increased AI-related investment and usage cost” at 43.7% (up 7.4 points from 36.3% the previous year). The fact that growth in cloud ranks high among the reasons budgets go up is itself a sign that the assumption “cloud will make it cheaper” is breaking down in the field.
The third is that the swing back after moving to the cloud is actually happening. A compilation of figures on cloud repatriation reports that in Flexera’s 2025 State of the Cloud, 21% of workloads or data have already been repatriated, and that 84% cite “managing cloud spend” as their biggest challenge. Research from IDC found that 59% of organisations exceeded their budget in 2024. However, IDC research as of October 2024 also found that only about 8% of organisations pull all workloads back; for most, repatriation is selective.
That word “selective” is the argument of this article in miniature. Even the companies that are moving things back are not moving everything back. They are moving back only the functions that did not fit. If that is so, then you may as well decide the location function by function from the start. The way the question “cloud or on-premise?” is framed — one answer for the whole system — is itself wrong.
A production management system is not a single function. Label printing, inventory allocation, shop-floor data entry, master maintenance, purchasing, cost accumulation. Their required response times differ, the damage caused when they stop differs, and the nature of the data they handle differs. It is because you are trying to give one answer to things that are not alike that nothing gets decided.
The difference between cloud and on-premise production management is measured in allowable downtime, not cost
So if the decision is made function by function, what is the criterion for sorting them? Not cost. Allowable downtime, that is, RTO (Recovery Time Objective).
The reasoning is shown by the cost model later in this article, but here is the conclusion up front. At a factory of this size, if you make cost the deciding axis, cloud wins almost every time. And the point at which that reverses is more than 10 years out. Because the gap does not fall inside the usual investment-appraisal horizon of 5 to 7 years, cost does not function as a deciding axis. The meeting never ends because it is being held around an axis that does not work.

Allowable downtime by function
Allowable downtime is not something the IT department settles on paper. It is set by the people on the floor who are actually in trouble when the function stops. The table below shows typical levels from an exercise we actually ran at a Japanese-owned factory in Thailand. Your own figures should be filled in again from scratch, together with the people on the floor.
| Function | What happens on the floor when it stops | Downtime that can be tolerated |
|---|---|---|
| Label and part-tag printing | Finished goods cannot move to the next process or to shipping | 5-15 minutes |
| Inventory allocation and issue instructions | Picking stops | 15-30 minutes |
| Shop-floor data entry and process progress | Records pile up, and accuracy falls with after-the-fact entry | 2-4 hours |
| Master maintenance (part numbers, BOM) | New part numbers cannot be registered | 1 business day |
| Purchasing and ordering | Today’s orders slip to tomorrow | 1 business day |
| Cost accumulation and analysis | The monthly close is delayed | 3 business days |
Looking at this table, you can see that inside one and the same “production management system” there are functions that can only be down for 5-15 minutes living alongside functions where 3 business days of downtime still leaves the business running. The gap is a factor of several hundred.
Why is label printing so short? Because if a label cannot be applied to a finished item, that item physically cannot move to the next process. It has to be stacked in a holding area, and once the holding area fills up, production itself stops. On top of that, relabelling later carries the labour of identifying which lot each item belonged to. 15 minutes of downtime can turn into several hours of rework the following day.
Cost accumulation is the opposite. If it cannot be run today, it can be processed in bulk tomorrow. Except immediately before the monthly close, 3 business days of downtime does not stop the business. Spending the same amount on availability for this function as for label printing is excessive.
Alongside RTO, RPO (Recovery Point Objective — how much data loss is acceptable) should be set function by function as well. Shop-floor data entry has a short RPO: if the last 30 minutes of entries disappear, nobody can reconstruct them. Master maintenance is different — if the last update disappears, the person who made it remembers and can enter it again. How to think about these 2 indicators, and how to turn them into an actual design, is covered in our article organising business-system backup around RTO and RPO, which is worth reading before you fill in the table internally.
One RTO for the whole system produces over-investment and under-provisioning at the same time
What we often see in the field is a requirements document with a single line: “system-wide RTO: 4 hours.” That one line triggers 2 failures simultaneously.
The first is over-investment. An RTO of 4 hours is far too strict for cost accumulation and purchasing. If you build a 4-hour recovery capability for a function that would be fine with 3 business days, the redundancy cost and the operational effort behind it are wasted in full. And this waste is invisible in money terms, because it is buried under a line item called “system-wide availability.”
The second is under-provisioning. An RTO of 4 hours is far too loose for label printing. If you design a function that can only be down for 5-15 minutes to a 4-hour target, then when a real failure occurs the floor has nothing to do for 3 hours and 45 minutes. And it is only when the failure happens that anyone realises “we cannot wait 4 hours for this.” At that point, the cost of redoing the design is several times what it would have been during the initial build.
Putting a single RTO on the whole system means ignoring the requirement of the strictest function while over-investing in the loosest one — a decision that misses in both directions. What should be written is not one line, but as many lines as there are functions.
This work does not finish inside the IT department alone. The people who can answer “what happens if label printing is down for 15 minutes” are in manufacturing, and the person who can answer “how long can we hold out in that state” is the plant manager. It is worth having both in the room at the very first requirements session. What tends to happen to projects that proceed without them is set out in our article on the typical patterns behind failed production management system implementations.
In a cloud migration, the first thing to stop is not the server but authentication
Now to the core of it. Suppose you have built that table of RTOs by function. The table has a decisive gap: authentication is not in it.
Authentication is not recognised as a “function.” It does not appear on the list of business functions, it is not drawn on the business flow diagram, and therefore no RTO is assigned to it. In reality, though, it sits in front of all 6 functions in the table above. If you cannot log in, no labels come out and no shop-floor entries go in.
In other words, the RTO for authentication has to match the shortest value in the table. If label printing is 5-15 minutes, then authentication is 5-15 minutes too. Because this drops out of the design, the first trouble after a cloud migration occurs, almost without exception, in authentication.

The 3 authentication configurations and how they behave when the line drops
When a production management system moves to the cloud, authentication in practice falls into 3 configurations. Each behaves differently when the line connecting the factory to the outside world is cut.
Configuration 1: cloud IdP only. Authentication runs against a cloud identity provider such as Microsoft Entra ID (Microsoft Corporation) or Google Workspace (Google LLC), with no domain controller on the on-premise side. Operationally this is the lightest. But when the line drops, nobody inside the factory can log in. Sessions already logged in keep working until the access token expires, and stop the moment it does. If the lifetime is 1 hour, then within at most 1 hour of the line dropping, every terminal falls over one after another. The awkward part is that they do not all fall at once. Because the times differ, the floor perceives it not as “the line is down” but as “the system is unstable,” and isolating the cause is delayed.
Configuration 2: on-premise AD plus federation to the cloud. Active Directory (Microsoft Corporation) sits inside the factory and authentication is linked out to the cloud. The behaviour is the exact opposite depending on whether you use password hash synchronisation or federation (a method that forwards authentication requests to the on-premise side). With password hash synchronisation, authentication completes on the cloud side alone even if the on-premise AD is down. With federation, logins to cloud services stop the instant the on-premise authentication platform goes down. The two are almost indistinguishable from the look of an architecture diagram. When a vendor’s proposal says only “integrated with AD,” it is worth confirming which of the two it is.
Configuration 3: on-premise AD only. Cloud-side applications are used only from terminals joined to the on-premise AD domain. If the line drops, logins inside the factory still work, but of course nothing can reach the cloud applications. Authentication is alive while the application is dead — and from the floor’s point of view, that is indistinguishable from configuration 1.
On top of these 3, what multi-factor authentication (MFA) depends on also needs checking. Where a second factor is confirmed through an authenticator app or SMS, authentication will not pass if the mobile network is unwell, even when the factory line is alive. The availability of authentication is not determined by the internal network alone.
Label printing and shop-floor data entry stop at the login screen
Here is what that looks like concretely.
At 07:30 the shift starts. Handheld terminals are switched on and an operator holds up an ID card. Authentication does not go through. The cause is that a piece of equipment on the line side failed during the night. The server is running. The application is running. The data is intact. But nobody can log in.
The first thing the floor does in this situation is, in most cases, to write on paper. That in itself is the right response. The problem is the labour of entering those paper records afterwards, and the labels that could not be printed in the meantime. With no labels, finished goods go into the holding area. By late morning the holding area is full, and by the afternoon somebody has to decide whether to stop production.
And there is a second hill after recovery. When 3 hours’ worth of handwritten records are entered after the fact, the timestamps all tend to become the time of entry. The time data for process progress is off by 3 hours, and that day’s lead-time analysis is unusable. A day of missing data is less troublesome than a day of wrong data.
What is worth noticing here is that the cause of this incident is not the production management system. The package was not chosen wrongly, and the server was not underspecified. No RTO had been assigned to authentication — that is all. However much time you spend comparing package features, that failure will not be prevented. The axes on which to compare the packages themselves are collected in our comparison of production management systems, but the authentication design is worth progressing in parallel with the selection, not after it.
4 design decisions that keep authentication up
There are 4 countermeasures. None of them requires special technology. It is a matter of deciding what needs to be decided.
1. State an RTO for authentication and match it to the shortest function. Write one line in the requirements document: “RTO for authentication: 15 minutes.” That alone changes what the vendors propose. A vendor told to design authentication that recovers in 15 minutes can no longer put forward a design that leans on a single IdP. Conversely, without that line, authentication remains nobody’s area of responsibility right through to the end of the build.
2. Duplicate the authentication path. The aim is a state in which authentication still passes on one side when the other — cloud or on-premise — is down. The password hash synchronisation described above is one example. What matters is not “it is duplicated” but “we have taken one side down and confirmed that it still passes.” A failover test during a planned outage once a year is worth building into the operating calendar. A redundant configuration that has never been tested is not redundant.
3. Design token and cache lifetimes by function. Shop-floor terminals and office terminals do not need the same settings. A handheld terminal on the floor is used in a fixed place by a fixed set of operators. There, a longer token lifetime lets it keep working for a certain period even when the line is down. Office terminals, which can also be reached from outside the company, are better kept short. The trade-off between security and availability is struck at the level of the terminal and the function, not at the level of the system.
4. Provide an offline authentication fallback on shop-floor terminals. Only while the line is down, authentication passes using locally cached credentials or an IC card on the terminal, so that label printing and shop-floor data entry alone can continue. Alongside that, prepare an administrator account for emergency use only (a break-glass account) and an alternative for when MFA is unavailable, and aim for a state in which at least 2 people — the plant manager and the IT contact — know that it exists and where it is kept. An account known to only one person does not work on the day that person is away.
What matters is round trips, not bandwidth | how handheld terminals get along with the cloud
The next most common question after authentication is “won’t the floor get slower if we go to the cloud?” Here too, the thing being looked at is often the wrong thing.
The one-scan, 5-round-trip problem
First, some indicative latency figures. The numbers below are levels that are commonly observed. They are indicative only, not fixed values — measure from your own factory’s line. Even within the greater Bangkok area they vary considerably with the type of line contracted and the route it takes.
| Destination | Indicative RTT (round-trip latency) |
|---|---|
| Region inside Bangkok | 5-15ms |
| Bangkok to Singapore | 25-40ms |
| Bangkok to Tokyo | 70-90ms |
Looking at these numbers and thinking “70ms, no human will notice” is the first mistake. The issue is not one instance of latency, but how many times a single operation talks to the server.
When a handheld terminal scans the barcode on one part tag, several communications occur inside the system. Typically: validating the authentication token, looking up the part-number master, allocating inventory, registering the transaction, and requesting the label print. Depending on the implementation, that is 5 round trips.
If the system sits in a Tokyo region, and taking an indicative RTT of 80ms, the round trips alone add 5 x 80ms = 0.4 seconds. Server processing time is on top of that. With a region inside Bangkok, taking an indicative 10ms, it is 5 x 10ms = 0.05 seconds. With the same implementation, location alone produces a difference of 0.35 seconds.
Whether 0.4 seconds feels like “nothing much” is determined by the count. At a site that scans, say, 1,000 times a day, the round trips alone come to 0.4 seconds x 1,000 = 400 seconds, about 6.7 minutes. And this is waiting time, so for all of it the operator is standing there holding the handheld terminal.
The important part here is that adding bandwidth does not improve this number. Round-trip latency is determined by the speed of light and the length of the path, so going from 100Mbps to 1Gbps changes nothing. There are only 2 ways to improve it: shorten the distance (use a region inside the country), or reduce the number of round trips.
So the thing to check with a vendor is not “how many Mbps of line do we need?” It is “for one instance of each main shop-floor operation, how many round trips to the server occur?” Not many vendors can answer that on the spot. Where they cannot, the figure can be measured during a PoC (proof of concept); logging on the handheld side is enough to capture it.
Note that the number of round trips is determined by the design philosophy of the package, so reducing it afterwards is not easy. Modifications such as “pre-fetch the masters in bulk at scan time” or “make transaction registration asynchronous” are possible, but both are additional development. This is an item for the selection stage.
Where to put the offline buffer
Even after reducing round trips, if the line is cut the number of round trips becomes zero. What can be done while it is cut has to be designed in advance.
There are 3 options for where to put the buffer.
On the terminal. The handheld application holds a local database, writes to it while the line is down, and synchronises after recovery. No additional hardware is required, so it is the cheapest. The weaknesses are that losing or breaking a terminal loses that terminal’s unsynchronised data, and that inventory cannot be kept consistent across several terminals. It does not suit processing that needs exclusive control, such as inventory allocation.
On an edge server inside the factory. A small server is placed in the factory and takes charge of replicating the masters, printing labels, and receiving transactions temporarily. Terminals always talk to this edge server, and the edge server synchronises with the cloud. Because consistency across several terminals holds inside the factory even while the line is down, inventory allocation can continue as well. This is the core of the hybrid configuration described later. The weaknesses are that this single machine becomes a single point of failure, and that the synchronisation logic is additional development.
On the printer. Limited to label printing only, you can hold templates and recent print data on the label printer and, while the line is down, send print instructions to the printer directly from the terminal. It is the cheapest way to protect only the function with the shortest RTO. However, unless you decide in advance on a procedure for feeding information about printed labels back into the system afterwards, duplicate printing becomes impossible to trace.
Which one to choose comes back, again, to the RTO table. If label printing is 5-15 minutes, a configuration that could take 15 minutes or more to recover is not enough. If shop-floor data entry is 2-4 hours, a buffer on the terminal may well be sufficient. The answer differs by function, so here too the decision is framed as “what do we do per function,” not “what do we do as a system.”
Five-year cost comparison of cloud and on-premise production management systems [original model]
Cost does not work as a deciding axis, as stated earlier. To show that, here is the actual calculation. What follows is a simplified model based on the assumptions we use in projects at Japanese-owned factories in Thailand; it is not the set of figures from any single real company. It is meant to be used with your own quoted figures substituted in.
Note that this model does not calculate any savings, ROI or payback period whatsoever. It compares costs only. The assumptions behind savings vary too much from factory to factory, and presenting them in a model would be misleading.

Model factory assumptions
- Location: greater Bangkok, single site, with a Japanese head office
- Employees: 180
- Users of the production management system: 40 (25 office, 15 shop floor)
- Shop-floor devices: 15 handheld terminals plus 10 tablets
- Currency: THB (Thai baht)
- Comparison period: 5 years
The cost of shop-floor devices (750,000 THB) is included at the same amount in all 3 scenarios, because the devices the floor uses are needed whether you go cloud or on-premise. We do sometimes see quotations that include this on one side only, but that is not a comparison.
5-year totals for the 3 scenarios
Scenario A: full cloud (SaaS)
| Category | Item | Amount (THB) |
|---|---|---|
| Initial | Implementation support, requirements definition, master migration, training | 1,800,000 |
| Initial | Line redundancy, initial | 120,000 |
| Initial | Shop-floor devices (15 handhelds + 10 tablets) | 750,000 |
| Initial | Initial total | 2,670,000 |
| Annual | Licences (40 users x 1,800 x 12 months) | 864,000 |
| Annual | Line (12,000 x 12 months) | 144,000 |
| Annual | Additional development | 300,000 |
| Annual | Annual total | 1,308,000 |
Five-year running cost: 4,320,000 + 720,000 + 1,500,000 = 6,540,000 THB
Five-year total: 9,210,000 THB
Scenario B: full on-premise
| Category | Item | Amount (THB) |
|---|---|---|
| Initial | Perpetual licences | 2,800,000 |
| Initial | Implementation support | 1,800,000 |
| Initial | 2 servers + storage + UPS | 900,000 |
| Initial | OS and database licences | 450,000 |
| Initial | Shop-floor devices | 750,000 |
| Initial | Internal network | 80,000 |
| Initial | Initial total | 6,780,000 |
| Annual | Maintenance (licence maintenance 420,000 + infrastructure maintenance 180,000) | 600,000 |
| Annual | Additional development | 300,000 |
| Annual | Annual total | 900,000 |
Five-year running cost: 3,000,000 + 1,500,000 = 4,500,000 THB
Five-year total: 11,280,000 THB
Scenario C: hybrid (shop-floor functions on-premise, management functions in the cloud)
| Category | Item | Amount (THB) |
|---|---|---|
| Initial | Implementation support (including integration design) | 2,200,000 |
| Initial | 1 edge server + UPS | 380,000 |
| Initial | OS and database licences | 180,000 |
| Initial | On-premise licences for shop-floor functions | 1,200,000 |
| Initial | Shop-floor devices | 750,000 |
| Initial | Line redundancy, initial | 120,000 |
| Initial | Initial total | 4,830,000 |
| Annual | SaaS (25 users x 1,800 x 12 months) | 540,000 |
| Annual | Line | 144,000 |
| Annual | Maintenance (180,000 + 100,000) | 280,000 |
| Annual | Additional development | 400,000 |
| Annual | Annual total | 1,364,000 |
Five-year running cost: 2,700,000 + 720,000 + 1,400,000 + 2,000,000 = 6,820,000 THB
Five-year total: 11,650,000 THB
Side by side, the 3 scenarios look like this.
| Scenario | Initial total | Annual total | 5-year running | 5-year total |
|---|---|---|---|---|
| A Full cloud | 2,670,000 | 1,308,000 | 6,540,000 | 9,210,000 |
| B Full on-premise | 6,780,000 | 900,000 | 4,500,000 | 11,280,000 |
| C Hybrid | 4,830,000 | 1,364,000 | 6,820,000 | 11,650,000 |
The differences are as follows.
- B − A = 2,070,000 THB (over 5 years, on-premise costs 2,070,000 THB more than cloud)
- C − A = 2,440,000 THB
- C − B = 370,000 THB
3 things are worth reading out of this.
First, over 5 years cloud is the cheapest. The gap is 2,070,000 THB — not as large as the gap in initial cost (6,780,000 − 2,670,000 = 4,110,000), but not an amount that can be ignored either.
Second, scenario C, the hybrid, is the most expensive. This is the point at which many meetings turn over. Hybrid tends to be thought of as “in between cloud and on-premise, so the cost is in between too,” but in reality it carries both sets of initial costs and both sets of running costs. The edge server, the SaaS licences and the line redundancy are all needed. Implementation support and additional development both grow by the amount of integration design involved. The result is 2,440,000 THB more than full cloud and 370,000 THB more than full on-premise.
Third, as 5-year totals, these 3 figures sit within the same order of magnitude. The gap between 9,210,000 and 11,650,000 is not a difference in digits large enough to settle the decision in one stroke. Which is precisely why an axis other than cost is needed.
The crossover is 10.1 years, or 13.4 years once a refresh is included
On-premise, with its high initial cost and low annual cost, will necessarily fall below cloud if you take a long enough period. In which year, then?
The cumulative cost formulas are as follows (n is the number of years).
- Cumulative A = 2,670,000 + 1,308,000n
- Cumulative B = 6,780,000 + 900,000n
The difference in annual cost is 1,308,000 − 900,000 = 408,000. The difference in initial cost is 4,110,000. Equality holds when 408,000n = 4,110,000, that is, at n = 10.1 years.
Checking the arithmetic: at n = 5, A = 9,210,000 and B = 11,280,000. At n = 10, A = 15,750,000 and B = 15,780,000. B is still higher in year 10, and the crossover can be confirmed to occur just after that.
To this, add one cost that will certainly occur in reality: the hardware and OS refresh. Server maintenance runs out after 5 years, and operating systems have support end dates too. Insert a one-off 1,350,000 THB in year 6 — 900,000 for servers plus 450,000 for OS and database licences — and:
- Cumulative B = 8,130,000 + 900,000n
- 408,000n = 5,460,000 → n = 13.4 years
13.4 years. What that number means is clear. Factory system investments are normally appraised over 5 years, 7 at the most. A cost difference that reverses in year 13.4 does not fall inside the investment-decision horizon. In other words, comparing cloud and on-premise on cost does not produce an answer at this scale. The meeting never ends because it is being argued anyway.
To be clear, this is not an argument that “cloud is cheaper, so go cloud.” The opposite. It is an argument that if the difference is not decisive at 5 years or even 10 years, cost should be taken out of the decision inputs. Take it out, and then decide on allowable downtime.
Over a fixed 5 years, the crossover is not years but headcount (about 59 users)
There is another axis on cost that gets overlooked: the number of users.
The licences in scenario A are metered at 1,800 THB per person per month. As users grow, the cost grows in proportion. A perpetual on-premise licence, up to a point, does not change in total as headcount grows. Therefore, when you look at a fixed period of 5 years, the crossover happens not in years but in headcount.
With u users, the 5-year total for scenario A is as follows.
A 5-year total = 4,890,000 + 108,000u
The breakdown is initial 2,670,000 + 5 years x (annual licences 21,600u + line 144,000 + additional development 300,000) = 2,670,000 + 108,000u + 2,220,000.
This draws level with scenario B’s 11,280,000 when 108,000u = 6,390,000, that is, at u ≈ 59 users.
Checking the arithmetic: at u = 40, the total is 9,210,000 (matching scenario A). At u = 70 it is 12,450,000, above scenario B’s 11,280,000.
The practical implication is significant. The model factory has 40 users, but the user base of a production management system grows after go-live. Quality is added, purchasing is added, people at the Japanese head office are added as viewers, part of it is opened to subcontractors. A system contracted for 40 users reaching 70 3 years later is a common story.
So if you are going to use cost as a deciding axis, what to look at is not “in how many years does it cross over” but “at how many users does it cross over, and how many users will we have in 5 years’ time.” Having produced the figure of 59, it is worth confirming the projected headcount in 5 years with your executives. If there is a clear prospect of exceeding 59, negotiating the licence structure (a flat plan by headcount band, a cheaper category for view-only users) becomes a bargaining point during selection. That negotiation cannot be had after the contract is signed.
Data residency and the PDPA in Thailand | when cross-border transfer is an issue and when it is not
Back to data residency. Bangkok now having regions widens the options, but it is still worth working out which of the data in a production management system is actually within scope of the regulation.
On Thailand’s PDPA (Personal Data Protection Act), the rules on cross-border transfer (the whitelist notification and the notification on BCRs and appropriate safeguards) came into force on 24 March 2024. There are 3 bases for transfer: (1) transfer to a country with an adequacy finding, (2) BCRs (binding corporate rules), and (3) appropriate safeguards such as standard contractual clauses.
What matters in practice is that as of 2025 the PDPC (Personal Data Protection Committee) has not published a list of countries with adequacy findings. For the time being, therefore, all cross-border transfers are treated as transfers to a non-recognised destination, and safeguards under (2) or (3) are required. The reasoning “Japan has its own personal information protection act, so we are fine” does not hold, at least not on basis (1).
Looked at calmly, most of the data a production management system handles is not personal data. Part numbers, BOMs, processes, equipment, inventory quantities, lot numbers, costs. These fall outside the PDPA. Placing them in an overseas region does not raise the cross-border transfer question under the PDPA.
The items that do become an issue are the following.
- Operator IDs and names attached to shop-floor entries (the record of who made what and when)
- Attendance data, where entry and exit records or timekeeping are integrated
- Biometric data, where fingerprint or facial recognition is used (generally treated as sensitive personal data, which calls for more careful handling)
- Contact names and details of business partners (these sit in the purchasing module)
Of these, biometric data warrants particular care. The earlier discussion of how to handle authentication and the discussion of data residency intersect here. A request to “log in on the floor with fingerprints” is operationally reasonable, but at that moment you take on data that requires more careful handling. Where the templates will be stored, whether they leave the country, and how consent will be obtained are all worth confirming with legal at the design stage.
There is one further point that tends to be missed. Even when the data sits in a region inside Thailand, if operations staff access it from a support base outside the country, that can itself amount to a transfer. SaaS vendors with support based in Singapore or India are not unusual. Both the data processing clauses in the contract and the actual access paths need checking. A salesperson saying “the data is inside Thailand” is a statement about where it is stored, not about where it is accessed from.
Final judgement on legal interpretation should always be confirmed with your legal department or a local lawyer. What this article can offer is a list of the items to check.
4 deadlines to fix if you choose on-premise
The article has read as cloud-leaning up to this point, but there is of course a rational case for on-premise. It works when the line is down, the response is fast, and the data is physically at hand. For a factory carrying many functions with short allowable downtime, it is the right choice.
If you do choose on-premise, though, 4 deadlines are worth fixing in advance. Implement without them and, 5 years from now, you end up with a server that runs but that nobody can touch.
1. The OS support end date. Extended support for Windows Server 2016 (Microsoft Corporation) ends on 12 January 2027. If you have a server in the factory running on 2016 today, the remaining grace period can be worked back from there. When building new, the extended support end date of the chosen OS is worth writing into the approval request. With that one line in place, the timing for the next refresh budget is set automatically.
2. The database support end date. Mainstream support for SQL Server 2019 (Microsoft Corporation) ended on 28 February 2025, and extended support ends on 8 January 2030. Mainstream support having ended means that, as a rule, neither new features nor fixes involving specification changes will be provided. It runs, but it cannot answer new requirements.
3. The hardware maintenance end date. Server maintenance contracts normally run 5 years. Past 5 years, the supply of maintenance parts stops, and the recovery route in the event of a failure becomes “find the same model second-hand.” Maintaining an RTO of 15 minutes in that state is not possible. It is worth putting the maintenance contract end date and the RTOs by function into the same table and checking them together.
4. The people deadline. The most overlooked of them all. Will the person who configured that server still be at this site in 5 years? IT staff at Japanese-owned factories are either expatriates with fixed terms, or local staff who change jobs. An on-premise environment built around one individual turns into a black box the moment that person leaves. It is worth tying the storage of configuration information and passwords to the organisation rather than to an individual. Concretely: collect the architecture diagram, the IP address register, the list of administrator accounts and the vendor contacts into a single document, keep copies in 2 places — the Japanese head office and the local site — and update it once a year. Whether or not that practice exists changes what your options are in 5 years’ time.
Note that of these 4 deadlines, the first 3 move into the provider’s area of responsibility if you choose cloud. What the cost of cloud includes is not just compute resources; it is also the outsourcing of this deadline management. Whether the 408,000 THB annual difference in the model looks expensive or reasonable depends on whether you have the internal capacity to keep doing that management.
3 hybrid patterns
As scenario C showed, hybrid is not cheap. It is 2,440,000 THB more than full cloud and 370,000 THB more than full on-premise. The reason to choose hybrid anyway is not cost but allowable downtime. Unless this is shared at the outset, the argument “we went hybrid and it is still expensive” will come up partway through the approval process without fail.
What you are buying is not a lower figure but allowable downtime. That belongs on the first line of the proposal.
With that established, there are 3 hybrid patterns in practice.
Pattern 1: shop-floor functions on-premise, management functions in the cloud. This is scenario C. Functions with short allowable downtime — label printing, shop-floor data entry, inventory allocation — sit on an edge server inside the factory, while purchasing, costing, analysis and master maintenance sit in the cloud. The boundary is clear in business terms, so the division of responsibility is clear too. The weaknesses are holding licences on both the shop-floor side and the cloud side, and the integration layer becoming additional development. That is why additional development in the model is 400,000 THB, higher than the 300,000 THB for full cloud.
Pattern 2: the system in the cloud, with only a read-only cache and printing inside the factory. The whole system sits in the cloud, and the factory holds only a replica of the part-number master and inventory, plus label printing. Because writes always go to the cloud, data consistency is easier to keep. While the line is down, you can “look” and “print” but not “register.” It can be built more cheaply than pattern 1, but since shop-floor data entry cannot continue, it is not enough for factories where the RTO for shop-floor data entry is short.
Pattern 3: production in the cloud, on-premise for degraded operation only. Everything runs in the cloud in normal times, and the server inside the factory is there as a degraded-mode facility to be started only during a disaster or a long line outage. It carries the DR (disaster recovery) way of thinking over directly. The biggest pitfall is that because the degraded side is not used in normal times, it does not work when the moment comes. A standby that is not used will rot, without exception. A drill running half a day on the degraded side alone, once a quarter, is worth building into the operational design. If such a drill cannot be built in, this pattern is better not chosen.
For all 3 patterns, the method of deciding is the same. Put the table of RTOs by function alongside, and judge line by line: “on which side does this function need to sit in order to meet the value in the table?” That is all. There is no point at which technical preference decides it.
7 items to decide before you request quotations (a checklist)
Here are the items to settle internally before you go out for quotations. If these 7 items are filled in, you can compare proposals from different vendors on the same footing. If they are not, each vendor quotes on a different premise and no comparison is possible.
| # | Item to decide | Who decides | What happens if it is left blank |
|---|---|---|---|
| 1 | RTO and RPO by function | Manufacturing + plant manager | One RTO is applied to the whole system, and over-investment and under-provisioning occur together |
| 2 | RTO for authentication (matched to the shortest function) | IT + plant manager | The first incident after migration happens in authentication |
| 3 | Round trips per instance of each main shop-floor operation | Have the vendor measure it | You end up with “we added bandwidth and it is still slow” |
| 4 | What can and cannot be done while the line is down | IT + manufacturing | During an incident the floor switches to paper on its own judgement |
| 5 | Location of personal data, and where support staff access from | IT + legal | PDPA questions surface after the contract and go-live is delayed |
| 6 | Five-year totals submitted in the same format (initial + annual x 5) | IT + finance | The proposal that is only cheap up front gets selected |
| 7 | Exit terms (data export formats, return of data on termination, price-increase clauses) | IT + purchasing | You cannot take your data out at the next refresh, which amounts to lock-in |
A note on item 7. If you go with SaaS, it is worth confirming in the contract what format you can receive your data in at the end of the agreement. “Exportable as CSV” is not enough. Whether the master hierarchy, BOM parent-child relationships, attachments and change history are included in the export needs confirming individually. Run for 5 years with this left vague and, at the next refresh, you land in a state of “migration is too expensive to move.” The same applies to price-increase clauses: if no annual cap on revisions is written into the contract, the 5-year total in your model does not hold together in the first place.
Item 6 also matters in practice. Without a common format, a proposal that is cheap only in initial cost and expensive annually looks favourable. As the model shows, at this scale the difference in annual cost becomes a difference of millions of THB over 5 years. The submission format is best specified by the buyer.
Frequently asked questions (FAQ)
Which is cheaper for a production management system, cloud or on-premise?
It depends on the comparison period. For the model factory in this article (180 employees, 40 users), the 5-year total is 9,210,000 THB for full cloud and 11,280,000 THB for full on-premise, making cloud 2,070,000 THB cheaper. Cumulative costs cross over in year 10.1, or in year 13.4 if a one-off hardware and OS refresh of 1,350,000 THB is added in year 6. Because that does not fall within the usual investment-appraisal horizon of 5 to 7 years, cost does not produce an answer at this scale. The decision is better made on allowable downtime.
What is the going rate for a cloud production management system?
It varies greatly with the size and requirements of the Thai site, so unit prices alone are not enough to judge by. For reference, price ranges for cloud production management systems in Japan are given as starting from roughly JPY 30,000-50,000 per month, around JPY 10,000 per user per month where charging is per user, with initial costs in the range of several hundred thousand yen. For on-premise, the indicative figures are roughly JPY 1,000,000-3,000,000 initial and about JPY 30,000 per month. These are Japanese domestic rates, however, and an implementation in Thailand carries additional effort for requirements definition, master migration, Thai-language support and shop-floor training. That is what the 1,800,000 THB of implementation support for full cloud in this article’s model represents. Comparison is better made on the total of initial + annual x 5 than on licence unit prices.
Can a factory in Thailand really run a cloud production management system over its line?
It is determined by the product of round trips per operation and RTT, rather than by bandwidth. Indicative RTT figures are 5-15ms for a region inside Bangkok, 25-40ms for Bangkok to Singapore, and 70-90ms for Bangkok to Tokyo. These are levels commonly observed and are indicative only, not fixed values — measure from your own line. With an implementation that makes 5 round trips per scan, a Tokyo region adds 5 x 80ms = 0.4 seconds in round trips alone. With a region inside Bangkok it is 5 x 10ms = 0.05 seconds. Adding bandwidth does not close that gap. There are only 2 levers: reduce the number of round trips, or shorten the distance.
In what cases is staying on-premise fine?
When functions with an allowable downtime of 5-15 minutes are at the centre of operations, and the site is somewhere that line redundancy is not realistic. In addition, it is conditional on having the capacity at the site to keep managing the end dates for OS, database and hardware maintenance. Maintain on-premise without that capacity and, each time a deadline arrives — such as extended support for Windows Server 2016 (Microsoft Corporation) ending on 12 January 2027 — you are forced into an emergency refresh. Conversely, there is no rational case for keeping functions that tolerate 3 business days of downtime, such as cost accumulation and purchasing, on-premise as well. The question is better framed as “which functions stay on-premise” rather than “should we stay on-premise.”
How long does a core-system cloud migration take?
Within the scope of a production management system, many projects allow roughly 1 year from requirements definition to go-live. That said, it swings widely with the size of the factory and the state of the existing data, so it is not a figure that generalises. A quotation under your own conditions is the way to find out. What extends the timeline is not the number of functions but the state of the masters. Duplicate part numbers, BOMs that do not match the physical product, units and supplier codes that differ by department. Projects where cleaning this up takes twice as long as expected are the most common kind. What determines whether a migration succeeds is the quality of the existing data, not the features of the new system. How long to run in parallel is also worth deciding in advance. Start without fixing the date on which the old system is switched off, and parallel running never ends.
Summary
In 2026, Bangkok acquired cloud regions and “keeping data inside Thailand” became a standard option for the first time. But only the data-residency constraint changed; the framework for the decision has not.
What to stop doing is trying to pick “cloud or on-premise” once for the whole system. Inside a production management system, label printing that can only be down for 5-15 minutes lives alongside cost accumulation that can be down for 3 business days. It is because one answer is being sought for both that the meeting never ends. The unit of decision is the function.
The deciding axis is not cost. In the model factory, the 5-year totals are 9,210,000 THB for full cloud, 11,280,000 THB for full on-premise and 11,650,000 THB for hybrid, with cumulative costs crossing over in year 10.1, or year 13.4 with one refresh included. A difference that does not fall inside a 5-to-7-year appraisal horizon does not work as a deciding axis. If you do look at a fixed 5 years, though, what settles the crossover is headcount rather than years: once users exceed about 59, the total rises above full on-premise. The user count in 5 years’ time is worth confirming first.
And the thing most easily left out is authentication. Authentication does not appear on the list of business functions, yet it sits in front of every one of them, and nobody has assigned it an RTO. That is why the first incident after a migration happens in authentication. The RTO for authentication is best matched to the shortest value in the function table.
If you choose hybrid, it is worth sharing at the outset that it is not because it is cheaper. In the model it is 2,440,000 THB more than full cloud and 370,000 THB more than full on-premise. What you are buying is not a lower figure but allowable downtime. Proceed without sharing that premise and there will be friction after go-live.
The next thing to do is not to collect quotations from 3 vendors. It is to have manufacturing and the plant manager fill in the allowable downtime for each function. With that table in hand, which side each function sits on is decided mechanically. Without the table, no package choice will let you make the judgement.
Sorting out which side each function belongs on is not something that can be judged without looking at how the floor actually operates and what the line actually does. Even with the same “15 minutes for label printing,” the time you can genuinely hold out changes with the size of the holding area and the departure time of the shipping truck. TOMAS TECH is based in Bangkok and has worked with Japanese manufacturers on factory IT and FA, including the selection and migration of production management systems and the operational design on the shop-floor side. If the package is not chosen yet, or even whether to go cloud is not settled yet, we are happy to be involved from that sorting-out stage. If you have your current function list and network configuration to hand, we can start by filling in the allowable downtime table together. You can reach us via our contact page.