Blog

2026.08.11

System Backup and Disaster Recovery 2026 | What You Protect Is Recovery, Not Copies

System Backup and Disaster Recovery 2026 | What You Protect Is Recovery, Not Copies

“The backup job completes successfully every night.” When an IT staff member at a factory in Thailand gives you that answer, there is exactly one thing to check next. Has anyone ever restarted the business from that backup? What a system backup posture for business systems actually protects is not the number of stored copies, but recovery itself. This article sets out the conditions for a posture that can genuinely recover, using two yardsticks – RTO and RPO – and a breakdown of the real cost, framed for the conditions of a site in Thailand.

Why “Backups Are Running” and “We Can Recover” Are Not the Same Thing

Whenever backup comes up, the first thing on the table is almost always the acquisition status. What time the job runs each night, how many generations are retained, where the success notification is sent. Audits check the same three points, and once the green checkmarks line up, the matter is treated as closed.

What an actual incident asks is something else entirely. To which point in time can you roll back? How many hours does the restore work take? Is the person who can execute that procedure in this country right now? Once the data is back, does the business actually run? A job success notification answers none of these questions.

This is where the central argument of this article sits. “A backup is running” and “we can recover” are two different things. Backup is the means; recovery is the goal. However thoroughly you measure the health of the means, that is not a guarantee that the goal can be reached. And yet, at most sites, the only thing being measured is the means.

That gap shows up as a hard number in the JIPDEC Corporate IT Utilization Trends Survey 2026. The survey ran from January 16 to January 20, 2026, covering 1,107 companies in Japan with 50 or more employees, with responses from staff working in IT strategy and information security. In that survey, the rate of companies that had experienced a ransomware infection was 45.8%, and the share that could not recover their systems and data without paying a ransom was 13.0%. The latter figure rose from 10.5% in the previous survey.

How you read that 13.0% matters. Assuming these companies simply had no backups is almost certainly wrong. They had backups, but could not restore from them – or nothing in a restorable state remained. The fact that the share unable to recover went up rather than down tells you that the spread of backup and the ability to recover are not moving in the same direction.

So the first decision in a selection or design exercise is not a product name. It is a design judgment – what values to set for RTO (Recovery Time Objective) and RPO (Recovery Point Objective) for each business tier. Compare products before those two numbers exist and the only remaining axes of comparison are price and feature count, and whether you can actually recover goes unverified right through to go-live.

Why Sites in Thailand Remain Ransomware Targets in 2026

Raise this at a site in Thailand and you sometimes hear, “surely we are less of a target than headquarters in Japan.” The actual distribution does not match that intuition.

According to figures compiled by Check Point Software Technologies, of the more than 1,230 ransomware attack cases published between July and September 2024, 13% occurred in the Asia-Pacific region. That is not a negligible share at the regional level. On top of that, the National Cyber Security Agency (NCSA) of Thailand reports that ransomware damage inside Thailand in 2023 increased by a factor of 1.5 over 2022. The upward trend is no exception in a country where Japanese-affiliated sites are concentrated.

Damage at Japanese-affiliated manufacturing sites themselves has also been reported. Incidents involving production stoppage, data encryption, and information theft occurred in Thailand from 2022 through 2024, in Malaysia from 2022 through 2023, in Indonesia in 2021 and 2024, and in Vietnam in 2024. A distribution spread across all of ASEAN, and across multiple years at that, cannot be explained by the special circumstances of one country or one year.

Being a Manufacturer Makes the Odds Worse

Break the figures down by industry and the picture gets harder. In the JIPDEC survey, the infection experience rate for manufacturing was 57.1%, well above the 45.8% overall figure. Manufacturing also showed that 18.2% failed to recover even after paying the ransom. The fact that payment is no guarantee of recovery shows up clearly as an industry-level number.

The content of the damage is worth looking at too. In the same survey, the most common impact was loss or corruption of data at 51.3%, followed by leakage of confidential information at 35.1%. Leakage rose about 6 percentage points from 29.3% in the previous survey. This is where posture design is affected. If data is only encrypted, a healthy backup gets it back. But stolen data cannot be undone no matter how many backup copies you hold. Draw the line that backup is an availability control, not a confidentiality control as early as the internal briefing material.

Then there is the recovery period. In the same survey, the most common answer for time to recovery was “one week to one month” at 34.7%. Is your plan one you could explain to management on the assumption that the factory stops for one week to one month? If you cannot answer that, the design work has not started yet.

Protecting a site in Thailand all the way down to the production line goes beyond the scope of information systems alone. Segmentation on the control network side is covered in OT Security for Factories in Thailand – A Practical Guide for 2026. This article narrows in on the backup and recovery design within that wider picture.

Three Classic Patterns Where the Backup “Was Running” but Recovery Failed

Break down cases where recovery failed and the causes fall into roughly three types. None of them can be detected by a job success notification.

Pattern 1 — Never Tested Once

This is the most common type. The backup job succeeds, but nobody has ever actually restored from it. Any number of defects go unnoticed in that state. A critical folder or database has dropped out of the backup scope. An online database backup is being taken in an inconsistent state. The data is encrypted but nobody knows where the decryption key is kept. The job succeeds but the files are corrupted.

None of these surface until you try to restore. And the first chance to restore is usually a real production incident. Performing a recovery procedure for the first time in production would be considered abnormal by any standard of process management, yet in the world of backup it has become the norm.

Pattern 2 — Sitting on the Same Network as Production

The second is a lack of isolation. The backup server sits in the same network segment as production, is managed under the same authentication platform, and can be logged into with the same administrator account. In that configuration, the moment production is compromised the backup falls inside the blast radius as well.

The first thing a ransomware operator looks for is the backup. Destroying the means of recovery before encrypting raises the probability of payment. A backup area mounted as a shared folder, a domain-joined backup server, and a privileged account shared with production – when those three line up, the backup becomes the first thing the attacker deletes.

Pattern 3 — The Restore Procedure Lives in One Person’s Head

The third is people. There is no procedure document, or there is one but it bears little resemblance to the actual work. Only one specific person can perform the recovery, and that person is at headquarters in Japan, or has already left the company. At sites in Thailand, this hole opens easily through vendor staff turnover and local staff attrition.

Incidents do not restrict themselves to weekday business hours. When one happens late at night or over a long holiday, who starts the recovery, following which procedure, using which privileges? A posture with only one contact path becomes a posture with infinite recovery time the moment that one path fails to connect.

What these three share is that none of them are visible from the acquisition side of backup. However much you strengthen the machinery for measuring the health of acquisition, the 3 things that matter — testing, isolation, and procedure — will not improve. The object of measurement has to move to the recovery side.

What the Nichirei Case Shows About a Backup That Could Actually Recover

So what does a backup that can actually recover look like? A case published in 2026 is instructive.

The Nichirei Group detected an event suspected to be a cyberattack at around 6:50 a.m. on July 13, 2026, immediately established an emergency response headquarters, and performed an emergency shutdown of systems across the entire group. As a result of that response, shipping operations at cold storage warehouses and food plants restarted in stages from July 17, 2026 – four days after the incident – and a return to normal operation was in sight after roughly one week.

In the JIPDEC survey seen earlier, the most common recovery period was “one week to one month” at 34.7%. Set against that distribution, you can see how different it is to be entering staged resumption of shipping in four days and to have normal operation in sight in about a week.

The Two Conditions That Produced That Speed

From what has been published, the conditions behind that speed can be organized into two.

The first is that a backup system logically and physically separated from the production network survived. In other words, the opposite of Pattern 2 in the previous section was in place. What matters is that both were present – not logical separation alone, and not physical separation alone. Logical separation means separation of authentication and network paths; physical separation means separation of equipment and location. With only one of the two, there remains a possibility of collapse the moment an attacker obtains administrative privileges.

The second is an immediate company-wide shutdown as the first response. The faster the decision from detection to shutdown, the narrower the range that gets encrypted. What matters here is not technology but the design of decision-making. Who can decide, on what information, to stop operations up to what scope? If that authority has not been delegated in advance, several hours disappear into confirmation and reporting.

Degraded Operation Ran in Parallel

One more thing that should not be missed is that handwritten slips were used to run degraded operations in parallel, minimizing the logistics stoppage. Rather than halting the business until the system is fully back, whatever can be run on paper is run on paper. This is not a feature of a backup product; it is a design decision on the business side.

Factories need the same thinking. Of production instructions, goods receipt and issue, shipping, and inspection records, which ones could you run on paper and for how many days? Are the forms needed to run them printed and stored? Once the system is back, how do you load the paper-processed volume into the system? Settle those three points and the demands placed on RTO drop to a realistic level. Designing degraded operation is, in effect, a way to buy RTO.

RTO and RPO — Two Yardsticks, and Why They Must Differ by Business Tier

Here we define the two terms at the center of posture design. RTO and RPO design is work that comes before choosing a backup product.

RPO (Recovery Point Objective) states the point in time you need to be able to return to. Put differently, it is the range of data that can be lost while the business still holds together. An RPO of 1 hour means you build the business on the assumption that the last hour of data before the incident is lost. Shortening RPO requires raising acquisition frequency, and the higher the frequency, the higher the cost in bandwidth, storage, and processing load.

RTO (Recovery Time Objective) states how long it takes to resume the business. Having the data in hand and being able to resume the business are different things. Procuring hardware for the recovery target, rebuilding the OS and middleware, data transfer time, functional verification, and notifying users. RTO is the sum of all of these. A plan that estimates only the restore time from backup and tells management “we will be back in 2 hours” is almost certain to be wrong.

A Single Company-Wide Setting Always Breaks

Try to set these two uniformly across the company and you break in one of two directions.

Align to the strict end and cost explodes. Setting RPO to 1 hour and RTO to 4 hours for every system means protecting the shared folder and the test environment at the same level as everything else. The value protected and the cost do not balance.

Align to the loose end and the business stops. Put the core system RPO at 24 hours and every order, shipment, and inventory movement from the day of the incident disappears. In a factory, that triggers work to reconstruct one day’s records from paper, and no new production advances while the reconstruction is under way.

Therefore, varying RTO and RPO by business tier is the only realistic answer. When the numbers differ by tier, the means deployed and the money spent differ too. That becomes the skeleton of the posture design.

Decide Who Decides, First

There is one more practical caution. RTO and RPO are not numbers for the IT department to set. The people who know “how much loss occurs after how many hours of downtime” are the business departments. Numbers set unilaterally by IT get overturned the moment a disaster happens, with “nobody told us those were the conditions.”

As an approach, what works is to put two questions to the business departments – “when this operation stops, how many hours can you endure?” and “how many hours’ worth of the most recent data could you reconstruct from paper or memory if it disappeared?” – and then have IT translate the answers into cost. Once the cost is visible, the business departments revise their demands to a realistic level. Running that round trip once makes later consensus dramatically faster.

Table A — RTO and RPO Design Matrix by Business Tier

Now we put the thinking so far into an actual design table. The following is the framework TOMAS TECH uses in posture design for factories. It is not an external statistical survey but a design guideline derived from the nature of each business tier. Use it as a starting point, on the assumption that you will move the numbers to fit your own circumstances.

Business tierRPORTORecommended backup method
Core systems (ERP and production control)Within 1 hourWithin 4 hoursCloud and on-premises duplication, immutable snapshots
Production line control and SCADAWithin 24 hoursNext business dayLocal NAS with weekly cloud sync
Shared files and drawingsWithin 24 hoursNext business dayCloud backup (3-generation retention)
Email and groupware (Microsoft 365 and similar)Within a few hoursWithin a few hoursBackup service built for SaaS
System Backup and Disaster Recovery 2026 | What You Protect Is Recovery, Not Copies - figure 1

How to Read the Table

Core systems are set at an RPO within 1 hour and an RTO within 4 hours because when this tier stops, receiving and shipping themselves stop. Orders, goods issues, and inventory movements all converge on this one line. This is the one tier built on cloud and on-premises duplication as a premise, combined with immutable snapshots – copies that cannot later be deleted or rewritten. Doubling up the means raises the cost, but if you are going to raise cost anywhere, raise it here.

Production line control and SCADA are set at an RPO within 24 hours because the main data in this domain is setpoints, recipes, and ladder logic, and the volume that changes day to day is small. Changes happen at changeover or modification, not every hour. However, the RTO of next business day is not because recovery is easy – it is because recovery does not complete over the network and involves work on site. If an equipment vendor has to attend, the time stretches further. Confirm who can physically come before you argue about method. Note also that the method is given as “local NAS with weekly cloud sync” because the local NAS side carries the daily acquisition that meets the 24-hour RPO, while the cloud side is narrowed to the role of holding an independent copy at a remote location. The weekly sync is positioned as the last line of defense if the site is lost in its entirety.

Shared files and drawings are a tier with high volume, moderate change frequency, and an impact from loss that usually amounts to redoing work. Retention is set at 3 generations because corruption or accidental overwrite is often not detected in the most recent generation and is noticed several days later.

Email and groupware is the only tier with both RTO and RPO within a few hours, and the reason is that this tier is less a business operation in itself than the communication path for every operation. If the means of communication is down during incident response, every other recovery task slows. And, importantly, a SaaS such as Microsoft 365 may have high availability on the provider side yet offer only limited standard capability to roll back over long periods against user error or destruction from inside. That is precisely why a separate backup service built for SaaS is required.

How to Turn This Table Into Your Own Version

Do not use this table as is – rebuild it as your own version. The procedure has three stages. First, classify your own business systems into whichever of these 4 tiers they fall under. Anything that fits nowhere – quality inspection data or a cost accounting aggregation platform, for example – gets added as a fifth row. Next, collect target RTO and RPO values through the two questions to the business departments. Finally, lay out the methods and cost required to meet those targets, and negotiate looser numbers only on the rows where cost is excessive.

The deliverable of this exercise is a single table. It can take 3 meetings to produce one table, but choose a product without it and you will not be able to explain the basis of the selection afterwards.

The 3-2-1 Rule and Immutable Design — Why You Need a Copy That Cannot Be Deleted

Once RTO and RPO are set, the discussion moves to means. The reference point here is the 3-2-1 rule – keep 3 copies of the data, on 2 different media types, with 1 of them stored at a different location.

The rule itself is old, but its meaning changed once the ransomware era began. What 3-2-1 used to protect against was equipment failure, fire, and flood. What it has to protect against now is an attacker who has taken administrative privileges and deliberately hunts down the backup to destroy it. When the adversary changes, the same rule requires a different implementation.

System Backup and Disaster Recovery 2026 | What You Protect Is Recovery, Not Copies - figure 2

Do the 3 Copies Share the Same Fate?

The first thing to check is whether the 3 copies are genuinely independent. A common failure is a production server, a backup server in the same rack, and a NAS in the same building. There are 3 copies, but a fire or a compromise takes all of them at once. We call a state that satisfies the count without achieving independence “apparent 3-2-1.”

The test is simple. Check whether the privileges of a single administrator account can delete all 3. If they can, it is the same as having 1 copy. Attackers move after taking privileges, so the reach of those privileges is the effective blast radius.

Immutable, Meaning a Copy That Cannot Be Deleted

The direct answer to this problem is immutable design. A copy is stored with a retention period set, and during that period not even an administrator can delete or alter it. Object lock features in cloud storage, and write-once areas on dedicated appliances, fall into this category.

With an immutable copy, even if an attacker takes administrative privileges, that one copy survives. Put the other way round, in a configuration with no immutable copy at all, loss of administrative privileges means loss of every copy. Immutable snapshots are specified only for core systems in Table A because that is the tier that must be the last thing standing when privileges are lost.

Offline as the Last Resort

The other option is offline. Data is written out to tape or removable disk and physically detached from the equipment for storage. What is not connected to the network cannot be deleted over the network. It takes effort, but in terms of certainty it remains at the top even today.

The caution when adopting this at a site in Thailand is the storage environment. Long-term storage in high heat and humidity accelerates media degradation. Secure a locked cabinet in the office or a separate air-conditioned site. And perform a read from the offline media at least once a year. Media you cannot read is the same as media you do not have.

Write Out the 3 Layers of Isolation

In practice we recommend writing isolation out across 3 layers. Network isolation – the backup area cannot be reached directly from production. Authentication isolation – a production administrator account cannot log in on the backup side. And location isolation – there is no dependency on the same building or the same power supply. Put these 3 into a table and fill in “achieved” or “not achieved” for each, and your weak points become obvious.

Cloud Versus On-Premises — Which Is Realistic for a Site in Thailand

Once the means are set, the discussion becomes where to put the data. Get a feel for cloud backup cost first, then apply it to the conditions of a site in Thailand.

What the Pricing Actually Looks Like

Organize the published price examples and billing models split broadly into 2 types.

Billing modelExample servicePrice
Capacity-basedAcronis powered by (100GB)From JPY 3,000 per month, JPY 36,200 per year
Capacity-basedAcronis powered by (1TB)From JPY 20,000 per month
User-basedVeeam Data Cloud for Microsoft 365JPY 9,822 to 10,913 per user per year (depends on scale)
User-basedAvePoint Cloud BackupFrom JPY 3,840 per year at 3-year retention, from JPY 6,000 per year at unlimited retention
User-basedSysCloudJPY 5,890 per user per year

As a general market feel, the figures cited are from around JPY 1,000 per month per PC, JPY 20,000 to the JPY 50,000 range per month for a 1TB server, and roughly JPY 300 to 900 per month for user-based pricing.

There are 2 things to take from this list. The first is that prices differ between user-based services, and one driver of that difference is retention period. The gap between 3-year retention and unlimited retention at AvePoint Cloud Backup is the clearest example, and some services such as Veeam Data Cloud for Microsoft 365 also move on unit price by scale. Retention period is determined not by RPO but by when you notice. Accidental deletion and insider misconduct sometimes come to light months later. The second is the width of the monthly figure for a 1TB server under capacity-based pricing, from JPY 20,000 to the JPY 50,000 range. The same 1TB moves substantially depending on acquisition frequency, generation count, and how transfer charges at recovery are handled. When you request a quote, always confirm not just capacity but what recovery itself will cost. A structure with cheap storage during normal operation and expensive retrieval creates a reason to hesitate over recovery at exactly the wrong moment.

The Decision Axes at a Site in Thailand

There are 3 axes for deciding cloud or on-premises at a site in Thailand.

The first is bandwidth. Backup acquisition can be slow, but recovery is meaningless unless it is fast. Measure how many hours it takes to pull 1TB of data back from the cloud. If that time exceeds RTO, a cloud-only configuration does not hold. That calculation is the reason Table A specifies cloud and on-premises duplication for core systems. The division of roles is to restore quickly from the local copy, and to use the cloud as the last line of defense if the local copy is lost.

The second is local operations staff. On-premises equipment requires someone to physically look after it – media swaps, hardware failure response, firmware updates. Choose an on-premises-centric configuration at a site with no resident staff member in charge and operations can sit stalled for months.

The third is power and installation environment. Outages and voltage fluctuation affect both equipment lifespan and data consistency. If you choose on-premises, treat an uninterruptible power supply and air conditioning at the installation site as prerequisites and put them in the cost.

Using Both Is the Baseline

The conclusion is that this is not a question of choosing one. Combining local storage to buy recovery speed with cloud storage to secure independence is the baseline form. That configuration also happens to be a natural implementation of the 3-2-1 rule. What you have to decide is not “which one” but which business tier gets restored from which.

Table B — The 5 Cost Layers of a Backup Posture

Look at cost as a single quote and you cannot tell whether a cheap proposal is cheap or merely narrow in scope. Split it into layers.

Assumptions for the Model Factory

The following is a model factory estimate built on the following assumptions. It assumes a manufacturer with a site in Thailand, 1TB of data in its ERP and production control systems, 40 Microsoft 365 users, and 2 restore drills per year.

LayerContentIndicative annual cost
Layer 1Storage charges (cloud backup for 1TB of ERP and production control data)JPY 240,000
Layer 2Backup of Microsoft 365 accounts (40 users)JPY 153,600
Layer 3Restore drills and operational effort (2 per year)JPY 256,000
Layer 4Disaster recovery standby environment (remote replication and standby systems)JPY 720,000
Layer 5Audit and PDPA documentation (annual external review)JPY 300,000
TotalJPY 1,669,600

These amounts are an estimate based on the model factory assumptions and are not a guarantee of prevailing market prices. Change any one of data volume, user count, site configuration, or required RTO and the amounts move considerably.

System Backup and Disaster Recovery 2026 | What You Protect Is Recovery, Not Copies - figure 3

The Basis for Each Layer and How to Read It

Layer 1 is based on the capacity-based pricing seen in the previous section. The level of JPY 20,000 per month for a 1TB server is counted over 12 months to give JPY 240,000. This scales with data volume, so measuring your own volume lets you substitute directly.

Layer 2 is based on user-based pricing. The level of JPY 3,840 per user per year for a service with 3-year retention is multiplied by 40 users to give JPY 153,600. Choose to extend retention to unlimited and this layer rises.

Layer 3 is restore drills and day-to-day operational effort. Making this explicit as a cost is one of the key points of this table. A drill is not “work we could do if we felt like it” – it consumes staff time. Leave it out of the numbers and the proposal looks cheap on paper, then gets deferred at execution time because “the day job is busy.”

Layer 4 is the disaster recovery standby environment. It is the largest of the 5 layers, and also the one most easily cut. What happens when it is cut is covered in the next section.

Layer 5 is audit and PDPA documentation, assuming one external review per year. This layer is a cost for accountability rather than technology, and it connects directly to the Thai legal requirements discussed later.

Which Layers Can Be Cut and Which Cannot

An argument about where to cut under a limited budget is inevitable. What to watch here is that a layer carrying no visible effect gets cut. The Layer 4 disaster recovery standby environment produces no effect at all during normal operation. That is why it is the first candidate.

The decision to cut is not wrong in itself. But if you cut it, write back into Table A how far RTO stretches as a result. In a configuration with no standby environment, an incident begins with procuring hardware. Factor in the time it takes to procure and install server hardware in Thailand and the design of an RTO within 4 hours for core systems no longer holds. What you are cutting is not cost, but the RTO promise. Drop Layer 4 alone without documenting that correspondence in the table and you are left with RTO 4 hours on the posture diagram and several weeks in reality.

Note that these 5 layers are a slice covering the backup posture only. To see where they sit within the total cost of keeping systems running, read them alongside the maintenance cost breakdown covered in System Maintenance Cost for Factories 2026. Put backup alone into an approval request and there is no yardstick left for judging whether the amount is reasonable.

How Thailand’s Personal Data Protection Act (PDPA) Relates to Backup Posture

Discussion of backup and disaster recovery tends to proceed as a technical matter, but in Thailand it connects directly to legal requirements.

The 72-Hour Constraint

Thailand’s Personal Data Protection Act (PDPA) obliges data controllers to report a data breach to the Personal Data Protection Committee (PDPC) within 72 hours of becoming aware of it. Further, where there is risk to the rights of individuals, notification to data subjects must be made at the same time as the notification to the PDPC, and where multiple individuals are affected, this is required to be done through public channels such as the media.

Those 72 hours impose concrete requirements on posture design. To report, you have to identify what leaked. Which data in which system, as of what point in time, was compromised to what extent. That identification work is only possible when data from before the breach is in your hands. In other words, backup is not only a means of recovery – it is also the evidence for determining the scope of the breach.

Overwrite a compromised server in the course of recovery and the material for reconstructing what happened is gone. Rushing recovery, destroying the evidence, and then lacking the information needed for the 72-hour report is a sequence that really does happen. The countermeasure is simple – build into the plan the step of taking one disk image of the compromised state before recovery work begins.

Where ISO 27001 Fits

The other point to note is that Thailand’s Ministry of Digital Economy and Society (MDES) recognizes ISO 27001 certification as evidence of meeting a minimum security standard. Certification does not exempt you from the legal obligation, but it can be used as a means of demonstrating that a standard is met.

In practice, borrowing the document formats ISO 27001 requires is more useful than deciding whether to certify. An asset inventory, a risk assessment, a record of controls, and a procedure for when an incident occurs. If those 4 exist as documents, most of the information needed for the 72-hour report can be pulled from them. Layer 5 of Table B stands audit and PDPA documentation up as its own layer because these documents are not “done once written” – they have to keep being updated as the system configuration changes.

Drawing the Line Between Disaster Recovery and Backup

Let us settle the terminology here. Backup means making copies of data; disaster recovery is a posture that lets the business resume in a different location. The former is about data, the latter about the business.

The scope of disaster recovery to consider at a site in Thailand is not just systems. If a flood makes the factory itself inoperable, where, by whom, and which operations get continued? Restoring the data from backup does not resume the business without somewhere to run it and someone to run it. When you evaluate the Layer 4 standby environment, write down not only the hardware but who will operate that hardware and where.

How to Stop Recovery Drills From Becoming a Ritual

Even if the design in the preceding sections is correct, without testing you are back at the first pattern. Now to drills.

The Classic Route to Becoming an Empty Ritual

Left alone, restore tests always become empty ritual. The route is predictable – year 1 is done seriously, year 2 traces the previous year’s procedure document, and year 3 leaves only a record saying “same configuration as last time, therefore deemed performed.” Change the responsible person and nobody even notices that the procedure document has diverged from the actual configuration.

The way to prevent this is not to run more drills. It is to define pass and fail conditions as numbers.

Write Pass and Fail Conditions in Terms of RTO and RPO

The purpose of a drill is not to confirm “we got it back” but to confirm “we got it back within the time we set.” So the pass condition is pulled directly from Table A. For a core system drill, the pass condition is whether you reached a state where the business could resume within 4 hours. Fail to reach it and the drill is a fail, and a fail means fixing either the means or the number.

In that form, the drill result feeds back into the design. “It did not come back in 4 hours” becomes the input for deciding whether to add duplication or to relax RTO. A drill with no pass condition produces nothing but a record that it was performed.

The 4 Items to Record

Fix the items recorded at each drill at 4. The measured elapsed time from start to business resumption. The point in time of the data you were able to recover to (the measured RPO). The tasks that got stuck along the way and how long each took. And the places where the procedure document and the actual work diverged.

The fourth is the most valuable. A divergence from the procedure document will occur in exactly the same place at the next incident. Make fixing the procedure document immediately after the drill part of the same set. Head into the next drill without fixing it and you will rediscover the same divergence every time.

Who Should Perform It

Another useful technique is having someone other than the author of the procedure document perform it. Authors fill in the gaps unconsciously, so defects in the document never surface. Hand only the procedure document to local staff and ask them to execute it, and every ambiguity is exposed at once. At a site in Thailand, this method also flushes out language issues at the same time. A procedure document written only in Japanese, when the people responding late at night are local staff, is itself a recovery time risk.

The frequency of 2 per year is the level assumed in Layer 3 of Table B. You do not have to cover every system each time. Running them in rotation with a different scope each time – core systems first, shared files and email second – is more realistic.

Issues Specific to Production Control and Core Systems — Tying Design to the Cost of a Stopped Factory

There is one decisive difference between backing up a general information system and backing up a factory core system. When it stops, production stops.

Produce Your Own Downtime Cost Figure

When explaining the validity of RTO to management, the strongest basis is downtime cost. Industry averages from other companies are meaningless here; you need your own number.

The calculation requires 4 inputs. First, the scope of production that halts when the target system stops – all lines, or only certain processes. Second, the hourly production value of that scope. Third, the costs that keep accruing even while production is stopped, meaning labor and fixed equipment costs. Fourth, the additional costs arising from delivery delay, such as expedited freight and customer penalties.

Multiply these 4 together and you get the hourly cost of downtime. Only when that number sits next to the annual figure for Layer 4 of Table B can you decide whether to hold a standby environment. Judging that “JPY 720,000 is expensive” without knowing the downtime cost is a judgment with nothing to compare against.

What Makes Production Control Recovery Harder

There is another technical issue. A production control system does not run on its own. Data collection terminals, handheld terminals, label printers, measuring instruments, the accounting system above it, EDI with customers. The business only runs when all of these are connected.

Consequently, restoring the database alone does not resume the business. You have to verify that every peripheral connection is alive and that data is consistent with the connected systems. This is where it differs sharply from restoring a general file server. It is also the reason the pass condition for a drill should be “the business could resume” rather than “the database came back.”

Decide the Recovery Order in Advance

When multiple systems go down at once, decide in advance which gets restored first. The order can be derived from Table A. The principle is to restore the tiers with the shortest RTO first, but dependencies also matter. Sequencing constraints such as “no other system can be logged into until the authentication platform is back” have to be identified beforehand.

This ordering table is not something you can produce in the middle of an incident. Produce it during normal operation and confirm its validity in a drill.

Where Headquarters Standards and Local Sites Diverge — 3 Holes at Overseas Sites

At Japanese-affiliated sites in Thailand, the backup policy often comes down from headquarters in Japan. Three holes open up in that arrangement.

Hole 1 — The Standard Does Not Assume the Local Configuration

Headquarters standards are written on the premise of the headquarters system configuration. Systems that exist only at the local site – a mechanism built for local accounting requirements, or a production control system introduced locally – are sometimes outside the standard’s scope. It would be better if being outside scope were stated explicitly, but in most cases there is no mention at all. Anything not mentioned ends up being protected by nobody.

The remedy is simple – build an inventory of local systems and match each one against the items in the headquarters standard. Whatever does not match is the scope you must decide locally.

Hole 2 — Reporting Stops at “We Are Doing It”

The second is reporting granularity. When the reporting form to headquarters only asks yes or no on “are backups being performed,” nothing about the quality of that performance is conveyed. As covered above, performing backup and being able to recover are different things.

Add to the reporting form the date of the most recent recovery drill and the measured RTO and measured RPO. Those 3 items alone change the report from “are you doing it” to “can you get it back.” For headquarters, it also becomes a form that allows comparison between sites.

Hole 3 — Recovery Authority Does Not Sit Locally

The third is the most serious. Administrative privileges over backup, and the credentials needed for recovery work, exist only at headquarters, so the local site cannot begin recovery on its own judgment. Incidents do not consider time zones. Late night in Thailand is the small hours in Japan, and the several hours until someone answers get added straight onto RTO.

There are cases where handing full privileges to the local site is difficult from a security standpoint. In that case, prepare a procedure usable only in an emergency. Sealed credentials held in a local safe, with an after-the-fact report when used, is a perfectly acceptable arrangement. What is needed is the mindset that control during normal operation and reachability during an emergency are designed separately.

Sequencing the Rollout — The First 90 Days

A posture review that widens its scope never lands. Here is an approach bounded at 90 days.

The First 30 Days — Count What You Have

Do not talk about replacing systems; count what exists. There are 5 things to count. The inventory of target systems and the data volume of each. The current backup frequency and number of retained generations. Where backups are stored and how many administrator accounts can reach that location. The number of restores actually performed over the past year. And the scope of production that halts when the target system stops.

These 5 only need to be gathered from existing records. Aim for a perfect inventory and 30 days will not be enough. The purpose is not precision but building the base for the discussion.

Days 31 to 60 — Produce 2 Tables

Produce your own version of Table A, and the inspection sheet for the 3 layers of isolation. Table A involves interviews with the business departments, so it is the slower of the two. Bring one person each from sales, manufacturing, purchasing, accounting, and quality, and fill it in together in the same room. Simply discovering that the answers differ by department has value.

The isolation inspection sheet can be produced by IT alone. For the 3 layers of network, authentication, and location, fill in the current state as “achieved” or “not achieved.” Any cell you cannot fill is a place you have not investigated.

Days 61 to 90 — Drill One System Line

Take the tables you produced and run a recovery drill on just the single system line with the tightest RTO. Do not try to cover every system. Even with one line, actually doing it surfaces five or six unforeseen issues.

Judge the drill against the numbers in Table A. If it fails, separate whether the cause was insufficient means or an unrealistic number. The result of that separation becomes the direct input for the next investment decision. Only once you reach this stage does requesting product and service quotes begin to mean anything.

Beyond 90 Days

Once you reach a drill on one system line at 90 days, extend to the remaining tiers in sequence. In parallel, put your own numbers into the 5 layers of Table B, calculate the annual figure, and produce material setting it beside the downtime cost. That material becomes the basis for the following year’s budget request.

5 Common Failures

Here are the failures seen repeatedly in posture building.

FailureWhat happensHow to avoid it
Treating job success notifications as evidence of healthA state where acquisition succeeds but restore fails goes unnoticed for a long periodRun periodic recovery drills with pass conditions defined in terms of RTO and RPO
Managing backup under the same privileges as productionAll copies are lost simultaneously the moment administrative privileges are takenIsolate across the 3 layers of network, authentication, and location, and keep 1 immutable copy
Setting RTO and RPO uniformly across the companyEither cost becomes excessive, or core systems fail to meet the required levelVary the numbers by business tier and secure agreement from the business departments
Cutting the disaster recovery standby environment on price aloneThe RTO on the posture diagram diverges from reality, and an incident starts with hardware procurementIf you cut it, rewrite the RTO promise at the same time and reflect it in the table
Keeping recovery authority and procedures at headquarters onlyFor a local incident at night or on a holiday, the wait for contact is added on top of RTODefine a procedure the local site can start in an emergency, and how credentials are stored

Of these, the first two have the largest impact. The first creates a state where you cannot even notice there is a problem. The second creates a state where, by the time you notice, no means remain. Neither is detectable through normal-operation indicators.

The fourth needs a supplementary note. There are genuinely many sites that cannot hold a standby environment under budget constraints. The problem is not going without it, but going without it while leaving the RTO on the posture diagram unchanged. If it has been rewritten, the judgment management makes during an incident changes too.

Frequently Asked Questions

Why does a backup posture for business systems matter?

What matters is not that backups are taken but that you can recover from them. In the JIPDEC Corporate IT Utilization Trends Survey 2026, the ransomware infection experience rate was 45.8% and the share unable to recover without paying a ransom was 13.0%, up from 10.5% in the previous survey. In manufacturing the infection rate was higher still at 57.1%, and 18.2% failed to recover even after paying the ransom. Payment is no guarantee of recovery. Securing a recoverable state under your own power is a precondition for business continuity.

What is the difference between RTO and RPO?

RPO states the point in time you need to be able to return to – the range of data that can be lost while the business still holds together. RTO states how long it takes to resume the business. Shortening RPO requires raising acquisition frequency; shortening RTO requires preparing the recovery environment and procedures in advance. As a principle these two should differ by business tier, and the design table in this article sets core systems at RPO within 1 hour and RTO within 4 hours, and shared files and drawings at RPO within 24 hours and RTO of the next business day.

How much does cloud backup cost?

Published price examples put capacity-based pricing at 100GB from JPY 3,000 per month (JPY 36,200 per year) and 1TB from JPY 20,000 per month. Among user-based services for Microsoft 365, offerings range from JPY 3,840 per user per year up to JPY 9,822 to 10,913 per year. As a general market feel, the figures cited are from around JPY 1,000 per month per PC, JPY 20,000 to the JPY 50,000 range per month for a 1TB server, and roughly JPY 300 to 900 per month for user-based pricing. The model factory estimate in this article puts the annual figure at JPY 1,669,600, covering storage charges and Microsoft 365 backup plus drill effort, standby environment, and audit documentation.

What is the difference between backup and disaster recovery?

Backup means making copies of data; disaster recovery is a posture that lets the business resume in a different location. The former is about data, the latter about the business. Restoring the data from backup does not resume the business without the hardware, the location, and the people to run it. The 5 cost layers in this article stand the disaster recovery standby environment up as Layer 4 in its own right because it is a different kind of investment from the storage costs in Layers 1 and 2. At a site in Thailand, set the scope to include the case where a flood makes the factory itself inoperable.

How often should recovery drills be run?

The model factory in this article assumes 2 per year. You do not have to cover every system each time, though – running them in rotation with a different scope each time, core systems first and shared files and email second, is more realistic. More important than frequency is the pass condition. The purpose of a drill is not to confirm you got the data back but to confirm you got it back within the time you set, so the pass condition is pulled directly from the RTO and RPO defined in Table A. On top of that, having someone other than the author of the procedure document perform it flushes out both documentation ambiguity and language issues at once.

Summary

A backup running and being able to recover are two different things. A job success notification answers none of the questions of what point in time you can return to, how many hours it takes to resume the business, and who can execute it. The fact that in the JIPDEC Corporate IT Utilization Trends Survey 2026 the share unable to recover without paying a ransom rose from 10.5% in the previous survey to 13.0% tells you that the spread of backup and the ability to recover are not moving in the same direction.

There is no reason a site in Thailand would be an exception. Of the more than 1,230 ransomware attack cases published between July and September 2024, 13% occurred in the Asia-Pacific region, and damage inside Thailand in 2023 rose by a factor of 1.5 over 2022. Damage at Japanese-affiliated manufacturing sites has been reported across Thailand, Malaysia, Indonesia, and Vietnam over multiple years. The infection experience rate in manufacturing is 57.1%, and the share that failed to recover even after paying the ransom is 18.2%.

The causes of failed recovery come down to 3 – never tested, not isolated, and procedures locked inside one person. Conversely, in the Nichirei Group case, a backup system logically and physically separated from the production network survived, and a company-wide shutdown immediately after detection meant that from the detection on July 13, 2026 the group entered staged resumption of shipping operations four days later on July 17, with normal operation in sight in about one week. Running degraded operations on handwritten slips in parallel also helped narrow the scope of the stoppage.

The center of the design is RTO and RPO. Rather than one company-wide setting, vary the numbers by business tier – core systems at RPO within 1 hour and RTO within 4 hours with cloud and on-premises duplication plus immutable snapshots; production line control and SCADA at RPO within 24 hours and RTO of the next business day with a local NAS and weekly cloud sync; shared files and drawings at the same level with cloud backup at 3-generation retention; and email and groupware within a few hours with a backup service built for SaaS.

Cost splits into 5 layers. In the model factory estimate, storage charges are JPY 240,000, Microsoft 365 backup JPY 153,600, restore drills and operational effort JPY 256,000, the disaster recovery standby environment JPY 720,000, and audit and PDPA documentation JPY 300,000, for a total of JPY 1,669,600. Layer 4 is the one most easily cut, but write back into the table that what you are cutting is not cost but the RTO promise.

In Thailand, PDPA requirements bear directly on the posture as well. There is an obligation to report to the PDPC within 72 hours of becoming aware of a breach, and reporting requires identifying the scope of the breach. Build into the plan the step of taking one image of the compromised state before rushing into recovery. The fact that MDES recognizes ISO 27001 as evidence of meeting a minimum security standard is also a useful reference when choosing a documentation format.

So what to start on tomorrow is not a product comparison. Count the target systems and their data volumes. Agree RTO and RPO per business tier with the business departments and put them on a single table. Inspect the state of isolation across the 3 layers of network, authentication, and location. And run a recovery drill on the single system line with the tightest RTO, judging pass or fail by the numbers. Request quotes once those 4 are filled in and every vendor’s proposal lands on the same footing.

Even at the stage where your own RTO and RPO have not been set, it is perfectly fine to engage in a narrower way – looking only at your current backup configuration together and checking which of the 3 layers of isolation has a hole in it. TOMAS TECH supports Japanese-affiliated manufacturers in Thailand end to end, from measuring the current state and designing the posture to technical advice on product selection and the design and execution of recovery drills. Inspection sometimes reveals that fixing the operation of what you already have is enough, so please feel free to contact us even for a conversation that does not assume introducing a new product.

References

  • Corporate IT Utilization Trends Survey 2026 Press Release – JIPDEC Survey period January 16 to January 20, 2026, 1,107 companies in Japan with 50 or more employees, ransomware infection experience rate 45.8%, manufacturing 57.1%, share unable to recover without paying a ransom 13.0% against 10.5% in the previous survey, most common recovery period one week to one month at 34.7%, data loss and corruption 51.3%, confidential information leakage 35.1% against 29.3% in the previous survey, manufacturing failing to recover even after paying the ransom 18.2%
  • Trends in Ransomware Damage Across ASEAN – Risk Management Navi Check Point Software Technologies figures showing more than 1,230 published cases between July and September 2024 with 13% in the Asia-Pacific region, National Cyber Security Agency of Thailand figures showing damage inside Thailand in 2023 at 1.5 times the 2022 level, damage at Japanese-affiliated manufacturing sites in Thailand from 2022 to 2024, Malaysia from 2022 to 2023, Indonesia in 2021 and 2024, and Vietnam in 2024
  • Nichirei Cyberattack Response and Backup – london3.jp Detection at around 6:50 a.m. on July 13, 2026 and establishment of an emergency response headquarters, emergency shutdown of group-wide systems, survival of a backup system logically and physically separated from the production network, staged resumption of shipping operations from July 17, 2026, normal operation in sight after about one week, degraded operation using handwritten slips
  • Guide to the Thailand Personal Data Protection Act PDPA – LOGON International Obligation to report to the PDPC within 72 hours of becoming aware of a data breach, notification to data subjects where there is risk to individual rights, notification through public channels where multiple individuals are affected, and recognition by the Ministry of Digital Economy and Society of ISO 27001 as evidence of meeting a minimum security standard
  • Cloud Backup Cost Benchmarks – c-compe Security Software Comparison Capacity-based pricing from JPY 3,000 per month and JPY 36,200 per year for 100GB and from JPY 20,000 per month for 1TB, Veeam Data Cloud for Microsoft 365 at JPY 9,822 to 10,913 per user per year, AvePoint Cloud Backup from JPY 3,840 per year at 3-year retention and from JPY 6,000 per year at unlimited retention, SysCloud at JPY 5,890 per user per year, from around JPY 1,000 per month per PC, JPY 20,000 to the JPY 50,000 range per month for a 1TB server, and roughly JPY 300 to 900 per month for user-based pricing