Blog

2026.10.06

Industrial Historian Selection: RFP and FAT/SAT in Thailand

Industrial Historian Selection: RFP and FAT/SAT in Thailand

“The SCADA is supposed to keep history, but when we try to pull up the temperature on the day last year’s defect occurred, it is slow, there are gaps, and we cannot tell which tag is which.” We often hear this from control and instrumentation engineers and factory IT staff at Japanese-owned factories in Thailand. Here is the short answer: what you should gain from an industrial historian implementation is not a large amount of storage, but “time series data that can be trusted, carries context and can be replayed later”. To get there, you need to decide 4 things: (1) whether you need a historian at all, (2) what to record, at what accuracy and with what timestamps, (3) where on the network to place it, and (4) what to verify in acceptance testing. Comparing products goes better once these 4 have been decided.

All figures in this article such as tag counts, scan intervals, data volumes and buffer capacities, except those with a cited source, are original assumptions created for this article, based on the model factory described later. They are neither industry averages nor survey results. Product features and plans are presented as what each company has published (vendor announcements). None of them are TOMAS TECH products; they are all products of other companies.

Why Industrial Historian Implementation Is Being Questioned Now

What was announced at ICC 2026 in September 2026

There have been notable announcements around historians in the past few weeks. In a press release dated 23 September 2026, Inductive Automation announced that at ICC 2026, its annual conference held in Sacramento, USA (more than 1,800 attendees and more than 60 sessions, according to the company), it previewed a new “Enterprise Historian Module” for TimescaleDB as part of its next version, “Ignition 2027”. The company also stated that Tiger Data, the developer of TimescaleDB, has become a strategic partner, that the two companies will develop “TimescaleDB for Ignition”, and that this will be sold and supported directly by Inductive Automation.

According to a recap article the company published on 30 September, the module is intended to meet requirements for query performance, high availability and storage capacity. Ignition 2027 is planned for release in February 2027, and the company is said to be moving to an annual release model. The company also explains that the current Ignition 8.3 will be supported until 2030 as a long-term support (LTS) version. All of this comes from vendor announcements, and the Enterprise Historian Module is at the preview stage. No availability date or price has been given.

Note that Tiger Data is the company that changed its name from Timescale on 17 June 2025. The open-source time series database extension (an extension to PostgreSQL) continues to be called “TimescaleDB”. Take care not to confuse the company name with the product name. The two companies also announced a collaboration at Hannover Messe 2026 on 20 April 2026, and that announcement cited “closed, per-tag licensing” and “proprietary data formats” as challenges in the market. This is a claim by the parties involved, but it shows that how licences are counted and how easily data can be extracted have become issues in historian selection.

What this means for factories in Thailand

These announcements are not in themselves a reason to replace anything right away. However, it is worth noting the shift: historian options are no longer a binary choice between “a dedicated product or the history feature bundled with SCADA”, and architectures that use a general-purpose time series database underneath are becoming an official option even from industrial software vendors.

For factories about to introduce a historian, or to replace an old history system, decisions such as “do we implement on the current long-term support version or wait for the next one” and “are we locking our data into a particular product’s format” can be considered more concrete than before. The foundation for those decisions is the 4 decisions covered in this article.

What Is an Industrial Historian: How It Differs from SCADA Logs and General-Purpose Time Series Databases

The NIST definition, and its place within the control system

The OT security guide “SP 800-82 Rev.3” (September 2023) from the US National Institute of Standards and Technology (NIST) defines a data historian in its glossary as “a centralized database supporting data analysis using statistical process control (SPC) techniques”. The same document gives examples of control centre configurations in which an HMI, engineering workstations and a data historian are connected over a LAN, and in which PLC data is stored in the historian.

Its uses are not limited to SPC. The same document also states, as something specific to OT networks, that a data historian can be a “supplemental” source of event data that helps fill in the picture of a cyber incident. However, it is “supplemental”; this does not mean it can replace a security logging platform. It is best understood as a record for tracing what happened in the factory along a timeline, from quality investigations and analysis of equipment abnormalities to energy analysis and post-incident reviews.

How it differs from SCADA logs

SCADA and HMI systems also have history functions for trend displays. In this article, we consider the following 3 separately.

AspectSCADA/HMI history functionGeneral-purpose time series databaseDedicated industrial historian
Main purposeTrend display on operator screens, review of the recent pastFast writing and aggregation of time series data, used with your own designLong-term storage of process data across the whole factory, used for quality, equipment and energy analysis
CollectionOften limited to the tags that SCADA holdsCollection mechanism must be built separately (gateways, MQTT, custom integration, etc.)Often has collection functions for various control devices (interfaces, adapters, connectors)
Data handlingConfigured per product; retention is sometimes set relatively shortCompression and retention policy depend on your designThe product provides deadband, compression, quality flags, handling of out-of-order data, etc.
Suited toOperational monitoring of 1 line is the main goal and analysis requirements are smallThe IT department has experience running time series databases and can design collection and operation in-houseYou want to keep data from multiple lines and systems for the long term, independently of operations

The table summarises general tendencies as organised by this article, and products differ greatly. What matters is not which is superior, but “where your own requirements lie”. The approach to selecting and upgrading SCADA itself is explained in “SCADA Selection for Thailand Factories: RFP, FAT/SAT and TCO“.

“Can be stored” is not the same as “can be replayed”

The value of a historian is not that it can accumulate large amounts of data. It lies in being able to replay later, when a quality problem is found six months on, which tags had which values at that time on that day, and whether those values were trustworthy (not a sensor fault or a communication loss). For that, not only the values but also accurate timestamps, quality information (good, missing or substitute value) and the meaning of each tag (which instrument on which equipment, and in what unit) need to be kept together with the values. This is what this article means by “trusted, with context, and replayable”.

Historian Selection Comparison: 3 Options and How to Read Product Trends

The first decision: do you really need a historian?

This article considers that if 2 or more of the following situations apply, it is highly likely that the SCADA history function alone is no longer enough.

  • You want to compare data from multiple lines and multiple SCADA systems or PLCs on the same timeline
  • You need to keep process data for several years for quality investigations
  • You do not want to put heavy analytical queries on the SCADA used for operations
  • People outside the control network, such as head office or the quality assurance department, want to see the data
  • You want to fill gaps in the data afterwards even if communication is lost

Conversely, if the goal is operational monitoring of 1 line and a short retention period is acceptable, reviewing the SCADA history settings first may be quicker.

Using a general-purpose time series database

The options for general-purpose time series databases are widening. As an example, InfluxData announced the general availability of “InfluxDB 3 Core” and “InfluxDB 3 Enterprise” on 15 April 2025. According to the company, Core is open source under MIT/Apache 2 licences and is positioned as a “fast recent-data engine for real-time use cases”. Enterprise is said to add high availability, read replicas, automatic failover and more. The fact that Core is described as being “for recent data” is a point to check when considering whether to use it as a long-term storage historian.

TimescaleDB (developed by Tiger Data), an extension to PostgreSQL, is also one of the representative options among general-purpose time series databases. As mentioned above, Inductive Automation has announced that it will incorporate it as one of its own historian options.

If you choose a general-purpose database, you need to design and operate collection, quality flags, out-of-order data, filling of missing data, compression, retention and permission management yourselves. Whether the IT department has the experience and manpower for this can be considered a major dividing line in selection.

Trends in industrial software and dedicated historian products

Historian functions are also being reworked in industrial software products. Inductive Automation released Ignition 8.3 on 16 September 2025 and announced that it had assembled an “Industrial Historian Solution Suite” consisting of the Historian Core Module, the SQL Historian Module, and a Historian API for implementing your own historian. According to the company’s documentation, the Core Historian is a database with QuestDB embedded, with partitioning, deduplication, archiving and aggregation functions, and the handling of old data can be chosen from “None (keep indefinitely)”, “Prune (delete old partitions)” and “Archive (move to another location)”. Memory allocation is said to default to 10% of the system’s total memory. QuestDB’s documentation explains that this replaced the previous SQLite-based internal historian.

As an example of a dedicated historian, Canary Labs explains on its official website that its storage method uses “lossless compression” and that store-and-forward is a dedicated component that delivers data reliably even while offline. AVEVA’s PI System is a product with exception (deadband) and swinging door compression, described later. AVEVA is said to have published improvements such as better handling of out-of-order data in PI Server 2024 R2.

These are examples of options on the market, and there are many other products. This article makes no recommendation or ranking here. When comparing, what you should look at is less the performance multiples or tag limit figures, and more whether that product and configuration can meet “what to record, how, and how to accept it”, covered in the following sections.

Checking standard interfaces

When selecting a product, one thing worth checking is the standard for extracting historical data. The OPC Foundation’s “OPC 10000-11 UA Part 11: Historical Access” is a specification that defines how data is stored in and retrieved from historians and databases. It covers the concepts of data history and event history, the behaviour for creating, retrieving, updating and deleting archived data and annotations, and security including access rights and auditing. The latest version is V1.05.04 (released 29 November 2024). For handling aggregates such as interpolation and time averages, there is a specification called “Part 13: Aggregates”, and version 1.05.07 (published 15 April 2026) is available on the OPC Foundation reference site.

Even if a product says “OPC UA compatible”, it may support only reading and writing current values and not historical retrieval (Part 11). In the RFP, confirm individually which specifications are supported for historical retrieval and aggregation.

What to Record and How: Tag Register, Sampling, Timestamps and Store-and-Forward

Industrial Historian Selection: RFP and FAT/SAT in Thailand - figure 1

The tag register is the real design document

The first deliverable to create in a historian implementation is not a server configuration diagram but the tag register. The register should hold at least the following information.

  • Tag name, and which equipment and which instrument it belongs to (equipment hierarchy)
  • Unit, measurement range, instrument accuracy
  • Collection interval (scan interval)
  • Deadband and compression settings, and the rationale for them
  • Retention period
  • Meaning of quality flags
  • Description (in Japanese, English and Thai)

If tag names remain as PLC addresses, a few years later nobody will understand what they mean. How you name the equipment and tag hierarchy is also related to topic design when sending data over MQTT. The approach is covered in “MQTT Sparkplug B implementation in Thailand: RFP, FAT/SAT and a 90-day pilot“. Also, if the list of devices the tags come from is out of date, you cannot build the register. For taking inventory of control devices, “OT Asset Management Implementation: Inventory, RFP and 90-Day Acceptance for Thai Factories” is a useful reference.

Decide the sampling interval from “what the data is used for”

Collecting every tag at the shortest interval may seem safe, but it only increases capacity and communication load, and piles up data that is never used. Decide the interval by what the data will be used for. A slowly changing value such as temperature needs a different interval from a value where you want to capture sudden pressure changes. For digital states (running/stopped, alarms and so on), recording only when they change can be considered more natural in many cases than collecting them at a fixed interval.

Keep timestamps and quality together with the values

To make data replayable later, it must first be correct about “when the value is from”. The meaning changes depending on whether the timestamp is applied at the PLC or sensor, at the gateway, or when the data arrives at the historian. If delayed data is given its arrival time, cause and effect can appear to be in the wrong order. We believe it is important to unify the time source within the factory and to write into the specification where each timestamp is applied. If you store in Coordinated Universal Time (UTC) and convert to Thai local time for display, there is less confusion when comparing data with overseas sites or head office.

Keep the quality of each value as well. If sensor wire breaks, communication losses, manually entered substitute values and the like are treated the same as good values, the conclusions of your analysis will be wrong.

Prevent missing data with store-and-forward

Factory networks do go down. There are many reasons, such as equipment failure, construction work, and the order in which things come back after a power cut. Store-and-forward is a mechanism in which the collection side (an edge gateway or collection server) temporarily holds data it could not send and, once communication returns, sends it oldest first. Without it, data from the period when communication was down is lost permanently.

What needs care here is how late-arriving data is handled. According to an AVEVA presentation, in the PI System “out of order data bypasses compression”. Products differ in whether late-arriving data can be inserted at the correct point in time, and in what happens to compression and aggregation when it is. How to estimate the buffer capacity is covered in the calculation later.

The Key to Long-Term Process Data Storage: Decide Deadband and Compression from Instrument Accuracy

Industrial Historian Selection: RFP and FAT/SAT in Thailand - figure 2

Exception and compression: the example of AVEVA PI System

In long-term storage, the most debated topic is the settings that thin out data. Here we take up AVEVA’s PI System as an example whose method is explained in public material (based on a presentation at AVEVA World San Francisco in 2023. The terms are specific to the PI System and cannot be generalised to other companies’ historians).

The PI System has a 2-stage mechanism. The first, “Exception”, is a simple deadband for removing noise. It runs only on the collection side in PI Interfaces, and reduces the number of events sent to save network bandwidth. The settings are the deadband width in engineering units (ExDev), the percentage of span (ExDevPercent), and the maximum number of seconds between events (ExMax). The same material notes that AVEVA Adapters can perform deadband filtering but ignore tag exception settings, and that PI Connectors do not perform exception testing. In other words, even within the same product, thinning behaviour differs depending on the collection method. This is a point to confirm in the RFP.

The second, “Compression”, is an algorithm called swinging door, which reduces the amount of data stored by discarding points that can be safely reproduced by interpolation. The material quotes its purpose as “remove instrument and process noise while still recording significant process changes”.

Both “not configuring” and “over-configuring” cause problems

The same material contains easily overlooked cautions: the default values are not zero, and “turning compression off” is not the same as “turning compression on with a deviation of zero”. If you start operating with the default values at implementation, unintended thinning may be taking place.

As an example of the effect, for a state tag taking only 0 and 1, with no exception and a 1-second scan, the number of events over 48 hours is given as 172,800 with no compression and 15 with minimal compression. 172,800 is every point over 48 hours x 3,600 seconds. However, this is an example under a specific condition, a digital state tag taking only 0 and 1, and this ratio cannot be applied to analogue values.

On the other hand, in a case from an offshore oil and gas customer in the US, more than 87% of tags reportedly used neither exception nor compression, and the onshore server was periodically unstable. The tags were classified, new settings were decided based on instrument accuracy and the views of experienced site staff, and when these were tested on a pilot rig, stored events were reported to have fallen by about 94%. This too is the result of a single customer’s pilot, not an effect for factories in general. What you should take from this is not the reduction figure, but that “there are problems both with continuing to store every point and with over-configuring and losing changes”.

Decide based on instrument accuracy, and verify the difference from the raw data

So what should the settings be based on? In a case from a US refinery in the same material, using conservative exception and compression settings reportedly resulted in a maximum difference from the original data below the instrument’s measurement accuracy for all tags analysed. What this case shows is the idea that “there is no point storing changes finer than the instrument can measure; conversely, thinning that creates differences beyond the instrument’s accuracy is corrupting the data”.

This article recommends the following procedure.

  1. Classify tags by instrument type and use (quality investigation, control analysis, energy, etc.)
  2. Write the instrument accuracy into the tag register, and use it as the basis for provisional deadband and compression widths
  3. On test tags, store the full raw data and the thinned data side by side
  4. Calculate the maximum difference from the original data when the thinned data is interpolated, and check that it falls within the instrument accuracy
  5. For important quality-related tags, check visually on trends, together with experienced site staff, that no sudden changes have been lost
  6. If there are no problems, roll out to tags of the same class

Also note that methods differ by product. Some products advertise “lossless compression” that does not thin data, while others discard points that can be reproduced by interpolation. The question is not which is better; what matters is to confirm in acceptance testing whether the chosen product’s method and settings can preserve the accuracy you want to store.

Estimating Time Series Database Capacity with a Factory Model (Hypothetical Model)

All figures from here on are assumptions set independently by this article. They are not product performance figures or industry averages. Use them as a calculation template and replace them with your own tag counts and intervals.

Assumptions (Model Factory H)

ItemAssumption (hypothetical value)
FactoryJapanese-owned resin/chemical factory in eastern Thailand
Tag count2,000 tags (1,500 analogue, 500 digital)
Collection interval1 second for all tags, to keep the comparison simple
Storage size per point16 bytes (including timestamp, value and quality flag; assumed value before compression)
1 year365 days
Unit1 GB = 1 billion bytes

Raw data volume

The number of points 1 tag produces per day is 24 hours x 3,600 seconds = 86,400 points (the 172,800 points in the AVEVA example above correspond to 2 days of this). For 2,000 tags, that is 2,000 x 86,400 = 172,800,000 points/day, and for 1 year, 172,800,000 x 365 = 63,072,000,000 points/year.

In bytes, 1 day is 172,800,000 points x 16 bytes = 2,764,800,000 bytes, or about 2.76 GB; 1 year is 63,072,000,000 points x 16 bytes = 1,009,152,000,000 bytes, or about 1,009 GB; and 10 years is about 10,092 GB.

Differences by retention rate

How much deadband and compression reduce the data varies greatly with the nature of the signal and the settings. So rather than predicting a reduction rate, we set provisional retention rates for comparison to get a sense of the volume.

Retention rate (assumed)Storage for 1 yearStorage for 10 years
100% (all points stored)about 1,009 GBabout 10,092 GB
50%about 505 GBabout 5,046 GB
25%about 252 GBabout 2,523 GB
10%about 101 GBabout 1,009 GB

What this table shows is that, even just across the provisionally assumed range of retention rates (10 to 100%), the required capacity for the same factory differs by a factor of 10. 1 year of storing every point equals 10 years at a 10% retention rate. We believe capacity estimates should be recalculated not from product catalogue values but from retention rates measured on test tags. Capacity for backups and replicas is also needed separately.

Estimating the store-and-forward buffer

Next is the buffer for when communication stops. Assume that a network device failed during a long holiday and nobody noticed for 72 hours. The number of points accumulated in that time is 2,000 tags x 3,600 seconds x 72 hours = 518,400,000 points, or in bytes 518,400,000 x 16 = 8,294,400,000 bytes, about 8.29 GB. Allowing a 2x margin, the edge side needs a buffer of about 16.6 GB or more.

Also estimate the time needed to send all the accumulated data once communication is restored. Assuming the edge can send 10,000 points per second, and since 2,000 new points per second are still being generated, the buffer shrinks at 8,000 points per second. 518,400,000 points / 8,000 points = 64,800 seconds, or 18 hours. These 18 hours are a period in which the data is still not filled in on the historian. If the sending capacity or the historian’s write capacity is insufficient, it will take even longer.

This calculation shows what needs to be verified in acceptance testing: “how many hours of disconnection it can withstand”, “how many hours it takes to catch up after recovery”, and “whether, in the meantime, data arriving out of order is placed at the correct time without breaking compression or aggregation”.

Where to Place the Historian: OT, DMZ and Enterprise Network

Industrial Historian Selection: RFP and FAT/SAT in Thailand - figure 3

The NIST principle: enterprise-to-OT communication goes through services in the DMZ

A historian is also a mechanism for people outside the control network to use control system data. Its placement is therefore directly tied to security design.

The security architecture example in NIST SP 800-82r3 places a DMZ to separate the OT environment from the enterprise network, and states that all communication between the enterprise level and the operations management level (including Level 3 of the Purdue model) goes through services within the DMZ. The field level is described as including devices at Levels 0 to 2 of the Purdue model. It further states that because the DMZ connects to the outside, services within the DMZ must be monitored and protected so that attackers cannot get into the OT environment undetected.

This article’s design example: a primary historian on the OT side and an enterprise-wide replica in the DMZ

From this principle, this article considers the following configuration as a design example. Note that this is this article’s own reasoning derived from the NIST principle; NIST does not specify historian placement in this form.

  • Primary historian: placed on the OT side (operations management level), collecting data from PLCs and SCADA. The master record for operations and quality investigations
  • Enterprise-wide replica: placed in the DMZ, receiving data from the primary historian. Users on the enterprise network side, such as head office, quality assurance and analytics teams, access only this replica
  • Enterprise network: BI, ERP and similar systems take data from the replica in the DMZ. They do not connect directly to the OT-side historian

This can be considered to make it less likely that heavy queries from the enterprise side affect the primary historian that supports operations, and to reduce the number of entry points into the OT side.

The weakness hidden in “direct automatic replication”

Note that the replication mechanism itself can become a path for attacks or errors. NIST SP 800-82r3 lists “improper data linking” as an example of a vulnerability, stating that database links that automatically replicate data historian data to other databases can, if not properly configured, become a vulnerability that allows unauthorised access or tampering.

We believe that, however convenient it may be, a design that links directly from IT-side BI or ERP databases to the OT-side historian should be avoided, and the direction of replication, the communication used, authentication and monitoring should be written into the design document. If you want to physically guarantee one-way flow from OT to IT only, a data diode is also an option. The approach is laid out in “Industrial Data Diode Selection for Thai Factories: A Practical OT-to-IT Guide“.

Protecting the historian server itself

Historian servers often run on general-purpose Windows or Linux. NIST SP 800-82r3 states that general-purpose systems used for engineering workstations, data historians, maintenance laptops, backup servers and the like can generally be protected in the same way as IT equipment, for example by distributing updates from an antivirus server placed inside the control network (with the qualifier “generally”; the same treatment cannot be extended to control devices such as PLCs). Decide at implementation who will handle OS updates, account management and backups, and by what procedure.

Data Integrity and Audit Trails in Regulated Industries

In factories handling electronic records subject to regulation by the US Food and Drug Administration (FDA), such as pharmaceuticals, medical devices and food for the US market, additional requirements may apply to how historian data is handled. US federal regulation 21 CFR §11.10 requires persons who use closed systems to create, modify, maintain or transmit electronic records to have procedures and controls that ensure the authenticity and integrity of electronic records. As one of these, (e) lists the use of secure, computer-generated, time-stamped audit trails to independently record the date and time of operator entries and actions that create, modify or delete electronic records.

This is a US FDA regulation and is limited to cases involving records subject to FDA regulation. It is not a legal obligation for factories in Thailand in general. However, in factories that are in scope, whether a record remains of who changed what, and when, when historian data is corrected afterwards (manual corrections, filling of missing data, adding annotations, etc.) becomes a product selection requirement. The OPC UA Part 11 mentioned above also covers security including access rights and auditing. Whether you are in scope, and what is required, should be confirmed individually with your quality assurance department or specialists.

12 Items to Write in a Historian RFP

These are the items you should at minimum write into a historian RFP (request for proposal) when requesting quotations from vendors or system integrators.

  1. Purpose and users: what it will be used for, such as quality investigations, equipment analysis, energy and post-incident reviews, and who will look at it
  2. Tag register: tag count, data types, units, instrument accuracy, equipment hierarchy, languages for descriptions
  3. Collection methods and target devices: types of PLC, SCADA and instruments, communication methods. For each collection method, whether deadband and exception settings take effect
  4. Interval, deadband and compression method and how they are configured: the method (thinning or lossless), default values, how the rationale for settings is recorded, permissions and history for changes
  5. Timestamps and quality: where timestamps are applied, the time source, the time basis for storage (UTC or local time), types of quality flags
  6. Store-and-forward: buffer capacity, length of disconnection it can withstand, sending capacity after recovery, handling of out-of-order data
  7. Retention period and capacity: retention period, handling of old data (deletion or archiving), basis for capacity estimates
  8. Placement and network: configuration of the OT-side primary historian and DMZ replica, direction and method of replication, access paths from outside
  9. Data extraction: support for OPC UA historical access (Part 11) and aggregates (Part 13), and how data is output to other systems via SQL, API, CSV and so on. Whether it is locked into a proprietary format
  10. How licences are counted and future versions: whether counted by tags, connections or servers, how expansion is handled, length of long-term support and migration to the next version
  11. Audit trail and security: records of data corrections, accounts and permissions, procedures for OS updates and backups (and the relevant requirements if regulated)
  12. Local support and FAT/SAT criteria: maintenance contact within Thailand, explanations in Thai, response time in the event of failures, test items and acceptance criteria, scope of witnessing

Items 3, 4 and 6 are especially important. If you do not write item 3, you will only find out after implementation that thinning behaviour differs by collection method even under the same product name. If you do not write item 4, unintended thinning will continue with the default values. If you do not write item 6, nobody can take responsibility when data from a disconnection period disappears. Item 10 also determines whether you will be able to take your data out in 5 or 10 years, as product versions and configurations keep changing, as in the ICC 2026 announcements. How to position the historian within the way manufacturing data is collected across the company is laid out in “Manufacturing Data Collection: Thailand 90-Day PoC and RFP“.

What to Check at Historian FAT/SAT

In factory acceptance testing before shipment (FAT) and site acceptance testing after installation (SAT), verify under your own conditions, rather than relying on catalogue values, that “data is not lost, is not corrupted, has correct timestamps, can be retrieved when needed, and can be restored if lost”.

Test itemMethodHow to set the acceptance criterionMain stage
Disconnection and backfilling of missing dataDeliberately stop communication between the collection side and the historian, and check whether data is filled in with the correct timestamps after resumptionWithstands the disconnection time set in the RFP with no gaps. Record the time from recovery until it catches upFAT/SAT
Out-of-order dataCheck that late-arriving data is placed in the correct position and that aggregation and compression are not brokenThe replayed trend matches the original dataFAT
Compression and deadband errorOn test tags, place full-point data and stored data side by side and calculate the maximum difference from interpolated valuesThe maximum difference falls within the instrument accuracy in the tag registerFAT/SAT
Retention of sudden changesOn quality-related tags, generate short peaks or steps and check that they remain after storageExperienced site staff check the trends and confirm no important changes have been lostSAT
Timestamp correctnessCompare device-side time with stored time. Also check behaviour when the time source is switchedWithin the tolerance set in the RFP. Where timestamps are applied matches the specificationFAT/SAT
Quality flagsCause a sensor wire break or communication loss and check that quality is recorded correctlyAbnormal values are not stored as good valuesFAT/SAT
Query performanceRun queries over long periods and many tags with data volumes close to productionMeets the response time users need for their work, against criteria set in advanceSAT
Replica and replicationCompare the contents of the primary historian and the DMZ replica, and check the direction and path of replicationContents match, and the path from the enterprise side to the OT side is closed as designedSAT
Backup and restoreActually restore from backup to a different server and check that data and tag settings come backCan be restored to the defined point in time within the defined timeSAT
Recovery after a power cutCut power and switch it back on, and check that collection and store-and-forward recover automaticallyRecovers without manual work, and gaps are filled from the bufferSAT

Numerical acceptance criteria differ depending on the factory’s use, so decide them at the RFP stage. “Restore” in particular is an item for which a plan is often written but never actually tested. Since historian data cannot be collected again later, we believe it is well worth actually confirming that it can be restored.

Issues Specific to Thailand and ASEAN

BOI efficiency measures

In the “Measure for Industrial Upgrades towards Smart and Sustainable Industry” in the Thailand Board of Investment (BOI) investment promotion guide (2023 edition), incentives for improving the efficiency of existing businesses (regardless of whether they are BOI-promoted) are set out, on condition of an investment of THB 1 million or more (excluding land and working capital). The measure for adopting digital technology is given as a 3-year corporate income tax exemption (capped at 50% of the investment), and conversion to Industry 4.0 under a plan approved by NSTDA as a 3-year corporate income tax exemption (capped at 100% of the investment). Requirements for adopting digital technology include introducing software or information systems that link data across at least 3 functions inside and outside the organisation, or using AI, machine learning, big data, data analytics and the like, and in some cases investment in software developed or modified by a provider certified in Thailand is required. Some industries are excluded.

However, this is material from the 2023 edition, and we have not been able to confirm whether it is open for applications as of October 2026 or whether the conditions have changed. Whether a historian implementation could be eligible should be checked against the latest application guidelines and confirmed individually with the BOI or specialists.

Thailand’s Cybersecurity Act

According to commentary from the law firm Tilleke & Gibbins (1 August 2025), Thailand’s Cybersecurity Act (B.E. 2562, 2019) currently applies to state agencies, supervisory and regulatory agencies, and designated critical information infrastructure (CII) organisations. The CII sectors are national security, essential public services, banking and finance, IT, telecommunications, transport and logistics, energy and public utilities, and public health; manufacturing is not included in the current sectors. CII organisations must make an initial incident report within 24 hours. A draft amendment published in July 2025 also proposed adding “industrial work”, but it is at the draft stage, and this article has not been able to confirm its subsequent enactment.

This law does not impose direct obligations on typical Japanese-owned manufacturing factories. However, there may be an impact if you have energy or public utility facilities, or depending on future amendments. Whether it applies should be confirmed individually with specialists. Since historian data can also serve as a record that helps fill in the picture during an incident, it can be considered one of the factors when deciding on retention periods and timestamp accuracy.

Points to watch in local operations

From here on is this article’s own view. In factories in Thailand, it is not uncommon for operators and maintenance staff to look at data in Thai, Japanese expatriates in Japanese, and head office or overseas sites in English. Whether tag descriptions can be held in 3 languages, and whether the language of screens and reports can be switched, determines whether the historian will actually be used. Other points worth checking before implementation are whether there is a maintenance setup that can respond locally in case of failures, and whether the recovery procedure after a power cut is written in Thai.

If you are also thinking of connecting it with data from multiple factories or from systems other than the historian (MES, quality, maintenance), there is also the approach of positioning the historian as “a component responsible for part of the data”. See also “Industrial Data Fabric for Manufacturing: RFP and Acceptance Guide“.

90-Day Plan for Industrial Historian Implementation

Days 0-30: Clarify the purpose and take inventory for the tag register

Decide what the historian will be used for, such as quality investigations, equipment analysis and energy, and identify the users. Then list the target equipment and tags, and write the units, instrument accuracy, required intervals and retention periods into the register. At this stage, also make a provisional first decision: whether the SCADA history is enough, whether to build with a general-purpose time series database, or whether to introduce a dedicated historian.

Days 31-60: Measure on test tags

Select several dozen representative tags, and store the full raw data side by side with data to which deadband and compression have been applied. Check with experienced site staff whether the maximum difference from the original data stays within the instrument accuracy and whether important changes are lost. At the same time, deliberately stop communication and test whether store-and-forward fills the gaps and whether timestamps are correct. From the retention rate obtained here, rebuild your own version of the capacity estimate.

Days 61-90: RFP, FAT/SAT criteria and ordering decision

Based on the measured results, finalise the 12 RFP items and the FAT/SAT test items and acceptance criteria, and obtain quotations from multiple companies under the same conditions. Agree the placement (OT side and DMZ replica) and network design with the factory IT and information security staff. The choice of product version (implement on the current long-term support version or wait for the next one) is also decided here, after checking the support periods each company has published.

The deliverables you should have in hand at the end of the 90 days are these 6: (1) a definition of purposes and users, (2) the tag register (with the rationale for accuracy, intervals, retention periods and compression), (3) records of compression error and disconnection tests on test tags, (4) your own version of the capacity estimate, (5) the placement and network design, and (6) the RFP and FAT/SAT criteria.

Frequently Asked Questions

What is an industrial historian, and how does it differ from SCADA data storage?

An industrial historian is a mechanism for storing a factory’s process data along a timeline over the long term and using it for analysis. NIST SP 800-82r3 defines it as a centralised database that supports statistical process control. The SCADA history function is mainly for trend display on operator screens, and may be limited to that SCADA’s tags or have a relatively short retention period. If you want to keep data from multiple lines and systems for the long term, independently of operations, a historian becomes a candidate.

How do you choose between a historian and a general-purpose time series database?

A general-purpose time series database offers a high degree of freedom, but you need to design and operate collection, quality flags, out-of-order data, compression, retention and permission management in-house. A dedicated historian is characterised by often having these functions as part of the product. Decide based on factors such as whether the IT department has the experience and manpower to run a time series database, and whether you want to avoid locking your data into a particular product’s format. Recently, industrial software vendors have also announced configurations that use a general-purpose time series database underneath.

For long-term process data storage, how should compression and deadband be decided?

Decide them based on instrument accuracy. On test tags, place full-point data and thinned data side by side, and check that the maximum difference between interpolated values and the original data falls within the instrument accuracy and that no important changes are lost, before rolling out. A product’s default values may not be zero, and settings may not take effect depending on the collection method, so avoid starting with the default values.

Where on the network should the historian be placed?

This article’s design example places a primary historian on the OT side and an enterprise-wide replica in the DMZ, with users on the enterprise network accessing only the replica. This is reasoning derived from the NIST SP 800-82r3 principle that communication between the enterprise and OT goes through services in the DMZ; NIST does not specify this placement. NIST also cites database links that automatically replicate historian data to other databases as an example of something that can become a vulnerability if not properly configured.

What should be checked in a historian RFP and at FAT/SAT?

In the RFP, write down the tag register, the thinning behaviour of each collection method, the compression method and default values, timestamps and quality, store-and-forward capacity and out-of-order data, retention period, placement, how data is extracted, how licences are counted and long-term support, audit trails, and local support. At FAT/SAT, verify backfilling of missing data after disconnection, compression error, timestamp correctness, quality flags, query performance, replica consistency, restore from backup, and recovery after a power cut.

How much does a historian cost?

Cost varies greatly with how licences are counted (tags, connections, servers, etc.), storage capacity and server configuration, the types and number of target devices for collection, whether there is a DMZ replica, integration with existing systems, and the scope of acceptance testing. This article does not estimate amounts because there is no primary information on prices, but capacity can be estimated from tag count x interval x retention rate, as in the hypothetical model in the main text. The reliable approach is to first create the tag register, measure the retention rate on test tags, and then compare quotations. Whether you could be eligible for BOI promotion measures should be confirmed individually with the BOI or specialists.

Summary

  • The value of an industrial historian is not the volume stored, but being able to keep “time series data that can be trusted, carries context and can be replayed later”.
  • The first decision is whether SCADA history is enough, whether to build with a general-purpose time series database, or whether to introduce a dedicated historian. Product comparison comes after that.
  • The real design document is the tag register. Decide intervals from how the data is used, decide where timestamps are applied, and keep quality flags together with values. Prevent gaps during disconnections with store-and-forward.
  • Decide deadband and compression based on instrument accuracy, and roll out only after verifying the difference from the original data on test tags. Do not start with the default values.
  • In the hypothetical model, raw data from 2,000 tags at a 1-second interval is about 1,009 GB per year. With provisional retention rates of 10 to 100%, the capacity differs by a factor of 10. Also estimate the buffer for a 72-hour disconnection and the time needed to catch up after recovery.
  • For placement, a design example of an OT-side primary historian and a DMZ replica can be considered. Avoid direct automatic replication.
  • At acceptance, verify backfilling of missing data, compression error, timestamps, query performance and restore under your own conditions.

TOMAS TECH supports Japanese-owned factories in Thailand, from preparatory stages such as clarifying the purposes of a historian, taking inventory for the tag register, checking compression error on test tags and estimating capacity, through to drafting the RFP, designing network placement, witnessing FAT/SAT and integrating with existing SCADA and production management systems. Even if you are at the stage of “not yet ready to choose a product, but wanting to know where to start”, please feel free to get in touch via our contact form.

References