Blog

2026.10.04

MQTT Sparkplug B implementation in Thailand: RFP, FAT/SAT and a 90-day pilot

MQTT Sparkplug B implementation in Thailand: RFP, FAT/SAT and a 90-day pilot

When a Thai factory asks vendors about MQTT Sparkplug B implementation, many can say that their products support MQTT or Sparkplug. Those labels alone do not show whether a PLC value will reach a MES or HMI with the right meaning and state. Procurement needs evidence of what is sent when equipment starts, stops, loses its connection, recovers, or changes configuration. This guide turns that question into an RFP, factory acceptance test (FAT), site acceptance test (SAT), and illustrative 90-day pilot for buyers, OT and IT teams, and production engineering.

Start with acceptance criteria, not a product name

Sparkplug defines a topic namespace, payload and session-state framework for MQTT clients. MQTT itself transports messages; it does not by itself standardize all the equipment meaning and state that an industrial subscriber needs. The Eclipse Foundation specification page dates Sparkplug 3.0 to October 2022. The 3.0.0 specification is the version reference used here. Do not present it as a new 2026 release. Name the applicable version in the contract.

Separate four decisions: the behavior required by the specification; a product’s compatibility claim and tested version; the factory’s integration design; and the evidence required at FAT and SAT. A product can satisfy the first two while a local time source, tag convention, firewall, or MES mapping still causes a poor implementation. Acceptance must be based on the actual boundary conditions.

This article addresses procurement and acceptance of Sparkplug when the underlying collection design is known. For source tags and acquisition methods, see PLC data collection design. For choosing the gateway itself, see industrial IoT gateway selection. Here the focus is the data contract, session state, and interoperability when several applications consume those readings.

Define the machines, signals and consumers on one page

Begin the RFP with a bounded example: two assembly lines, three machines with different PLC vendors, and run/stop, good-count, reject-count, product-code and alarm signals going to an HMI and a MES. These numbers are illustrative, not a standard scale or recommended target. State what is outside the pilot. If control commands and regulated records are excluded, write that explicitly rather than allowing them to be inferred from a demo.

For every signal, record its source, business meaning, unit, type, normal range, update trigger, event-time origin, handling of missing values and owner. The equipment team owns the meaning of source values; the network team owns the permitted routes; the application team owns storage and presentation; procurement ties deliverables to payment. A run bit that is reset on a PLC restart or a counter that wraps cannot be interpreted safely without those decisions.

For multi-site operations, agree on plant code, stable machine ID, display label and time notation. A Thai-language HMI can coexist with stable technical identifiers. Changing a machine name after historical data has accumulated affects joins and access rules, so record the naming convention and approval process before expanding the pilot.

Translate Sparkplug roles into purchasing scope

A typical architecture has a PLC or sensor, an MQTT Edge Node, an MQTT Server and Sparkplug Host Applications or other consumers. Specification roles do not always map one-to-one to boxes and licenses. Ask each bidder to show which product and version fills each role, who runs it, and who controls failover. A broker alone does not necessarily validate the business meaning of every Metric. Review the behavior of the whole chain.

Sparkplug defines node and device Birth, Data and Death message types and a topic namespace. Birth is more than the first reading: it supplies the model a subscriber needs for the session. A subscriber seeing a Death or state change should not continue presenting the last value as if it were live. Specify the effect on HMI status, alarms and MES processing. The specification describes message behavior; the factory must decide how a stopped machine differs from an unreachable one.

Even when all roles come in one product, ask about license expiry, certificate renewal and network changes. Which screen detects the failure, and who restores service? A “Sparkplug supported” checkbox does not remove single points of failure or define operations.

MQTT Sparkplug B implementation in Thailand: RFP, FAT/SAT and a 90-day pilot - figure 1

Keep compatibility claims, official listing and site acceptance separate

Eclipse maintains a Compatible Products list. Verify the exact product and version when a bidder cites it. Absence from the list alone does not prove a product is incompatible: the official Get Listed process includes TCK testing, membership, agreements and a listing request. Conversely, a listing does not certify the factory’s tag meanings, network, credentials or procedures.

Give the bid matrix separate columns for the claim, supporting evidence, product version and site test. Ask for a link to the official entry and relevant test details if a bidder uses the “Sparkplug Compatible” mark. Being able to speak MQTT v5 is not proof of implementing Sparkplug. A listed broker does not make an arbitrary Edge Node and Host combination automatically acceptable.

The Eclipse TCK process explains compatibility evaluation and states that the specification is the controlling source if a test conflicts with it. Treat TCK evidence as product evidence, not as a substitute for SAT. Include revalidation rules for major product changes in the contract.

Twelve RFP lines that permit real comparison

For each requirement, request an answer, demonstration method, pass criterion, submitted artifact and support owner. “Available” or “not available” is too coarse. This table is a purchasing template, not an endorsement of any product.

RequirementBidder responseAcceptance evidence
Scope and rolesSpecification version; Edge, Server and Host products and versionsArchitecture and configuration export
Data modelGroup, node, device and Metric names and typesActual Birth messages and dictionary
Birth/DeathBehavior on start, orderly stop and lost connectionTrace, screen recording and timeline
ReconnectionState recovery after network loss and restartFault injection and logs
Quality and timeSource/receipt time, invalid values and sync failureCorrelated logs and HMI result
Change controlAdding, deleting and renaming a tagChange record, diff and rollback
SecurityAuthentication, authorization, encryption, key renewalSettings and permission tests
MonitoringOffline, delay, stopped input and queue pressureAlerts and operator screen
ResilienceFailover behavior for each roleSwitch-over test record
RetentionLocal limit, discard and replay policyCapacity estimate and reconciliation
InteroperabilityDifferent vendors’ Edge and Host combinationTest configuration and constraints
HandoverBackup, recovery, logs and support contactsAs-built documents and training

Set numerical thresholds from the intended process. If an RFP says “within five seconds,” define the start and end timestamps and whether the criterion covers normal load, burst load or recovery. Five seconds is merely an example, not performance promised by Sparkplug. Compare candidates with representative types, rates, concurrent connections and fault conditions, not only maximum tag counts in a brochure.

Make the Metric dictionary a contractual deliverable

Give each Metric a row containing machine ID, technical name, business label, type, unit, source, update trigger, time origin, quality rule, HMI field, MES destination and approver. Define unknown, disconnected, out-of-range and maintenance states. A Boolean zero may mean “stopped” or “sensor fault” depending on the equipment. The protocol cannot make that business decision.

Separate an immutable identifier from a human display label so that a machine move or line renaming does not break historical joins. Use one technical identifier even if the explanatory labels have several languages. Version the dictionary and publishing configuration together so an analyst can tell which definition applied to older records.

Name the owner who signs off the dictionary. An equipment vendor may define the raw register, while production must approve its use in output or quality reports. Require a portable copy of the dictionary and configuration at SAT. Settings visible only in a proprietary screen make later migration and auditing harder.

Use FAT to reproduce state transitions

In a vendor test environment, start the Edge Node and compare the Birth model against the dictionary. Change input values and reconcile Data messages with HMI and stored records. Disconnect the network, check that consumers recognize an unavailable source, and examine whether the last reading is marked as old. Reconnect and verify that a coherent state is rebuilt without registering the same machine twice. This is a test-design example; validate exact protocol obligations against the specified version.

If possible, repeat the same sequence with components from different suppliers. Record whether an official compatibility claim exists separately from whether the factory’s required scenario works. If a second implementation is unavailable, mark the untested interoperability risk in the purchasing decision rather than treating one successful demo as universal proof.

The FAT evidence bundle should include architecture, versions, configuration export, message traces, action times, test input, expected and observed result, discrepancy, correction and retest. Keep timestamps in a consistent zone and retain machine-readable logs alongside video. Record how test credentials are replaced before production. Name the customer’s acceptance authority so the vendor’s own pass statement is not the only evidence.

MQTT Sparkplug B implementation in Thailand: RFP, FAT/SAT and a 90-day pilot - figure 2

Use SAT to test Thai factory constraints

Site testing adds real network segments, firewall rules, DNS, time sync, power, maintenance workstations and shift operations. A lab configuration can fail on site because a route is blocked, many nodes reconnect at once after maintenance, or the actual PLC tag differs from the sample. Write down the machine stop permit, affected production area, emergency contact and rollback before a site fault test.

Protect safety and production. Start with a test path and simulated signals. Fault injection affecting a live controller needs the factory’s production and maintenance approval. Determine disconnection duration and replay volume from actual acquisition rate and capacity. Sparkplug does not impose a universal number of hours of offline buffering. If an essential machine cannot be interrupted, perform part of SAT in isolation and record the remaining site risk rather than calling the test complete.

Let the shift supervisor and planner examine how the HMI distinguishes “old value,” “disconnected” and “recovering.” A correct message can still be misread by an operator. Test who receives an alarm at night, what the next shift sees, and what the incident ticket records. Acceptance includes that operational behavior.

Distinguish MQTT transport from Sparkplug semantics

The OASIS MQTT Version 5.0 standard specifies connections, publish/subscribe, QoS and sessions. Sparkplug defines industrial data structure and state above this transport. “QoS 1” alone is not a promise that a business event will be stored exactly once. Retries, application processing and database failures can create duplicate business records; the receiving system needs suitable keys and idempotent processing. Put each guarantee at the correct layer in the RFP.

Likewise, using MQTT retain or a Last Will is not automatically equivalent to handling Sparkplug Birth/Death correctly. Test what makes a consumer change status from valid to unknown. With multiple consumers, restart only one and check that it can obtain the model it needs. A demo with a single uninterrupted HMI does not expose stale state in a second application.

For broker failover, test whether the subscriber can rebuild its model. Specify target recovery time, tolerable loss and duplication, reauthentication and alarm-clearing rules. These are factory service objectives, not values implied by the word Sparkplug.

Review OT security through the network diagram and procedures

Sparkplug adoption does not complete OT security design. NIST SP 800-82 Rev.3 provides OT security guidance. Apply its principles to the local inventory, permitted routes, identity, change control, monitoring and recovery. This is guidance, not a claim about a Thai legal obligation; check contractual and legal requirements separately.

Ask how Edge Nodes, Server and Hosts receive separate accounts; which topics each may publish or subscribe to; who issues, renews and revokes certificates; and how secrets in backup files are protected. Management-console access and message-path permissions are different controls. Remote support needs an approved purpose, period, record and deactivation process.

Encrypted transport does not correct a wrong Metric name or an authorized bad change. Require peer review, pre-deployment validation, rollback, audit logs and a certificate-renewal rehearsal in the pilot deliverables. On contract termination, ensure factory-controlled configuration and identity management remain available.

Price the responsibility boundary, not just the box

There is no universal MQTT Sparkplug B implementation price. The scope depends on whether existing PLCs can be read, whether an existing device can act as an Edge Node, who operates the broker, whether the Host handles state, and how much OT review or downtime coordination is needed. Request separate lines for hardware, licenses, communications, dictionary work, integration, HMI/MES change, FAT, SAT, training, support and future machine additions. A single “turnkey” amount hides omissions.

Separate initial and operating cost. Check subscription and support terms, certificate renewal, log retention, monitoring, backup, incidents, regression tests and plant expansion. Clarify Thai local support, night/weekend contact, language and factory access. Remote troubleshooting does not replace safety and network approvals.

Rather than promise a savings percentage, measure the present manual work, unexplained stops, missing data, recovery time and report preparation using a method that can be repeated after the pilot. ROI and payback are site-specific. The business case should show both measurable outcomes and the cost of sustaining them.

An illustrative 90-day pilot with exit gates

Ninety days is a planning example, not a standards requirement. Extend it for machine stop windows or procurement lead times. Agree on machines, signals, consumers, owners, baseline, safety limits, pass criteria and withdrawal criteria before work starts.

Illustrative periodWorkExit evidence
Days 1–15Survey machines, signals and network; draft dictionaryScope, architecture, risk register
Days 16–30RFP replies, comparison and purchase scopeResponse matrix, versions, test plan
Days 31–50Vendor setup and FATTraces, deviations, correction plan
Days 51–70Site installation, SAT and operator trainingSite records and recovery procedure
Days 71–90Stable observation and expansion reviewAcceptance decision and investment case

Use gates, not dates alone. Do not start FAT with an ownerless dictionary; do not inject live faults without a factory permit; do not expand a line while serious missing-data issues remain. State who pays for retests and what is delivered at each milestone. Observe start-up, product changeover, weekend shutdown and post-maintenance recovery where they matter. An event that did not occur in 90 days remains an untested risk unless safely simulated.

MQTT Sparkplug B implementation in Thailand: RFP, FAT/SAT and a 90-day pilot - figure 3

Write measurable FAT/SAT criteria carefully

Every threshold needs a measurement point, a load condition and an exception rule. For latency, distinguish PLC change, Edge acquisition, broker arrival and Host persistence. For recovery, record disconnect, offline indication, reconnect and backlog catch-up. A measurement based on unsynchronized clocks is not credible.

Define the denominator of any missing-data rate. A change-only Metric does not produce one expected message per second. Reconcile injected events with send and receive logs. Test duplicate arrival separately from duplicate counting in the business database. Monitoring and quality applications may need different acceptance levels; this article does not invent one percentage for every factory.

“Not measurable” is not a pass. If traces cannot establish delay or loss, improve observability and retest. A changed threshold requires a written reason, business impact, price effect and customer approval.

Plan for change, scale and supplier exit

Adding machines introduces duplicate Metric names, different units and older equipment semantics. Make registration, duplicate-ID detection, dictionary review, permissions, capacity check and subscriber regression testing standard work. A meaning change should be versioned, not silently overwritten. Avoid reusing retired machine IDs.

For software upgrades, record the supported Sparkplug version and the exact combination of tested products. Check official listings for the relevant product version, then rerun the FAT/SAT scenarios affected by the change. Rollback includes dictionary and certificate alignment, not merely restoring a settings file.

At contract exit, arrange export of dictionaries, architecture, configuration, history, account deactivation and license terms. Standards adoption does not automatically make migration effortless. Portable artifacts and a clear handover obligation are purchasing requirements.

The purchasing meeting’s final checklist

Confirm ten items on one sheet: scope and consumers; specification version; product roles and versions; evidence for compatibility claims; dictionary owner; Birth/Death fault tests; OT security and certificate renewal; FAT/SAT pass criteria; pilot exit; and support, change and supplier-exit terms. If several remain “to be decided,” price alone cannot settle the bid.

Normalize quotations to the same tested scope. One bid may cover only an Edge Node while another includes broker and Host configuration. Reconcile the architecture, include factory staff time, and identify gaps before comparing totals. A lower price may still be the right choice when responsibilities and tests are explicit.

FAQ: questions before implementing MQTT Sparkplug B

Does an MQTT gateway mean Sparkplug B is implemented?

No. MQTT connectivity is different from implementing Sparkplug topic, payload and state behavior in the required roles. Verify the product version and test the actual data and fault scenarios.

Can SAT be skipped when a product appears on the Compatible Products list?

No. Listing concerns product compatibility. Site-specific signal meaning, network paths, time, certificates, HMI display and operations need separate acceptance. Check the exact version and combination.

What determines MQTT Sparkplug B implementation cost?

Machine scope, existing device capability, Edge/Server/Host purchases, dictionary creation, HMI/MES work, OT review, FAT/SAT, training and support. Compare the same scope and test conditions rather than relying on a universal figure.

Is 90 days enough to approve a full plant rollout?

It is an example pilot period. Observe the relevant operating changes, simulate safe fault cases, document remaining risks and then decide the next stage. Shutdown windows and security review may change the schedule.

Summary

MQTT Sparkplug B implementation is a procurement and acceptance exercise as well as a protocol choice. Connect the chosen specification version, product roles, Metric dictionary, session state, compatibility evidence and OT operations to reproducible FAT and SAT tests. The pilot’s exit criteria matter more than the number of days. Clear evidence and ownership make expansion beyond the first Thai production line easier to judge.

If you are still shaping the RFP or FAT/SAT scope, you can contact TOMAS TECH with a sketch of the equipment and data consumers. We can help identify the comparison and acceptance questions that fit your current constraints.

Sources