When a factory in Thailand plans industrial network redundancy, “we will make it a ring” is not an adequate purchase specification. The team must decide which failures are in scope—cable, switch, power, or controller—and how many missed control cycles a process can tolerate while remaining safe. This guide compares MRP, RSTP, PRP and HSR, then turns recovery, PTP, QoS and application behavior into RFP and FAT/SAT evidence.
Start with the process impact, not the protocol name
Redundancy describes several different outcomes. A remote I/O link may need to survive a cable cut without changing the PLC state. A MES collector may tolerate a short interruption if it buffers and resends records. An HMI may go dark while local control remains safe. Those cases need different limits, owners and tests. For each production step, write down what a loss of communication does to personnel safety, product quality, throughput and delivery. Do not assign one universal recovery time to a line whose PLC can continue locally and another whose remote I/O must enter a safe state immediately.
The requirement table should identify the failure point, affected traffic, interruption as seen by the application, tolerated missing frames or control cycles, detection method, recovery action and acceptance owner. “The ring changed path” is a network observation. “The conveyor kept making conforming product” is a process result. A demonstration of the first does not prove the second. Machine safety functions need their own safety engineering and validation.
If the installed network lacks an asset inventory or a defined OT/IT boundary, begin with our industrial network construction guide. The focus here is how to procure and prove fault tolerance for the paths that matter.
Translate “single failure” into a failure list
A single failure might mean one cable in a ring, an entire switch, one power source, or a shared cable tray damaged in one incident. Two logical paths through the same cabinet and breaker remain exposed to a common cause. The RFP should list device, port, transceiver, cable, tray, supply, UPS, time source and management dependencies by failure domain. Color the duplicated and single components on the as-built drawing and test the relevant boundaries.
Redundancy also consumes budget, capacity, engineering time, monitoring effort and spares. A factory need not duplicate every flow. Protect a control loop with high downtime cost, while using buffering and replay for non-control history. Make that distinction before selecting a topology.
MRP, RSTP, PRP and HSR at a glance
| Method | Typical topology | Behavior after a single covered fault | Prerequisites | Main caution |
|---|---|---|---|---|
| MRP | Ring of supported PROFINET devices and switches | A path change with a possible interruption | MRP support and configuration at every ring participant | Compare interruption with the control watchdog |
| RSTP | General Layer 2 switch mesh or ring | Interruption while the topology reconverges | Bridge interoperability and root design | Convergence varies with layout, implementation and fault |
| PRP | Two independent LANs, A and B | Duplicate frames let the receiver use the surviving LAN | PRP endpoints or RedBox and genuinely separate LANs | Cost, supervision and common-cause faults |
| HSR | Ring of compatible HSR nodes | Duplicate frames travel in both directions | Compatible endpoints and a defined interconnection boundary | Bandwidth and interoperability |
PROFIBUS & PROFINET International (PI) describes MRP as non-seamless media redundancy. It says ordinary PROFINET real-time traffic can use MRP, while demanding isochronous real-time applications may need MRPD as an additional mechanism. PI’s system description gives a typical MRP reconfiguration time below 200 ms. That is not a guarantee that a particular factory application will be interrupted for less than 200 ms. Confirm support, ring roles and ports, node limits and firmware for the delivered devices, then measure the installed configuration. MRPD, controller system redundancy and dual power inputs are separate concepts.
IEC 62439-3:2021 defines PRP and HSR for seamless switchover with zero recovery time for a covered network-element failure. PRP attaches nodes to two separate LANs; HSR carries duplicated frames around a ring or suitable mesh. This standard property does not mean zero production downtime after PLC CPU failure, shared power loss, a common cable route, software error or multiple faults. Do not convert “zero network recovery time” into “the factory can never stop.”
RSTP is a broadly interoperable Layer 2 choice, but the presence of compliant switches alone cannot promise a plant-specific convergence time. Root bridge, edge ports, topology changes and the difference between unplugging a cable and losing a switch must all be examined. Cisco’s industrial design guide gives illustrative convergence comparisons; a procurement specification should instead state the measured application limits for the actual site.

MRP for a PROFINET ring
In a supported ring, MRP keeps one route logically blocked to avoid a loop and uses an alternate path when the ring breaks. Draw the manager, clients, configured ring ports, non-ring attachments and alarm destination. During replacement, restore the saved ring configuration and check for loops and a second fault before returning the plant to normal operation.
Expect some interruption during MRP path change. Compare the PLC I/O period, watchdog, drive and application timeout on paper, then recreate representative tags and traffic in FAT. A catalog “typical” figure should not be treated as unconditional approval for servo synchronization or a safety function. Check the exact firmware and role of every delivered ring participant.
RSTP for mixed upstream networks
RSTP is useful when existing switches from several vendors serve an aggregation or supervisory network. Fix the root bridge, priorities, permitted attachments and port roles in a controlled drawing. Evaluate traffic with application buffering separately from tight control loops. A large bridge domain makes path behavior harder to diagnose.
Do more than pull one cable during testing. Consider switch power loss, flapping ports, reconnection after maintenance and an unexpected root candidate. Create potentially dangerous loops only in an isolated FAT lab or within an approved SAT plan. A restored physical path may still leave MAC learning, multicast membership, PLC sessions or application retries incomplete.
PRP when two LANs can remain independent
A PRP source transmits duplicated information over LAN A and LAN B; the receiver discards the duplicate. If one LAN fails, the other copy can arrive without a waiting period for path reconvergence. This requires compatible endpoints or properly designed RedBoxes for singly attached devices, plus genuine separation and health monitoring. A hidden fault on one LAN must be repaired before the second fault removes the remaining path.
Two VLANs on the same switch stack are not two independent failure domains for switch loss. Review cabinets, breakers, cable routes, maintenance practice and simultaneous firmware changes. Include the second cable route, optics, switches, spares, configuration backup and fault drills in the cost comparison.
HSR for a compatible device ring
IEC 62439-3 HSR sends duplicated frames in both directions around the ring. It is not a permission to mix any ring protocol or endpoint. Design the nodes, ports, connection to other networks, multicast and capacity as one system. Measure normal and one-side-failed loading.
Some vendor literature also uses “HSR” for a different “High Speed Redundancy” feature. Specify IEC 62439-3:2021 High-availability Seamless Redundancy in the RFP and demand a product-specific compatibility statement. An acronym alone is insufficient for acceptance.
Recovery time is only one layer of availability
Record four timestamps: when the physical link failed, when the redundant network path became usable, when the PLC or SCADA again received valid data, and when production and quality behavior returned to normal. A fast path change can still be followed by TCP reconnects, PLC timeouts, HMI refresh, MES replay or clock resynchronization. Calling all of these “switchover time” prevents meaningful bid comparison.
Choose metrics per application. For remote I/O, count missed update cycles and PLC state changes. For vision inspection, count lost images and quarantined products. For data collection, check missing sequence numbers and replay completion. For genealogy, verify ordering, duplicates and timestamps. For SCADA, test alarm history and operator visibility. Derive the limits from the process risk assessment instead of copying a vendor default.
PRP and HSR duplicate frames, so capacity planning and the duplicate-elimination boundary matter. MRP and RSTP can also shift normal traffic onto a constrained path after a fault. Model normal and worst credible single-fault load, including video, backups, multicast and control bursts. Measure queues and port utilization, not merely nominal line speed.
Design PTP and QoS alongside redundancy
If a process depends on precise timestamps, uninterrupted data delivery is not enough when the clock becomes unstable. For PTP or IEEE 802.1AS, document the grandmaster, boundary or transparent clocks, delay asymmetry on each route, failover behavior and holdover. IEEE 802.1AS addresses synchronization across normal operation and changes or failures in network components. Test time offset and resynchronization after a fault, not just the arrival of packets. Our factory PTP time synchronization guide explains the clock design and acceptance measures in more detail.
QoS is not a checkbox that automatically saves control traffic. Classification, trust boundaries, 802.1p or DSCP mapping, queues, shaping, policing and multicast treatment must match at every hop. IEEE 802.1Qbv scheduled traffic uses timing derived from 802.1AS. If you deploy that function, verify the product and profile’s behavior when timing degrades. A “TSN capable” label without named profile, endpoint support, configuration method and test cases is not an engineering specification.

Example acceptance evidence for a one-side failure
The following is a template, not a universal numerical target. Fill the limit for each application from its process analysis and use separate rows for normal and fault operation.
| Observation | Evidence to retain | How to set acceptance |
|---|---|---|
| Clock offset and resynchronization | PTP logs and comparison with reference time | Set limits from the application’s timestamp tolerance |
| Control latency and jitter | Send/receive timestamps, switch queues and PLC diagnostics | Meet the application requirement under load and after failure |
| Missing or duplicate frames | Sequence numbers and PLC/SCADA events | No unacceptable control or quality effect |
| Detour-path utilization | Port counters and packet capture | Within engineered peak and maintenance margin |
| Alarm and diagnosis | Monitoring alert and service ticket | Fault and accountable owner can be identified |
A green “synchronized” indicator alone is weak evidence; inspect the time-stamped quality records for discontinuity or reversed order. Conversely, do not force expensive seamless redundancy on a data stream that can safely buffer and replay a short interruption.
Physical paths, power and devices need separate decisions
Two logical lines that merge into the same cable tray can disappear together in a forklift accident or a localized fire. For an operating Thai plant, survey tray space, building-to-building routes, safe access, dust, heat, vibration, cabinet capacity and permitted shutdown windows. A/B routes on a drawing must match physical labels and the asset register.
Two power terminals on a switch do not create source redundancy when both connect to the same breaker. Verify independent supplies, UPS hold time and monitoring, breaker separation and replacement procedure. Even a dual-homed PLC can still have a single controller CPU, I/O module, actuator, time source or upstream server. Maintain separate drawings and test IDs for network, power, controller and application redundancy.
Serviceability is part of availability. During a planned update of one side, confirm that the surviving path handles the agreed load. A simultaneous erroneous update of both sides defeats separated cables. Put change approval, configuration diff, backup, rollback and spare replacement into the deliverables.
Five questions for choosing a method in a Thai factory
First, are you protecting a control period, continuous production, supervisory visibility or data integrity? Second, which industrial protocol and redundancy feature does each actual endpoint support? Third, is the scope one cable or also a switch, power source and shared tray? Fourth, can the plant fund and operate truly separate paths? Fifth, can the actual application be tested in FAT and SAT? Answering these questions narrows the choices before a vendor recommends a box.
For an existing PROFINET ring tolerating a short interruption, MRP is a natural candidate. Where such interruption is unacceptable, assess MRPD or another controller/application design as applicable. For mixed upstream monitoring traffic with buffering, RSTP may fit. Where frames must survive a covered single network-element fault without a reconvergence gap and two independent LANs are feasible, assess PRP. For compatible nodes and a controlled ring/interconnection boundary, assess HSR. None of them duplicates the PLC or power supply by itself.
If methods meet at a gateway or RedBox, check whether that boundary becomes a single point of failure. “Redundancy supported” may refer only to ports, a protocol, a supply, a CPU or a licensed feature. Demand an interoperability matrix for delivered part numbers and firmware, then test those exact versions.
Nine clauses to put in the RFP
- Protected process: Identify lines, PLCs, HMIs, remote I/O, SCADA, MES and maintenance devices, plus the safe behavior during communications loss.
- Failure model: Separate link cut, switch stop, power loss, transceiver fault, planned side maintenance and restoration; state how multiple and common-cause faults are handled.
- Application criteria: Specify measurement points and limits for cycles, watchdogs, alarms, quality records, replay and clock offset, not packet delivery alone.
- Method and compatibility: State standard edition, supported devices, firmware, configuration, gateways and ring or dual-LAN limits.
- Physical and power separation: Match A/B routes, cabinets, switches, breakers, UPS, grounding, labeling and access to the site survey.
- PTP/QoS: State clock source and profile, priorities, queues, failure behavior, test load and monitoring.
- Security: Specify OT/IT boundaries, allowed flows, management access, protected configurations, logs and change rights. NIST SP 800-82 Rev.3 is a useful primary guide for security that respects OT reliability and safety.
- FAT/SAT: Contract for fault-injection steps, representative load, timestamped evidence, pass/fail, retests, shutdown windows and rollback.
- Handover: Include as-built drawings, configurations, IP/VLAN/port registers, spares, alarms, restore runbook, training and support escalation.
Compare bids by equipment, second-route construction, panels and power, licenses, FAT, night SAT, shutdown coordination, training and maintenance. A low bid without fault injection has postponed the cost of proving availability. A generic price range would be misleading without a plant survey and a defined outage window.
Questions that produce useful vendor answers
Ask, “When we power off this LAN A switch, how many I/O cycles are missed at this PLC, and which logs prove it?” Ask, “After loss of the PTP grandmaster, when does this quality terminal return to its specified offset?” Ask where the two paths share a tray and which configuration backup a technician uses to replace a failed unit. Put the answers into the test specification rather than leaving them as meeting notes.
FAT and SAT acceptance sequence
FAT need not reproduce the whole plant, but it should use the specified hardware and firmware, actual switch settings, a representative PLC, traffic load and monitoring. First record normal operation; then inject one link cut, switch power loss, reconnection, one-side maintenance, clock-source loss and queue congestion. Align network counters and application events on one time axis. If a setting changes, record its version and repeat the same condition.
SAT must respect production and safety. Use an approved shutdown window, isolation procedure and contact tree. Do not create an uncontrolled loop or second fault on the live network. For failures that cannot safely be injected on site, combine FAT evidence with a physical site inspection. SAT adds real cable routes, supplies, endpoints, uplinks and operating load. Production, maintenance, quality, IT/OT and vendor owners should sign the same acceptance record.

| Test ID | Injection or check | Network evidence | Process evidence |
|---|---|---|---|
| F-01 | Single ring or LAN A link cut | Port event, route, missing/duplicate frames | PLC state, I/O cycles, product disposition |
| F-02 | Switch or one-side supply loss | Alarm, convergence, surviving-path load | HMI, alarm history, operation |
| F-03 | PTP source or time path loss | Offset, resynchronization, clock role | Timestamped quality records |
| F-04 | Peak load on the detour path | Queues, latency, drops | Control, inspection and MES completion |
| F-05 | Restoration and spare replacement | Configuration, ports and monitoring return | Technician demonstration and open issues |
If SAT fails, assign an owner, safe interim operation, repair deadline and retest. The network supplier saying “our switches passed” while production says “the line stopped” must be resolved with the four timestamps and shared contract criteria. After acceptance, verify alarms reach an owner and a hidden one-side fault does not remain open indefinitely.
Turn a one-side failure into a visible maintenance job
Redundancy helps when the first fault is absorbed during production and the maintenance team learns about it promptly. A plant may run for weeks on its last remaining path if a one-side failure produces no actionable alarm. A later cable job can then remove the final path. Convert the technical event “link down” into a maintenance record that names the affected line, failed side, owner and repair deadline. Define the notification destination, overnight contact, ticket, escalation and verification of restoration—not merely an NMS trap.
An operations dashboard should distinguish healthy, one-side failed, redundancy lost, repair in progress and multiple-fault states. For MRP, observe ring-open and manager status; for RSTP, root and port roles; for PRP, the receive condition of LAN A/B and duplicate frames; for HSR, each ring direction. Counter names vary by product, so hand over a model-specific mapping of “if this counter rises, inspect this component.” Verify in SAT that the alarm reaches the actual duty technician.
Spare inventory needs more than a quantity. Tie each replacement switch and optical module to its part number, firmware, license, backed-up configuration, port map, IP assignment, replacement steps and end-of-life substitute. A switch in the storeroom may be unusable if its firmware cannot join the ring, the password is missing or its configuration belongs to another cabinet. In F-05, have the local maintenance team perform replacement from the runbook and revise any unclear step.
“Update one side at a time” must be an executable procedure. Before isolating one side, confirm the survivor carries the required control and supervisory load. Approve the configuration diff, save a baseline, update one side, check alarms and normal traffic, then proceed to the second side. Confirm the intended pair of firmware and configuration versions, including behavior if device roles change on restart. Define the version to roll back to, the responsible engineer and the point at which the plant stops the update.
Compare total cost and untested scope in bids
Comparing PRP dual LANs, HSR devices, an MRP ring and an RSTP backbone by hardware price alone conceals cabling and test scope. Use separate rows for initial devices, copper/fiber route construction, cabinets, power and UPS, design/configuration, FAT/SAT, production shutdown coordination, monitoring, maintenance contracts, spares and future expansion. Accept a claim that existing fiber can be reused only with evidence of optical loss, actual route and available strands. Request an explicit exclusions list: clock synchronization, security configuration, machine-maker attendance, weekend tests and rollback may otherwise fall between contracts.
If bids use a five-year evaluation period, align assumptions for device refresh, licenses, UPS batteries, spares, monitoring and retesting. A business case is not just “years of downtime saved.” Product quarantine, work-in-process recovery, night maintenance and customer delivery effects also matter. Do not invent monetary values where the plant lacks evidence. Use past incident records to separate measurable duration and cost from qualitative residual risk. Document risks that remain after redundancy and connect them to operator drills and a restoration plan.
Frequently asked questions about industrial network redundancy
Is MRP faster than RSTP?
There is no universal ranking for a real application. PI provides a typical MRP reconfiguration figure; RSTP varies with topology and implementation. Compare the interruption and process effect measured on the delivered site, including the relevant failure type. Also ask whether the use case is a PROFINET ring or a mixed upstream network.
Do PRP or HSR guarantee that production never stops?
No. IEC 62439-3 addresses seamless handling of covered network-element failures. It does not protect every controller, power, software or common-route fault. Validate endpoints and process behavior in FAT/SAT.
Does a redundant ring automatically keep PTP in sync?
No. Path delay, clock role, port functions and the selected profile matter. Measure offset and resynchronization in normal and failure states, and treat time-source redundancy as a separate requirement.
How should bids be compared?
Use the same failure list, application limits, route independence, compatible devices, PTP/QoS design, FAT/SAT scope, spares and restoration drill. Switch count and catalog recovery time alone do not define equivalent bids.
Can an operating brownfield factory retrofit this design?
Often, but first verify drawings, shutdown windows, cable and cabinet space, supply capacity and old-device support. Prove the configuration in a test environment, migrate by line or building and retain a rollback route. Avoid improvised fault injection on a live production network.
Conclusion: accept the process result after a fault
MRP and RSTP require an application to tolerate measured interruption during path change. PRP and HSR aim for seamless frame delivery through covered single network-element faults. None eliminates every physical, electrical, controller, clock, QoS, operational or common-cause single point. A useful Thai-factory RFP starts from the failure model and process outcome, then measures the same criteria in FAT and SAT.
Even before choosing a protocol or part number, a team can inventory the current routes, identify production steps that cannot stop and draft an acceptance table. If you are planning an OT network upgrade in Thailand, you can discuss those constraints with TOMAS TECH.
Primary sources
- PROFIBUS & PROFINET International, PROFINET Planning Redundancy Guideline.
- PROFIBUS & PROFINET International, PROFINET System Description.
- IEC, IEC 62439-3:2021.
- IEEE 802.1, TSN standards and projects.
- IEEE 802.1, 802.1Qbv project description.
- Cisco, Industrial Automation Design and Implementation Guide.
- NIST, SP 800-82 Rev.3.