The ring reconverged in 50ms. The process failed anyway. The network did its job. The control system did not recover.


Industrial Ethernet switch with amber flashing LED and label reading Ring – Primary

One Night. Three Substations. One Breaker That Stayed Open.

A power event causes a voltage sag across a transmission corridor. Three substations lose communication. The ring network reconverges in 50ms. The SCADA shows everything is normal. But one breaker stays open. The line is dead. The cause is never found.

The network recovered. The devices did not. The transient that caused the failure disappeared. The evidence was gone. The investigation closed with no root cause. This scenario plays out more often than the industry acknowledges. Engineers accept "unknown transient" as an explanation because the evidence vanishes before anyone can capture it.

In a utility environment, a delayed breaker restoration may affect protection coordination across an entire corridor. In manufacturing, a brief communications interruption may force a production restart that takes hours to recover. The network event lasts milliseconds. The operational consequence may last the rest of the shift.

What Everyone Assumed

The engineers assumed the network was the problem. The network recovered. The problem must have been a transient. The transient passed. The system returned to normal.

The assumption was that recovery meant resolution. If the network returned to normal, the problem was solved. The assumption is common. It is also wrong. The network recovered. The problem did not. The transient that caused the failure is still there. It is just hidden.

Diagram showing vendor convergence time vs operational traffic restoration gap

The gap where failures live.

What Actually Happened

Vendor specifications measure convergence from the moment a link fails to the moment the last switch learns the new path. That is not the same as traffic restoration. During convergence, multicast traffic may pause. Unicast frames may loop. GOOSE messages may be delayed or dropped.

The ring is stable – but the control system does not know that. Convergence time is a vendor metric. It is not an operational metric. The control system cares about when traffic resumes, not when the switch finishes learning.

A switch that reconverges in 50ms sounds safe. But a protection relay expecting GOOSE delivery within 4ms will see that 50ms window as a communication loss. A drive expecting a position update every 10ms will fault. A SCADA master polling over the ring will time out. The network recovers. The devices do not.

The Consequence Layer

The gap between vendor metrics and operational reality has real consequences. A 50ms convergence delay may not cause an alarm. It may not trigger a fault. But it can push a protection calculation out of tolerance, delay a drive update, or time out a SCADA poll. The network appears stable. The process does not.

Engineers who measure only network recovery miss the operational recovery that matters. The network recovered. The problem did not.

Why Recovery Masks Root Cause

When the network recovers quickly, the investigation stops. The operator sees the network return to normal. The log shows no persistent fault. The root cause is never found. The transient that caused the failure disappears. The evidence is gone.

Root cause masking is the result of recovery that is too fast. The network recovers before anyone can see what happened. The engineer arrives after the evidence has vanished. The investigation records the event as "unknown." The fault returns. The cycle repeats. Each repetition costs time, money, and trust.

What the Vendor Metric Doesn't Show

The recovery time specified on the data sheet is measured in a controlled environment. Single failure. Clean traffic. No other events. The field is different. A restart storm – hundreds of devices reappearing simultaneously – can extend convergence by an order of magnitude. The same ring that recovers in 50ms during commissioning takes 500ms when the entire substation powers back on.

Lab versus field is the difference between what the vendor promises and what the environment delivers. The lab is predictable. The field is not. The test that passed in commissioning does not reflect what happens during an actual event. The gap between lab and field is where failures hide.

A Pattern That Repeats Across Industries

The same pattern repeats across industries. A manufacturer installs a ring network. The vendor demonstrates 50ms recovery. The system passes acceptance. Six months later, a drive faults during a restart. The logs show nothing. The network recovered in 50ms. The drive did not. The cause is never found. The drive faults again.

In a substation, a protective relay trips intermittently. The network recovers in 50ms. The relay logs show a communication loss. The network logs show normal operation. The cause is attributed to "network glitch." The relay trips again. In a factory, a robot loses communication. The network recovers in 50ms. The robot faults. The cause is never found. The robot faults again.

The pattern is the same. The equipment changes. The cause does not.

Where Westermo Fits

Capturing transient behaviour requires more than basic network monitoring. Engineers need timestamped events, persistent logging and visibility that survives power cycles. Industrial networking platforms such as Westermo are designed to preserve this information, helping operators investigate what happened before the evidence disappears.

Westermo's FRNT (Fast Recovery Network Topology) is tested under full load, not idle conditions. The recovery time specified is what you get when the network is busy – not when it is empty. That difference matters when your network is actually running.

What You Can Do Now

Test recovery behaviour before the next restart tests it for you. Simulate simultaneous power loss across multiple cabinets, not isolated single-device failures. Real events do not occur one device at a time.

Measure network latency during a scheduled restart. Compare it to steady-state latency. The difference is what your control system experiences but your monitoring never shows.

Configure your switches to retain event logs across power cycles. Test what happens during a restart. If the logs disappear, you have a visibility gap.

Measure operational recovery, not just network recovery. Measure how long the process takes to recover, not how long the ring takes to converge. The difference between those two numbers is often where hidden risk exists.

Monitor recovery behaviour continuously. A one-time test is not enough. Behaviour changes as the network ages. The recovery time that was 50ms last year may be 200ms this year. You will not know unless you monitor it.

A Final Thought

The network recovered. The problem did not. The transient that caused the failure is still there. It is just hidden. Engineers who design for predictable recovery design infrastructure that remains visible, manageable, and stable under stress. The sites that monitor recovery behaviour find the root cause. The sites that assume recovery means resolution will chase the same fault again.


DESIGN FOR PREDICTABLE RECOVERY

Throughput advises on network architectures that recover predictably, preserve forensic evidence, and reveal root cause.

Where is your network most unstable after restart?

Fill out the online form.

You May Also Be Interested In ...