Insights

Guides

Intermittent network faults: the live path, not another cable certificate

On this note

2026-08-19 · 2 min

Passive certification answers a different question: did this permanent link meet a category at a moment in time. Intermittent faults live in the active path: autonegotiation that reseats, PoE budget that sags when a camera heater starts, a GBIC that overheats after noon, a duplex mismatch that only appears when a particular host wakes, CRC that tracks a VFD two trays away. Active network diagnostics is that investigation.

How the investigation is actually run

Leave counters running with timestamps that operations can match to a shift log. A five-minute laptop visit will not catch a once-per-shift bounce. If you must disturb the path, do it at the end, after you have the error story. Check both ends of the hop: a clean switch port next to a dirty NIC is a different conclusion from two dirty ends, which often means the cable or the tray.

Power over Ethernet deserves its own clamp or PD measurement when cameras and access points share the complaint. Optics deserve temperature. Copper that was certified and then recoiled onto a drum deserves a field test under the failing heat and humidity, not a lecture about last year’s PDF.

When you write, name the hop in the same language the patch panel uses. A “network problem” that cannot be walked to is not finished. If the correlation is electrical, hand the file across; do not keep it as an IT mystery out of habit.

Symptoms

Interfaces that bounce. Spanning-tree topology changes. Hosts that “need a reboot.” OT timeouts at shift change. Wireless blamed because it is visible, while the wired uplink is the thing incrementing errors. See loss, latency, jitter for how those errors present to the process.

Causes

Physical: damaged pairs, water in a tray, a connector on strain, a patch that was certified then moved. Electrical: SNR collapse under noise — read when the network problem is electrical. Logical: speed/duplex, MTU, a storm from a loop, a controller that retransmits until the switch CPU dies. Power: PoE class vs actual draw, a midspan that browns out.

Investigation

Log interface counters with timestamps against process and electrical events. Capture errors on both ends of the hop. Check optics temperature and voltage where they exist. Do not recertify the whole floor first. If the copper is a suspect, a targeted field test under the failing condition beats a stack of old PDFs. Coordinate with OT/IT diagnostics when the “network” is a cell that IT does not own.

The conclusion names the hop, the counter, and the correlated condition. Replacing a switch because it is the only object in the ticket is not a mechanism.

Services

Related project

Related Insights