Reading a switch's error counters before anybody complains
A link that is failing does not go down. It goes slow, intermittently, in a way that gets blamed on the application for months.
24 July 20262 min readadmin
A cable that has been damaged rarely fails cleanly. Clean failure is easy — the link light goes out, somebody notices, the cable gets replaced. What actually happens is that the link stays up and starts corrupting a small fraction of frames, and everything above it deals with the loss by retransmitting.
The result is a network that works, with occasional slowness that nobody can reproduce. It gets attributed to the file server, then the internet connection, then "the system". The switch has been reporting the real answer the whole time.
The counters worth looking at
CRC errors and FCS errors. A frame arrived and its checksum did not match. This is the single most useful counter on a switch. It means physical-layer damage: a bad cable, a bad patch lead, a connector that was never properly terminated, or interference along the run. It is almost never the device at the end.
Input errors and runts. Frames shorter than the minimum. Often the same causes as CRC errors, sometimes a duplex mismatch.
Late collisions. On any modern switched network this should be zero forever. A non-zero count is a duplex mismatch — one end forced to full, the other auto-negotiating down to half — and it produces exactly the intermittent slowness described above.
Output drops. The switch had traffic to send and nowhere to queue it. Not a fault; a capacity signal. Consistent output drops on an uplink mean that uplink is the bottleneck.
Rates, not totals
A switch that has been up for two years will have accumulated errors, and the total tells you almost nothing — some of them may be from a cable that was replaced eighteen months ago.
What matters is whether the number is increasing. Clear the counters, wait a day, look again. A port with a rising CRC count has a live physical problem. A port with a large total and no movement had one, once.
The comparison that finds it fastest
Errors on one port and not on its neighbours points at that cable run. Errors across every port on one switch points at the switch, its power, or its earth. Errors on both ends of the same link point at the link itself.
That triangulation takes about two minutes and resolves most "the network is slow" reports faster than any packet capture.
Do it before the complaint
All of this is worth reading on a schedule rather than during an outage. A monthly look at error rates across the access switches turns a class of problem that presents as a mystery into a maintenance task with a cable at the end of it.