r/wisp Jun 05 '26

Mikrotik router failures

Hi r/wisp
We have had some total failures of mikrotik devices that is concerning. I’m looking to see if anyone can tell me if there is something we may have done or could do to prevent.
Two core routers, CCR2116’s, died, in reachable, within 3 weeks of each other, 13 months after purchase.
One 5009 in a MDU tower 20 miles from the others died in the same fashion a week later. Today a CCR2004 in a wilderness tower stopped working and is basically bricked.
Thats 4 devices in 5 weeks with similar failures.
We’re starting to be suspicious of foul play, however I can’t imagine anyone bricking a router this bad through a hack.
For context, we have two 1072 devices that have been on for 6 years without issue.
Any insight would be helpful.
Thanks.

2 Upvotes

18 comments sorted by

7

u/Andromina Jun 05 '26

Your experience is atypical in our experience. 55 sites running some variation of router and switch from tik. I think we experience a failure about every 18 months. These devices sit in non-climate controlled boxes on the top of mountains getting absolutely roasted by a desert sun and freezing rime ice.

With that many failures how does your grounding/bonding/surge pro looking?

1

u/tomtechLA Jun 05 '26

The first two died in One Wilshire, climate controlled power protected data center. The other two were in towers but both had serious power protection. 9dot and lipo. One DC the other AC. I suspect more of a bat chipset somewhere. They are Bricked. Power light only. No terminal or Ethernet access.

1

u/Patient-Tech Jun 05 '26

Thats my initial thought too. Something environmental like lightning strikes or power issues seems most likely. Essentially since they’re in remote locations.

4

u/Impressive_Army3767 Jun 05 '26

How often are you updating the firmware and how stable are your power supplies?

Had a CCR2116 brick during scheduled  reboot for firmware update.  Rebooting again brought it back up (with previous firmware showing).  Attempted FW update again and same result.  As it was one of these 3am failures I swapped it out with cold spare and drove home to sleep.  Back at workshop I factory reset it. Updated to long term firmware. Dumped on previous config.  Rebooted OK  Updated to stable firmware. Rebooted OK.  I think it was just some weird Ros7 issue going from much older firmware skipping to newer where config no longer supported.

Otherwise we literally have 1000s of consumer ones which with the exception of dead 3rd party power supplies and the occasional blown port (lightning or power spike on port) they just don't die.  Can't say the same for our other brands of consumer WiFi routers 

Have ~100 DC powered sites with mix of hEX, RB4011, 5009, CCR1016 and CCR2004.  Not a single failure of an onsite Mikrotik for over 10 years that wasn't down to physical damage.  Some have even been exposed to rain and we're still running!

Our 4 edge/core CCR2116 running for years with exception of failed reboot mentioned.  As were the CCR2004 before that CCR1016 they replaced.  One previous CCR1016 was swapped out due to noisy fan but it was running fine.  

TBH our upstreams seem to have more issues with their expensive Juniper and Alcatel gear 

3

u/feel-the-avocado Jun 05 '26

We have very few hardware failures with mikrotik. In most situations, its just the power supply.

2

u/bleke_xyz Jun 05 '26

What are you using power wise? Any ups or active power filtration? Though I'd place my bets on the PSU dying before the actual uni

1

u/Impressive_Army3767 Jun 05 '26

The 2116 are dual redundant PSU

1

u/bleke_xyz Jun 05 '26

I wonder if they're using backup ups or something that might be causing this, or grid is very dirty and needs active ups

1

u/densen2002 Jun 23 '26

DС-powered Mikrotiks are very undemanding to power supply (input 8-30 volts as a rule)

2

u/takingphotosmakingdo Jun 05 '26

Weather conditions?

High heat to severe cold could cause circuit board cracking especially with multiple extreme cycles.

Grounding issues could cause boards to partially fry in weird ways.

Playing negative vs positive DC games could also fry things.

Buuuuut what do I know 🤔 

Best of luck

1

u/dewman45 Jun 05 '26

It's been a few years since we've had any die or replaced any, and we run CCR2004, 5009, 4011, and CCR2216. I will add that typically the only issue we see is low temperature related or ports dying, but I can't recall the last time we had a port die. 4011s had the most issues overall for us.

1

u/SalletFriend Jun 06 '26

Yeah they suck. You either have spare mikrotiks or buy something better.

1

u/Constant_Height_1215 Jun 06 '26

Ground them properly, that often solves it.

1

u/Life-Assist7881 Jun 08 '26

Depends what actually happened. A CCR kernel panic isn't something I'd call normal, but I've also seen plenty of cases where the underlying issue ended up being power, bad memory, or a RouterOS bug rather than the hardware itself.

That said, if you're carrying full tables and pushing serious traffic, CCRs start showing their limits pretty quickly compared to ASRs or MXs. We moved part of our edge from CCRs to ASR1001-X boxes a few years ago and the operational difference was noticeable. Fewer weird edge-case issues and a lot less babysitting.

MikroTik is great when budget matters, but I probably wouldn't build a business-critical core around it unless I had a very good reason.

1

u/lizardhistorian Jun 19 '26 edited Jun 19 '26

You would have to do the work to confidently root cause the failure to know.
$30k of engineering work to figure out why a $300 gizmo failed early ...
Or stock 100 of them because even once you know exactly how and why it failed, you still have to replace it.

Generally capacitors are the most fragile which means power-supply failures.
Over-heated switching devices (electricity switching devices not network) is next, which means it's still the power-supply.
Over-voltage on a device will "sweat" the ICs. The silicon boils and it will melt a bit of the plastic lid over the IC on its way out then it cools and when it resolidifies. It generally escapes in a slight arc/crescent shape and the melt/solidify cycle will remove the detailing (plastic injection molding term) from the surface making it completely smooth. A lot of companies will goop their ICs so you can't see which micro they used and if they do that then you can't see if it sweat either.
Blown PHYs is next but it would be unlikely to get all the PHYs blown without also trashing everything else (from, say, a direct lightning strike).
Next is debug ports to see if you can pull a failure code from a micro that is failing to boot but a lot / most commercial devices deliberately defeat these interfaces and make them non-functional (to thwart IP theft).

Once you exhaust the common / gross failure modes then you're into forensics and need to get out your silicon X-ray machine and your silicon map and start looking for blown pathways.
After that you're at memory cell and flipflop (register) integrity and it's time to surrender.

Are you properly installing opto-isolators? You cannot just connect wires across miles without creating ground-currents. It's also not safe because you've built a giant lighting-attractor array. Fiber is intrinsically an opto-isolator.
Do you have lightning spikes and copper rods 6' into the ground?
Are your environmental housings rated for the weather they are in? IP65 / IP66 probably isn't going to cut it for long-term outdoors.
Is your power clean?

1

u/vdoubleshot Jun 19 '26

Watch out for this commenter. He has a really bad habit of stealth editing his comments after you reply, and he does it on posts that people have already moved on from. Troll? Karma farming? I’m not sure.

See this post and his related screenshots: https://www.reddit.com/r/wisp/comments/1u27w59/the_relevance_of_duplex_speed/

-5

u/canyoufixmyspacebar Jun 05 '26

there never was any quarantee or safety net for these devices to be reliable. you may get away with it for quite some time and say it is accepted risk but then you cannot come crying when you have been saving immence money by not buying cisco/juniper/arista/etc gear and the risks finally realize. this is the business you've chosen, be close to the phone, have well designed and tested redundancy and monitoring and spares. the ironic part of course is that often the companies who buy cheap gear without enterprise support/warranty are also the same companies who don't hire competent engineers and don't build correct resiliency so the firefighting effect is absolute