Before I was dealing with edge workers and cloud architectures, I worked tech support at MZTA. We landed a prestige government contract to move an entire region’s CHPP monitoring to our controllers.

Deploying in the city was seamless—a routine task. But when we started rolling out to the rural areas, the operational realities of dealing with legacy code and telecom providers hit hard.

Two massive issues completely paralyzed the rollout. Our main plant tech support couldn't handle either of them.

In rural areas, we used mobile modems as a reserve channel for telemetry. The hardware setup was straightforward, complete with a dedicated relay to physically hard-reboot the modem if it locked up.

When the rural plants came online, the traffic just stopped.

We could see our controllers firing the telemetry packets, but they were completely swallowed by the mobile provider's network. This was a battle-tested solution that worked everywhere else. I escalated the issue as high as possible with the telecom provider's tech support. They refused to provide network traces and played dumb.

After days of useless back-and-forth, I gave up on their support and started digging through the deepest pages of their technical documentation. I found the trap: they silently identify and throttle all M2M traffic on standard SIM cards. If you want your packets to actually route, you have to buy their dedicated "M2M Tariff."

The economics were insulting. A standard tariff was roughly $1 for 10GB of data. The secret M2M tariff? $1 per 1 KB of traffic.

We had never hit this issue before because our previous deployments happened to use a different provider on the controller side. The solution wasn't to argue, and it certainly wasn't to pay the extortionate rate. The fix was purely pragmatic: we threw their SIM cards in the trash and swapped to a different mobile carrier.

The second problem was technically much more interesting. The client wanted to scale from a single workstation to a multi-node client-server architecture. They paid our plant to pre-configure a heavy server, tested it locally at the factory, and shipped it to the site.

They plugged it into the local network, and half the system instantly broke.

Telemetry was received and logged perfectly. But not a single control command could be sent from the server to the plant controllers. The workstations could send commands fine, but the heavy server was dead in the water.

Our official plant tech support, who had just taken the client's money to pre-configure this exact machine, offered a brilliant solution: "Just change the server."

That answer was stupid, so I started mapping the problem myself. We were receiving telemetry, which meant the physical layer was fine. The issue was purely in the software transport layer.

I dug into the legacy source code of the server implementation. Out of the entire monolithic codebase, the original developers had chosen to use UDP for exactly one thing: sending outgoing commands.

I spun up a test environment, installed Wireshark, and took a packet dump from their server. No errors were logged. The software was firing the UDP packets, but they were vanishing.

Then I looked at how the UDP socket was actually configured in the code. The correct way to handle UDP on a multi-homed server is to explicitly bind the socket to the specific local IP address of the external-facing network interface before calling send.

The legacy code didn't do this. It skipped the explicit bind and just threw the UDP packet blindly at the OS routing table. When you fire a connectionless protocol like UDP without binding a specific source address, Windows simply looks at its adapter metric table and routes the packet out of whichever network interface is ranked first.

The client had two NICs on the server. Because of the random order they plugged the patch cords into the back of the machine, the "first" interface was their internal local network, not the external-facing network that connected to the controllers.

Our commands were being shot into a local LAN segment and dying silently.

The software fix would be to expose an interface binding setting in the config, but we needed the plant running immediately. The operational solution took five clicks: open the Windows adapter settings and change the Interface Metric number to prioritize the correct external NIC.

We wasted a month arguing with our own plant support, and I spent a week tracing packets, all to fix a problem caused by the physical order of patch cords. For over 30 years, every single customer who bought a multi-NIC server from us had simply gotten lucky and plugged the external cable into port 1. This client was the first one to break the streak and bring a three-decade-old legacy routing bug to the surface.