Why High Availability Starts in the Network Today
Failures are not an exceptional state. They are an architectural principle.
There is a remarkable difference between traditional infrastructure and modern edge architecture. It does not show up during normal operation. It becomes visible at exactly the moment something goes wrong. That may sound like an unusual perspective at first. After all, most platforms spend the majority of their lifetime not failing. That is precisely why it seems natural to optimize infrastructure for normal operation. The history of modern platforms, however, tells a different story. The larger systems become, the less likely it is that every component is in the desired state at all times. Disks fail. Switches restart. Fiber links are damaged. Routers lose peerings. Data centers go offline. Hardware is maintained. People make mistakes. Perhaps the actual task of modern platform architecture is therefore not to prevent failures. Perhaps it is to design systems in such a way that failures lose their exceptional character.
Many traditional high-availability concepts follow a familiar pattern. There is an active system. And a passive one. A primary site. And a secondary site. A production server. And a standby.
Primary
+----------+
| Active |
+----------+
|
|
|
+----------+
| Passive |
+----------+
takes over on failure
This model did not emerge by accident. Hardware was expensive. Additional resources were therefore used only when they were actually needed. Redundancy often meant keeping systems available that spent most of their lifetime waiting. This model still works today. But it has an interesting characteristic. Failure mode is fundamentally different from normal operation. Roles suddenly change. Systems take on new responsibilities. Traffic is switched. Dependencies change. Exactly at the moment when something has already gone wrong.
Edge platforms often follow a different approach. Not because active-passive is inherently wrong. But because the network already has a property that can be used for this purpose. Multiple points of presence can provide the same services at the same time. If all sites already announce the same prefixes, enforce the same security policies, and carry the same responsibilities, one simple question inevitably arises. Why should any of them wait?
That is why modern edge networks usually operate active-active. Not as an optimization. But as a fundamental principle.
Internet
|
+---------------+---------------+
| | |
v v v
+---------+ +---------+ +---------+
| Hamburg | |Frankfurt| | Alsbach |
+---------+ +---------+ +---------+
| | |
+---------------+---------------+
All sites active
Every point of presence handles production traffic. Every site terminates TLS. Every site performs security checks. Every site distributes requests to the compute platform. There is no "replacement data center." No cold standby. No second place. Every site is always a full part of the same platform.
The real difference perhaps only becomes visible when one of these sites disappears. Assume Frankfurt fails completely. Not planned. Not controlled. But right now. Which DNS records need to be changed? None. Which IP address needs to be updated? None. Which application needs to be restarted? None. From the platform's point of view, a network node simply disappears.
Before the failure
203.0.113.42
+--------+--------+--------+
| | |
v v v
HAM FRA ALS
After the failure
203.0.113.42
+--------+ +--------+
| | |
v v v
HAM ALS
Frankfurt route withdrawn
The Frankfurt point of presence no longer announces its prefixes. Its neighbors detect this. Routing tables are updated. New connections automatically reach the remaining sites. Not because an application reacts. Not because a script runs. But because the internet was designed for exactly these kinds of situations.
Of course, routing alone is not enough. Not every failure affects an entire data center. Far more common are much less dramatic scenarios. A Kubernetes cluster still responds. A service is still running. The pods still exist. And yet the application is no longer healthy. Perhaps response times suddenly rise from ten milliseconds to several seconds. Perhaps a backend returns only HTTP 500. Perhaps a database is blocked. Or a dependency no longer responds. From the routing perspective, the system would still be reachable. From the user's perspective, it has already failed. This is where the second layer of modern high availability begins. Health checks.
Health checks are often reduced to simple HTTP requests. In reality, they are much more than that. They form the memory of a platform. They continuously answer questions such as:
- Is this backend reachable?
- Does it respond within the expected time?
- Does it return valid responses?
- Is only one service affected, or an entire site?
- Should new traffic still be sent there?
The user's actual request no longer has to answer these questions. The platform already knows the answer.
Perhaps this is the decisive difference. A traditional load balancer distributes connections. A modern edge platform continuously evaluates the state of its entire infrastructure. It connects routing with health information. It connects network decisions with application state. It connects the global internet with local platforms. This also changes the meaning of high availability. It no longer comes from having a second system waiting somewhere. It comes from the platform always knowing,
- which parts are healthy,
- which responsibilities they have,
- and where the next request should reasonably be sent.
Perhaps this is the most important insight of this article. Failures are not exceptions. They are part of the normal operation of every larger platform. The real quality of an architecture is therefore not shown by whether components fail. It is shown by how little significance that failure has for the user in the end. And that leads directly to the next question. If a platform continuously decides whether and where requests are sent, what role does DNS actually still play? Is DNS really just the phone book of the internet? Or has it long since become part of modern platform architecture itself? That is what the next part of Building the Edge will be about.