Compute Is Not the Platform
Why modern applications today consist of two platforms.
There is a remarkable side effect of technological abstraction. The more successful it becomes, the harder it is for us to see where its actual boundaries are. Kubernetes is probably one of the best examples. Few technologies have changed the way we build and operate applications as fundamentally. Compute resources are abstracted, workloads are described declaratively, and containers are scheduled independently of individual hosts. From a developer's point of view, the underlying infrastructure appears almost limitless. A deployment is created, the scheduler takes over, and moments later an application is already processing production traffic. But perhaps that success creates an interesting misconception. The better Kubernetes has become at abstracting compute from the infrastructure underneath, the more often it appears that Kubernetes is the platform. In reality, the platform begins much earlier.
Let us look again at the path of a request. Not the part Kubernetes processes. The part before it.
Browser
│
▼
DNS
│
▼
Internet
│
▼
Routing
│
▼
Anycast
│
▼
Edge
│
▼
Kubernetes
│
▼
Application
Interestingly, Kubernetes has no influence over the first seven stages of this diagram. It does not decide which network path a request takes through the internet, nor at which point of presence a connection is terminated. It knows no BGP routes, no peering relationships, no TLS handshakes, and no web application firewall. Even the question of whether a TCP connection is allowed to reach the ingress controller has already been answered long before the first pod knows the connection exists. This is where it becomes clear that we often mix two completely different classes of infrastructure problems.
Compute answers questions about applications. Edge answers questions about connections. At first, those statements may sound obvious. Their consequences, however, are surprisingly far-reaching. A Kubernetes cluster cares about,
- which pod processes a request,
- which resources are available to a container,
- which version of an application is currently in production,
- how deployments are rolled out,
- and which services inside the cluster communicate with each other.
All of these decisions assume that a connection reaches the cluster in the first place. The edge, by contrast, answers an entirely different class of questions.
- Which network path leads to this platform?
- At which location is the connection accepted?
- Is the target system currently reachable?
- Should the TLS connection terminate here?
- Is the request legitimate?
- Is this a user or automated traffic?
- Should this connection be allowed to reach the compute platform at all?
These are not the same problems. And that is why they should not be solved by the same platform.
Historically, this separation was barely necessary. A web server answered HTTP requests. The application ran on the same system. A reverse proxy terminated TLS. Perhaps there was also a firewall. Often, no more infrastructure was required. A single server could assume nearly all responsibilities at once. Today, that same server would face tasks that conflict with one another. It would have to make global routing decisions, detect DDoS attacks, manage certificates, terminate TLS, perform Layer 4 and Layer 7 load balancing, reach Kubernetes services, and then also run the actual application. Not because any one of these tasks is especially difficult. But because each of them depends on completely different information.
The difference becomes clearer if we stop talking about products and talk about responsibilities instead.
| Edge | Compute |
|---|---|
| Accept connections | Run applications |
| Routing | Scheduling |
| TLS | Containers |
| DDoS mitigation | Deployments |
| Layer 4 load balancing | Services |
| Layer 7 routing | Business logic |
| WAF | Data storage |
| Traffic engineering | Autoscaling |
What is remarkable is not the split itself. It is that both sides can evolve almost completely independently. A new Kubernetes version does not change BGP routes. A new peering relationship does not change a deployment pipeline. A change to the web application firewall does not affect StatefulSets. And a new storage system has no impact on how a request reaches the cluster in the first place. That is the real value of clearly defined responsibilities. They do not reduce the number of components. They reduce the number of dependencies between those components.
This separation has another advantage that often becomes visible only during operations.
Consider two completely different failure scenarios.
In the first, a point of presence loses all upstream connectivity.
In the second, a Kubernetes deployment fails and all pods of an application enter CrashLoopBackOff.
From a user's perspective, both initially have the same result.
The application is unavailable.
Architecturally, however, they are completely different.
Scenario A
Internet
│
▼
PoP unreachable
→ Routing problem
Scenario B
Internet
│
▼
Edge works
│
▼
Kubernetes
│
▼
Pods failing
→ Compute problem
The causes lie in different platforms. They are analyzed by different teams. They require different tools. And they are solved at entirely different layers. That is exactly why it makes little sense to artificially combine both responsibilities into a single platform.
Perhaps this also explains why modern platforms increasingly consist of multiple platforms. Compute is a platform. Storage is a platform. Observability is a platform. And edge is one as well. Not because each layer necessarily has to be implemented by different products. But because each layer has its own information space. Routing requires routing information. Scheduling requires cluster state. Storage requires consistency models. Observability requires telemetry data. None of these platforms can sensibly assume the responsibilities of another without requiring knowledge it should not need to possess.
Perhaps this is one of the biggest misconceptions in modern cloud architecture. We often say that we run applications on Kubernetes. In reality, we run them on infrastructure whose first point of contact with the user lies far before Kubernetes. The cluster begins where execution of the application starts. The platform begins where the internet decides which path an individual request will take. That is why modern infrastructure does not end at the load balancer. And it does not begin in the Kubernetes cluster either. It begins at the boundary between the public internet and the first decision about whether a connection is allowed to become part of our platform at all.