Kubernetes Service Traffic Distribution: Topology Preferences Without Strict Locality
Use trafficDistribution to prefer same-zone or same-node endpoints, understand fallback and policy precedence, and verify capacity before production rollout.
Kubernetes Service traffic distribution expresses a preference for where a Service sends traffic, without making locality an absolute requirement. In the current Kubernetes 1.37 documentation, spec.trafficDistribution supports PreferSameZone and PreferSameNode; the older PreferClose value is deprecated in favor of the more explicit zone preference. The field is useful when cross-zone or cross-node traffic has measurable cost or latency, but operators must validate that local endpoints have enough capacity before enabling it.
This control is not a load-balancing guarantee, a session-affinity setting, or a replacement for strict traffic policies. It influences endpoint selection through topology hints. Traffic can still fall back beyond the preferred topology when local healthy endpoints are unavailable, and the actual dataplane must implement the Kubernetes Service behavior. Treat it as a routing preference whose outcome must be measured in the cluster that will run it.
Apply a preference to a Service
The field is part of the core Service specification. A minimal example for a multi-zone internal service looks like this:
apiVersion: v1
kind: Service
metadata:
name: catalog
namespace: production
spec:
selector:
app.kubernetes.io/name: catalog
ports:
- name: https
port: 443
targetPort: https
protocol: TCP
trafficDistribution: PreferSameZone
PreferSameZone asks the Service dataplane to favor healthy endpoints in the client’s zone. PreferSameNode favors endpoints on the client’s node and can reduce a network hop for suitable local request patterns. A client with no usable endpoint in its preferred topology does not necessarily lose access: the documented behavior falls back to the same zone and then to endpoints across the cluster for PreferSameNode, and cluster-wide when its zone has no available endpoint for PreferSameZone.
Omitting trafficDistribution leaves the default strategy in place. Adding the field does not change the Service selector, create Pods, spread workloads, or add replicas. It only expresses a preference for choosing among the Service’s eligible endpoints. A Service with an empty or incorrect selector, unhealthy backends, or an insufficient number of replicas remains broken; topology hints cannot repair a workload or its readiness configuration.
Understand hints, not a per-request contract
The EndpointSlice controller communicates topology preferences using endpoint hints. A Service proxy uses those hints in its routing decisions. This separates endpoint membership and readiness from the locality preference: the controller still publishes endpoints, while the service dataplane uses the topology information when selecting backends.
Do not read the field as a deterministic percentage split. PreferSameZone is not “send exactly 70 percent locally,” and PreferSameNode is not a guarantee that a request will never leave its node. The feature prioritizes endpoints in a topology, and may send a concentrated share of demand to that topology. Specific balancing details depend on the cluster’s supported Service implementation and endpoint state. Verify the behavior with the cluster version and networking implementation actually deployed rather than extrapolating from a manifest alone.
This matters especially during uneven client traffic. If most clients run in one zone, their preferred endpoints may receive most of the Service load even when the cluster has many healthy endpoints elsewhere. If clients are on a node with one local backend, same-node preference can create a hot Pod while healthy peers on adjacent nodes remain lightly loaded. The lower-latency route is not automatically the higher-availability or higher-throughput route.
Distinguish preferences from strict traffic policies
internalTrafficPolicy and externalTrafficPolicy express stricter locality semantics for their respective traffic types. When the applicable policy is Local, it takes precedence over trafficDistribution; an endpoint outside the required locality is not simply made eligible by the preference. When the policy is Cluster or unset, traffic distribution can guide endpoint selection for that traffic type.
This distinction should be explicit in a design review:
| Mechanism | Operational intent | Failure and fallback question |
|---|---|---|
trafficDistribution: PreferSameZone |
Prefer same-zone backends while preserving broader fallback | Can the local zone serve its normal peak load, and what happens when it has no healthy endpoint? |
trafficDistribution: PreferSameNode |
Prefer node-local backends, then fall back to wider topology | Can the local Pod absorb its share, and how will traffic behave on nodes without a local endpoint? |
internalTrafficPolicy: Local |
Restrict internal traffic to node-local endpoints | Is there a local ready endpoint for every client node that needs this Service? |
externalTrafficPolicy: Local |
Preserve source address and restrict external routing to local endpoints where supported | Does the external load balancer health-check and drain nodes correctly? |
sessionAffinity |
Prefer sending a client back to a prior endpoint | Is client-IP affinity the right identity and stickiness duration for this application? |
The names sound related, but they solve different problems. Do not add both a strict Local policy and a topology preference without modeling their precedence. A Service that intentionally sets internalTrafficPolicy: Local cannot use PreferSameZone to make cross-node endpoints a safe fallback for internal traffic. Likewise, sessionAffinity keys a client’s repeated connections according to its own rules; a zone preference is not sticky identity.
Compare with topology-aware routing
Kubernetes also documents the service.kubernetes.io/topology-mode: Auto annotation for topology-aware routing. Both that mechanism and PreferSameZone aim to favor local-zone traffic, but they have different distribution behavior. The documented Auto heuristic uses allocatable CPU information to proportionally distribute traffic across zones and includes fallback safeguards for small endpoint populations. PreferSameZone is simpler: it favors the local zone when suitable endpoints are available, which is more predictable but can concentrate traffic and overload a small local pool.
When service.kubernetes.io/topology-mode: Auto is present, the current Service proxy documentation says that it takes precedence over trafficDistribution. Avoid setting both as competing policies. During migration, choose one control plane for topology routing, test the exact behavior in a canary, and remove the older annotation only after observing endpoint selection and service-level latency. Kubernetes notes the annotation may be deprecated in favor of the field in the future; do not treat that statement as a removal date or assume every managed distribution has identical upgrade timing.
Capacity-plan for uneven zones and nodes
Before enabling a locality preference, model request arrival and serving capacity by topology. Total cluster replica count is not enough. For each zone or node that will receive local preference, estimate peak request rate, concurrency, CPU and memory demand, cache footprint, and the number of ready backends available during normal and degraded states.
For a simple zone-level capacity check, compare the peak client demand originating in zone z with the safe serving capacity of ready backends in z. If the zone is expected to receive 1,800 requests per second during peak and its local pool can safely handle only 1,200, the remaining cluster capacity may not automatically prevent overload while the preference keeps selecting local endpoints. The numbers are an example for planning, not a Kubernetes threshold or guarantee.
Use topology spread constraints or carefully designed zone-specific Deployments to make endpoint placement match the desired traffic pattern. Spread constraints can keep replicas from clustering in one failure domain, but they do not prove traffic is evenly distributed. Zone-specific scaling can be useful when client arrival differs by zone, but it adds deployment and capacity-management complexity. In either case, validate the result using real endpoint topology and representative traffic.
Test failure cases, not only steady state:
- A zone loses all ready replicas while clients remain active there.
- A node has no local endpoint, or its sole local endpoint fails readiness.
- A rollout temporarily reduces ready local capacity.
- Autoscaling adds endpoints in another zone before it can add them locally.
- One zone’s client demand grows faster than its replica count.
- A zone or node returns after an outage and receives a sudden burst of preferred traffic.
- A policy setting such as
internalTrafficPolicy: Localchanges which fallback paths are valid.
Check that recovery does not produce a synchronized surge. A preference may send clients back toward recovered endpoints quickly; if the workload has cold caches, startup penalties, or expensive connection establishment, measure the warm-up behavior and monitor the first minutes after capacity returns.
Roll out and inspect safely
Start with a non-critical Service or a limited canary. Confirm the API server accepts the field in the target cluster version, inspect the resulting Service object, and check that EndpointSlices contain the expected topology information. For example:
kubectl get service catalog -n production -o yaml
kubectl get endpointslices -n production \
-l kubernetes.io/service-name=catalog -o yaml
kubectl get pods -n production -l app.kubernetes.io/name=catalog \
-o wide --show-labels
These commands show configuration, endpoint membership, and node placement. They do not by themselves prove which endpoint served each request. Use application telemetry, proxy or dataplane metrics, controlled load, and per-zone request traces to verify actual routing. Compare request volume, error rate, latency percentiles, saturation, and readiness by topology before and after the change.
Ensure your labels and node topology are meaningful. The common zone and region labels are generally supplied by the platform or node configuration, but their exact presence and correctness depend on the cluster. A label that is missing, stale, or inconsistent across nodes defeats assumptions in the routing model. Inspect node labels rather than assuming every cluster provider exposes identical topology metadata:
kubectl get nodes -L topology.kubernetes.io/region,topology.kubernetes.io/zone
Also confirm the Service selector selects only the intended workload and that Pod readiness accurately reflects whether a backend can serve traffic. Endpoint readiness feeds eligibility; a process that responds to a shallow probe while its dependencies are unavailable can still be selected and fail requests.
For rollback, remove the preference or restore the previous routing configuration through the normal deployment path. Observe the return to the cluster’s default strategy and verify that no other traffic policy, topology annotation, or external proxy continues to enforce locality. A rollback is not complete merely because the YAML field disappeared; the dataplane and client behavior must converge.
Operational signals and acceptance criteria
No single Service field exposes a complete “locality worked” metric. Instrument the application and its network path with topology labels that are bounded and reviewable. Useful views include requests served by backend zone or node, local-versus-remote request ratio, endpoint count and readiness per topology, p95/p99 latency, connection errors, and CPU or queue saturation for the local pool. Avoid adding unbounded labels such as Pod UID, request ID, or arbitrary client identifiers to metric series.
Alert on the failure modes that matter to users: local-pool saturation, increasing error rate, loss of ready endpoints in a topology, or a sustained latency increase. A low cross-zone traffic ratio is not itself success if local backends are overloaded. Conversely, cross-zone fallback during a local outage can be the intended availability behavior. Define these outcomes in the SLO and runbook before treating telemetry as a pass/fail signal.
An operational rollout is ready when the exact cluster and dataplane accept the field, topology labels and EndpointSlice hints are present, normal and degraded tests show acceptable latency and load distribution, the fallback path is known, and the service owner can disable the preference without an emergency change. Keep the preference only if observed cost or latency improvement outweighs the additional skew and failure-domain considerations.
Related:
- Kubernetes Service Traffic Policies: Local Endpoints and Source IP
- Kubernetes Topology Spread Constraints: Design for Failure Domains
Sources: