Skip to content
SRE & DevOpsDeep Dive Published Updated 7 min readViews unavailable

Kubernetes Pod DNS: Search Domains, Policies, and Resolution Tests

Debug Kubernetes Pod DNS by inspecting resolver configuration, namespace search paths, ndots, DNS policies, headless Services, and CoreDNS behavior.

Kubernetes configures name resolution for Pods so applications can locate Services and other DNS names. Many apparent “cluster DNS outages” are actually a mismatch between the name an application queries, the Pod’s namespace, its DNS policy, and the resolver search list. Start with the resolver configuration inside the affected Pod, then test a known fully qualified Service name and compare it with the short name the application uses.

For a Service named database in namespace production, the fully qualified cluster DNS name is commonly database.production.svc.cluster.local, where cluster.local is the cluster domain configured for that cluster. Do not hard-code that suffix without verifying the cluster configuration. Kubernetes documents the record structure and search behavior, while the cluster DNS implementation and administrator configuration determine how upstream names are forwarded.

Read the Pod’s resolver configuration first

The kubelet configures /etc/resolv.conf in the container according to the Pod’s dnsPolicy, the node’s resolver settings, and any custom dnsConfig. A typical ClusterFirst Pod receives the cluster DNS service address, search suffixes for its namespace and cluster services, and an ndots option. The actual file is authoritative for the running Pod; do not infer it from a Helm values file or a different node.

kubectl exec -n production deploy/api -- cat /etc/resolv.conf
kubectl exec -n production deploy/api -- nslookup database
kubectl exec -n production deploy/api -- \
  nslookup database.production.svc.cluster.local

Use an image that includes the diagnostic tool, or launch a temporary approved debug container if the application image is minimal. Do not modify production Pods just to add tools. Record the Pod name, node, image, and DNS policy with the output because two replicas can run on different nodes with different upstream conditions.

Short names are resolved relative to the configured search path. A bare name such as database commonly resolves only within the caller’s namespace. A Pod in test should use database.production or the full service name to reach a Service in production. Partially qualified names can work through search expansion, but their behavior is less obvious and depends on resolver rules. Prefer a fully qualified name for cross-namespace dependencies when predictable resolution matters.

Understand DNS policies and custom settings

ClusterFirst is the normal Pod policy for cluster workloads: names that do not match the cluster domain are forwarded according to cluster DNS configuration. Default inherits the node’s name resolution configuration. None requires explicit dnsConfig values. ClusterFirstWithHostNet is intended for Pods using host networking that still need cluster DNS behavior; otherwise a host-network Pod can fall back to node-style resolution.

dnsConfig can specify nameservers, search domains, and resolver options, and its values are merged with the base configuration according to the selected policy. Use dnsPolicy: None only when the workload truly needs a separate resolver configuration and include the required nameservers and searches. A custom resolver is not a bypass for cluster networking: the Pod still needs network access to the chosen DNS servers, and NetworkPolicy or node firewall rules can block that path.

Treat ndots as a query-order control, not a DNS availability switch. With a high dot threshold, a query that looks like an external domain may first be tried against each search suffix before the resolver attempts the absolute name. This can generate extra queries and latency, particularly for applications that repeatedly query external names. Lowering ndots can reduce search expansions but may change how intended short internal names resolve. Measure the application workload and test both internal and external names before changing it.

Search lists also have implementation and size limits. Large custom lists can cause resolver truncation, startup validation errors, or incompatibility with older runtime components. Keep the list short and purposeful; do not append every namespace as a convenience. Explicit service names are usually clearer than multiplying global search suffixes.

Distinguish Service records from endpoint behavior

A normal Service DNS record resolves to the Service’s virtual IP. The Service data plane then selects a backend. If DNS resolves correctly but connections fail, inspect the Service selector, EndpointSlices, readiness, ports, and network path rather than repeatedly changing DNS. DNS answers and endpoint readiness are separate control-plane layers.

A headless Service sets clusterIP: None and has different DNS behavior: the DNS response can expose addresses for backing endpoints instead of one virtual IP. This is useful for peer discovery and StatefulSets, but clients then need to handle multiple records, endpoint churn, and cache expiry. An ExternalName Service returns a CNAME to the configured hostname; it does not create a proxy or connect to the target on behalf of the client.

Do not treat any hostname layout that happens to resolve as a stable API unless Kubernetes documentation specifies it. Use the documented Service and Pod DNS records. For StatefulSet peers, use the governing headless Service and the documented Pod identity pattern rather than synthesizing names from implementation assumptions.

Trace the failure from the Pod outward

When a name fails, compare these layers in order:

  1. Inspect /etc/resolv.conf inside the failing Pod and compare it with a working replica.
  2. Query the fully qualified Service name, then the short name; note whether the difference is namespace/search-path related.
  3. Check the Service object, its ClusterIP or headless configuration, and EndpointSlices.
  4. Check that the Pod can reach the configured DNS server and that the DNS service has healthy endpoints.
  5. If cluster names work but external names fail, inspect CoreDNS forwarding, node upstream resolvers, egress policy, and provider firewall behavior.

Record the exact query and response code. NXDOMAIN, SERVFAIL, timeout, and an empty answer point to different failures. A timeout suggests a reachability or overloaded-server path; NXDOMAIN can mean an incorrect name or a search expansion that changed the query; SERVFAIL can indicate upstream or resolver processing failure. Avoid collapsing all failures into “DNS is down.”

The Kubernetes DNS debugging guide uses tools such as nslookup from a Pod and inspection of the DNS add-on’s Pods and Services. Follow the distribution’s procedure for CoreDNS or another DNS implementation. Check logs and metrics for forwarding failures, throttling, and cache behavior, but avoid changing cluster DNS configuration until the affected Pod’s exact resolver path is understood.

Avoid masking resolver failures in application code

Applications should use bounded DNS timeouts and retry with backoff appropriate to the protocol. A retry storm from every replica can overload a resolver during an outage. Cache names only for a policy-aware duration; indefinite caching can keep a stale endpoint after a Service or external target changes. If the language runtime has its own DNS cache or uses a different resolver library, inspect that behavior separately from /etc/resolv.conf.

Use service discovery names for service identity, not fixed Pod IP addresses. A Service’s stable name allows endpoint membership to change without changing application configuration. For third-party hostnames, test from the actual workload namespace and account for egress controls and upstream DNS dependencies. Do not copy host resolver files into a Pod as a generic workaround; that can bypass the cluster search path and break in-cluster Service names.

Test changes under realistic conditions

Test short same-namespace names, cross-namespace names, fully qualified cluster names, external hostnames, and negative answers. Repeat on nodes with different resolver configuration if your cluster has heterogeneous node pools. Test hostNetwork, custom dnsConfig, headless Services, and rolling changes to DNS add-ons where the workload uses them.

After changing ndots, search domains, or DNS policy, verify both lookup count and end-to-end request latency. Ensure observability distinguishes application-level DNS errors from connection failures after a successful lookup. A reliable DNS runbook starts inside the affected Pod, follows its actual resolver configuration, and separates name synthesis from Service routing and upstream forwarding.

Related:

Sources:

Comments