Skip to content
WindowsDeep Dive Published Updated 8 min readViews unavailable

Windows Container Networking: Diagnose HNS Networks and Endpoints

Trace Windows container networking from runtime requests through HNS networks, endpoints, virtual switches, IPAM, and workload-level connectivity.

Windows container networking is built from several cooperating layers. A container runtime or orchestrator requests network resources; Host Networking Service (HNS) creates networks and endpoints; Host Compute Service (HCS) and the container runtime connect the container; and Windows networking components implement switching, address assignment, NAT, filtering, or overlay behavior. A container can be running while its endpoint, route, DNS configuration, or policy is missing. Diagnosing only the Docker or Kubernetes object hides the host network state that explains the failure.

This guide focuses on Windows Server containers and the HNS control plane. Windows and Linux container networking have different host stacks and driver support. WSL2’s virtual machine networking and a Hyper-V virtual switch are related infrastructure but not interchangeable HNS resources. Confirm the exact host OS, container runtime, orchestration layer, network driver, and Windows container image version before applying a network change.

Identify the network model before inspecting state

Windows container network drivers expose different connectivity and isolation models. NAT uses an internal host network and translates egress. Transparent connects container endpoints to a physical network through a virtual switch and expects addresses compatible with that network. L2Bridge attaches endpoints to a bridged network using host network services. Overlay provides a virtual network spanning hosts and uses encapsulation and orchestration-managed address allocation. The supported set and feature limitations depend on Windows Server version and driver.

Do not infer the driver from a Docker network name such as nat or overlay alone. Inspect the runtime network object, HNS network data, host virtual switches, and the orchestration manifest together. A cluster may create or reconcile the HNS network after the runtime starts. A hand-created HNS object can conflict with the orchestrator’s expected state and be removed during reconciliation.

Write down the expected path for one failing container: container namespace and endpoint, HNS network, host vNIC/vSwitch, route or NAT, node interface, upstream gateway, and destination. For a multi-host service, include the overlay or load-balancer control plane and the service address separately from the pod/container endpoint address. A reachable endpoint IP does not prove service discovery or load balancing works.

Capture runtime and HNS state without repairing it

Collect the container runtime’s view and HNS’s view before restarting the service or deleting a network. On hosts with the HNS PowerShell helper module, capture networks and endpoints as structured JSON, alongside visible network adapters and virtual switches:

$out = Join-Path $env:TEMP 'hns-network-case'
New-Item -ItemType Directory -Path $out -Force | Out-Null

Get-HnsNetwork |
    ConvertTo-Json -Depth 32 |
    Out-File -FilePath (Join-Path $out 'hns-networks.json') -Encoding utf8
Get-HnsEndpoint |
    ConvertTo-Json -Depth 32 |
    Out-File -FilePath (Join-Path $out 'hns-endpoints.json') -Encoding utf8
Get-NetAdapter -IncludeHidden |
    Select-Object Name, InterfaceDescription, ifIndex, Status, MacAddress |
    Export-Csv -NoTypeInformation -Path (Join-Path $out 'net-adapters.csv')
Get-VMSwitch -ErrorAction SilentlyContinue |
    Format-List * |
    Out-File -FilePath (Join-Path $out 'virtual-switches.txt') -Encoding utf8

The HNS helper cmdlets and Hyper-V cmdlets may not be installed on every host or may be exposed through a runtime-specific module. Check Get-Command Get-HnsNetwork, Get-HnsEndpoint, Get-VMSwitch -ErrorAction SilentlyContinue before automating collection. A missing helper command is not evidence that the service has no network state. Use the runtime’s supported inspect command as a separate capture, for example docker network ls and docker network inspect <network> where Docker is the active owner.

Compare network and endpoint IDs, subnet, policies, gateway, DNS, MAC address, and attachment state. Do not assume object names are globally unique or stable across runtime recreation. Identify which component owns each object. If the runtime says a network exists but HNS does not, or the reverse, retain both outputs and note when they were captured; the mismatch can indicate failed creation, stale state, or reconciliation delay.

Trace address assignment, route, and name resolution

For NAT, verify that the HNS network’s internal prefix does not overlap with the host’s physical, VPN, or other virtual networks. Overlapping subnets can route traffic to the wrong interface even when the container received a valid address. Inspect the endpoint’s IP, gateway, DNS information, host NAT configuration where used, and host routes. A working outbound connection from the host does not establish a working translated path from the container.

For transparent networks, confirm that the selected host interface and upstream switch allow the container endpoint MAC/address behavior required by the deployment. A DHCP-capable transparent network depends on the external network’s address assignment and admission controls. For an overlay, validate node-to-node underlay reachability, required encapsulation traffic, MTU, IPAM allocation, and the control-plane policy for the chosen orchestrator. Do not open broad UDP or overlay traffic without identifying the exact supported protocol and node set.

Test in order: container-local interface and route, host endpoint state, node-to-gateway reachability, cross-node endpoint path, service IP/load-balancer behavior, and DNS. A DNS failure can be caused by missing DNS settings in HNS, an unreachable resolver, or orchestrator service discovery; these are distinct from endpoint creation. Compare a new container with an existing healthy endpoint on the same network and a healthy endpoint on another node to localize the failure.

Interpret policy and version-specific boundaries

HNS applies networking policies that can include NAT mappings, access-control lists, load balancing, and encapsulation through Windows Virtual Filtering Platform components. The policy representation differs by network driver and host version. Treat the HNS JSON as implementation state to inspect, not as a supported hand-edit surface. Use Docker, Kubernetes, or the supported Windows management interface as the configuration authority.

Feature compatibility matters. Microsoft’s Windows container documentation lists unsupported combinations and constraints, including limitations for IPv6 across driver types and unsupported host-mode networking in documented configurations. A manifest copied from Linux may use flags that Windows does not support. Validate the exact Windows Server release and driver matrix rather than assuming Linux kernel network namespaces or iptables semantics map to HNS.

Windows Server updates, runtime upgrades, CNI changes, and host interface renames can all affect endpoint creation. Record the runtime and CNI versions, OS build, network driver, switch configuration, and interface identity before comparing behavior across nodes. If a node is reimaged or its NIC order changes, a script that expects Ethernet may bind the HNS network to the wrong interface even though the interface still exists under another name.

Diagnose failed or stale HNS objects conservatively

If no endpoint appears after a container starts, correlate the runtime operation and HNS events. Find whether network creation succeeded, endpoint creation was requested, IPAM allocated an address, and HCS attached the endpoint. Use the HNS/HCS operational logs available on the host and correlate timestamps with the container runtime and orchestrator. Enumerate channels with Get-WinEvent -ListLog '*Host-Network-Service*' and related HCS providers before assuming a particular event channel or ID exists.

If endpoints exist but traffic fails, compare policy and route state against a healthy container. Check vSwitch port status, host firewall rules, VFP policy, network profile, NAT mappings, and upstream ACLs. Pktmon or a supported ETW capture can help locate where packets stop, but capture one bounded reproduction and preserve the exact interface and filter configuration. Avoid changing both HNS and the host firewall simultaneously; that erases the evidence needed to locate the filtering layer.

Do not delete C:\ProgramData\Microsoft\Windows\HNS\HNS.data or clear all HNS state as an early recovery step. The file and service state can represent networks and endpoints owned by a running orchestrator. Destructive cleanup may remove unrelated virtual networks, disrupt every container on the node, or leave the runtime and HNS inconsistent. Follow the container platform’s supported node-drain and network-recovery procedure, capture state first, and test cleanup on a disposable node.

For Kubernetes, inspect both the Windows node’s network plugin and Kubernetes objects. The HNS network can be missing because kubelet or CNI initialization failed, while the cluster control plane still reports a Node object. Compare node events, CNI logs, kubelet logs, HNS networks/endpoints, and the pod’s assigned address. Microsoft’s Windows Kubernetes guidance documents scenarios where HNS network creation failure prevents the required virtual adapter from appearing. Do not manually create cbr0 or vxlan0 unless the current plugin runbook explicitly requires it.

Validate a recovery in a controlled sequence

Drain or otherwise remove the affected host from new workloads using the orchestrator’s supported procedure. Record current network objects, container state, runtime version, and active flows. Reconcile the configuration through its owner, then start one test container or pod on a known network. Confirm endpoint presence, expected IPAM, DNS, route, gateway reachability, same-node communication, cross-node communication, and service-level behavior. Reintroduce workload only after the platform reports the node and network healthy.

Test failure and recovery cases in a lab: a missing physical uplink, overlapping NAT prefix, blocked overlay path, failed CNI invocation, runtime restart, node reboot, and interface rename. Observe how the orchestration layer retries and whether old endpoints are garbage-collected. A single docker run that reaches the Internet does not verify the Kubernetes service path or a production overlay policy.

HNS incident checklist

  1. Identify Windows Server build, runtime, orchestrator/CNI, network driver, and owner of the network object.
  2. Capture runtime networks, HNS networks/endpoints, virtual switches, adapters, routes, and relevant logs before restarts.
  3. Compare IDs, subnets, gateways, DNS, IPAM, and policy with a healthy endpoint on the same node and another node.
  4. Trace packets from the container through HNS/vSwitch/NAT or overlay and upstream network boundaries.
  5. Check Windows-specific feature support for the exact driver and OS version.
  6. Reconcile state through Docker/Kubernetes or the supported management tool, not by hand-editing HNS data.
  7. Validate endpoint, DNS, cross-node, service, and recovery behavior before scheduling new production workloads.

HNS debugging is an ownership exercise: the runtime requests, HNS realizes, and Windows networking carries the traffic. Keeping each layer visible prevents a broad network reset from destroying the state that would have identified the actual fault.

Related:

Sources:

Comments