Windows DHCP Failover: MCLT, Partner States, and Lease Safety
Design Windows DHCP failover around MCLT and partner states, replicate scope settings safely, and verify lease continuity with observable acceptance tests.
DHCP failover is not two independent servers pointed at the same scope. It is a relationship between two Windows DHCP servers that maintain synchronized lease state for shared IPv4 scopes and coordinate which server may offer or renew a lease. The design reduces dependence on one DHCP service, but it does not remove the need for working relay paths, correct scope configuration, or an operator who understands what a partner state means.
The most important operational distinction is between lease-state replication and scope-configuration replication. Failover partners maintain separate synchronized lease databases. Changes to scope settings, such as options or exclusions, must still be replicated to the partner. Treat those as separate health signals and verify both before declaring the service resilient.
Start with the supported topology
Windows DHCP failover is configured for IPv4 scopes and supports two servers in a relationship. Microsoft documents current support for Windows Server 2016, 2019, 2022, and 2025. DHCPv6 scopes cannot be enabled for this failover relationship type. If DHCP is clustered, the cluster is treated as one server endpoint; configure the cluster name or address as the partner, not the name of a node that can change ownership.
The clients must be able to reach both partner servers, either directly or through a correctly configured DHCP relay. This is easy to overlook in routed networks: a relay configured to forward to only one server defeats the intended server redundancy for clients on that segment. Verify the relay destination list for every subnet, not only the subnet where the DHCP servers reside.
Get-DhcpServerv4Failover -ComputerName "dhcp01.example.com" |
Format-List *
Get-DhcpServerv4Scope -ComputerName "dhcp01.example.com" |
Format-Table ScopeId, Name, State, StartRange, EndRange
Run these read-only inventory commands against both servers and compare the relationship, partner, scopes, and scope ranges. Avoid relying on a successful console connection as proof that both DHCP services can communicate with one another or that the relay path reaches both.
Choose the mode based on placement
In load-balance mode, both servers can serve a scope. By default, the load ratio is 50:50, although the ratio can be configured. The servers use a hash of the client’s MAC address to select which partner should respond under normal operation. This is a distribution mechanism, not a guarantee that exactly half of all lease transactions or bytes will land on each server. A small client population or uneven renewal patterns can produce an apparently uneven load.
Hot-standby mode assigns one server the active role and the other the standby role. It is useful when the standby is intended to serve a remote site or provide reserve capacity only when the active server is unavailable. The default reserve percentage is 5 percent of the available addresses. That reserve is for new clients; a standby that cannot reach the active partner can temporarily renew a client’s existing address for the Maximum Client Lead Time (MCLT).
Both modes are configured for a relationship and its IPv4 scopes. A DHCP server may be active for one relationship and standby for another. Keep the design and the relay topology aligned: two servers at one site can use load balance, while a remote standby can be more appropriate when the WAN is expected to partition. Neither mode creates an IP address pool, a routing path, or a disaster-recovery plan that does not already exist.
Understand MCLT before changing partner state
MCLT is a safety bound on how far one server can extend lease information beyond what its partner has confirmed. It is not the configured client lease duration, a heartbeat interval, or a promise that a partner will take over immediately. During a communications interruption, the servers cannot assume that the other server is dead rather than temporarily unreachable. They therefore use constrained lease behavior to avoid both servers making unsafe assumptions about the same address.
When a server is deliberately placed in the Partner Down state, it waits for the MCLT period before it can assume responsibility for the full available address pool. The waiting period helps preserve address uniqueness when the prior partner may still have issued leases that were not synchronized. In hot-standby mode, the standby can use its configured reserve for new leases while waiting, and can temporarily renew known clients. In load-balance mode, loss of partner communication makes a server eligible to answer all client requests. It temporarily renews a lease assigned to the partner for the MCLT duration. For a client without an existing lease, it allocates from its own free pool first and then can use the partner’s free pool when its local pool is exhausted. Only after entering Partner Down and waiting through MCLT does it assume responsibility for 100 percent of the address pool; a broken TCP connection alone does not grant that authority.
Automatic state transition from communications-interrupted to partner-down is configurable. It can reduce manual intervention, but it also converts a network partition into a unilateral decision after the configured timer. Set it only after mapping expected link failures, partner reachability, and recovery ownership. A brief ACL or routing fault should not be mistaken for permanent server loss. For planned work, preserve a path to the partner and monitor state transitions rather than manually forcing Partner Down as a routine maintenance step.
Keep configuration in sync
Initial relationship creation copies the scope settings and active leases needed by the partner. Later scope edits are not a substitute for replication: Microsoft requires scope-parameter changes to be manually replicated to the partner. Replication can be initiated from either member. Review the result instead of assuming that a successful command changed both copies.
# Review the relationship before changing anything.
Get-DhcpServerv4Failover -ComputerName "dhcp01.example.com" -Name "Branch-01" |
Format-List *
# After an approved scope-configuration change, replicate the relationship.
Invoke-DhcpServerv4FailoverReplication `
-ComputerName "dhcp01.example.com" `
-Name "Branch-01" `
-Force
The replication command is an intentional configuration write. Run it from an elevated PowerShell session with the DHCP Server module on the member that holds the intended scope settings: replication copies settings from the initiating server and overwrites the partner’s values. Confirm the expected scope IDs in the command output, then inspect the relationship and scope options on the partner. Do not configure a second copy of a failover-enabled scope manually as if the pair were a traditional split-scope deployment.
If both servers update DNS dynamically on behalf of clients, configure them with the same DNS credentials. Otherwise the partner may not own the DNS records created by the original DHCP server and later updates can fail. Test a lease and its DNS update through each partner path; a successful address assignment alone does not validate dynamic DNS ownership and refresh behavior.
Monitor the relationship, not just the service process
Windows exposes failover state and configuration events in the DHCP Server Admin and Operational event channels. The Admin channel records local or partner state changes, communications up/down events, and time synchronization warnings. The Operational channel records relationship and scope configuration changes. The DHCP audit log also records failover-related client message behavior.
Useful failover counters include binding updates sent and received, acknowledgements, pending outbound binding updates, dropped updates, and transitions into Communications Interrupted, Partner Down, or Recover. Alert on sustained queue growth or dropped updates, not on one transient counter sample. A server process can be running while its relationship is unhealthy or its partner state is stale.
Correlate the event timestamp with routing, firewall, DNS, time synchronization, and relay logs. Both partners need a persistent TCP/IP connection for coordination, and Microsoft logs an event when that connection is established or lost. If server clocks drift, DHCP also records a time-out-of-sync event. Capture the relationship state from both sides before changing timers or state; two one-sided views can reveal a partition.
Acceptance test and failure cases
In a maintenance window or isolated test segment, validate each subnet through a representative relay. First verify a client can obtain a new lease and renew an existing lease with the relationship healthy. Confirm that the assigned address is in the expected scope, the router and DNS options are correct, and the client can reach the expected services. Then test one partner unavailable at a time and record how long new leases and renewals take to recover. Restore the partner and confirm both sides return to the expected state without manual pool edits.
Use measurable acceptance criteria:
- Both servers report the same relationship, partner, scope membership, and scope parameters.
- A new client and a known client obtain the expected service during a single-server outage, within the documented reserve and MCLT behavior of the chosen mode.
- No duplicate address is observed during the planned test, including after the failed partner returns.
- The DHCP state events match the intended transitions, and the relationship returns to normal after recovery.
- Pending or dropped binding-update counters do not continue to increase after recovery.
- Dynamic DNS updates succeed through either server when DHCP performs them on the client’s behalf.
Do not manufacture a Partner Down condition on a production subnet to see what happens. Use a lab or an approved test scope, with a known free-address margin and a rollback plan. If a server is unreachable because of a partition, investigate that path before assuming the server is powered off. Failover protects address allocation state; it does not guarantee application availability, repair bad scope options, or replace backups of DHCP configuration.
Related:
- Windows DNS Server Operations: Zone Replication, Dynamic Records, and Scavenging
- Fixing DNS Resolution Failures on Windows
Sources: