FreeBSD lagg Operations: LACP, Load Distribution, and Failover
Operate FreeBSD lagg with switch-matched LACP, realistic per-flow throughput expectations, member health checks, and controlled failover tests.
FreeBSD’s lagg(4) presents multiple physical interfaces as one logical link for aggregation or failover. It does not automatically double the throughput of every connection, create a redundant upstream switch, or prove that the network beyond the directly connected ports is reachable. A production lagg works only when the host protocol, member link properties, switch configuration, and routed topology agree.
The first design choice is the failure you intend to tolerate. LACP can aggregate compatible physical links with a cooperating switch; failover moves traffic to another port when the active member becomes unavailable; loadbalance hashes traffic across ports without negotiating LACP with the peer. Those modes have different switch requirements and traffic behavior. Changing protocols to “see which one is faster” on a live uplink can cause loops, MAC movement, or lost connectivity.
Establish physical and switch prerequisites
Map each FreeBSD NIC to its MAC address, PCI/driver identity, cable, switch port, speed, and duplex. Confirm the switch’s port-channel or LAG configuration, allowed VLANs, MTU, and LACP state before creating a host aggregate. LACP members must be compatible, including speed and full-duplex operation, and the switch must place the ports in the intended aggregation group. A switch configured for a static port-channel while the host speaks LACP is not the same configuration.
Capture interface and route state first:
ifconfig -a
ifconfig -m igb0
ifconfig igb0
ifconfig igb1
netstat -rn
Do not assume igb0 and igb1 are connected to the intended ports just because their names are adjacent. Check switch MAC tables and port status. If the aggregate carries the only remote-management address, schedule a console-backed change and record how to restore the original standalone interface configuration.
Create a runtime LACP aggregate
Bring physical members up, create the logical interface, select LACP, add the members, and assign the host’s Layer-3 configuration to the aggregate. The following syntax follows the lagg(4) example style; the address is documentation-only:
ifconfig igb0 up
ifconfig igb1 up
ifconfig lagg0 create
ifconfig lagg0 laggproto lacp laggport igb0 laggport igb1
ifconfig lagg0 inet 192.0.2.30/24 up
Do not leave the same address configured on a member and the aggregate. Inspect ifconfig lagg0 for protocol, master/active member, LACP partner state, and per-port collecting/distributing status where shown. On the switch, confirm the same partner system and the intended member set. A member with carrier but no LACP negotiation is not necessarily forwarding as part of the aggregate.
Verify Layer 3 only after Layer 2 state is correct: inspect the connected and default routes, resolve a known neighbor, and test traffic to a peer on the expected VLAN. If the parent is a trunk, place VLAN interfaces on the logical aggregate in the documented topology and coordinate the switch’s tagged VLANs on the LAG. Do not attach the same VLAN subinterface to both physical members independently.
Persist the logical interface
Once runtime behavior is accepted, persist the aggregate through rc.conf(5). The manual documents cloned_interfaces, interface configuration variables, and the laggproto/laggport options. Merge lagg0 into any existing clone list rather than replacing unrelated bridges, VLANs, or tap devices.
# /etc/rc.conf; merge with existing values and use actual NICs/addressing.
cloned_interfaces="lagg0"
ifconfig_igb0="up"
ifconfig_igb1="up"
ifconfig_lagg0="laggproto lacp laggport igb0 laggport igb1 inet 192.0.2.30/24"
The example address belongs to documentation space. For a DHCP-managed interface or a VLAN child, use the appropriate release-documented rc variables rather than copying the static address form. Avoid rewriting /etc/rc.conf by a script that discards preexisting clone values. Review what network service owns the NICs and ensure interface start ordering can create the aggregate before any dependent VLAN or service.
Apply changes in a maintenance window with out-of-band access. Restarting the network service can terminate SSH and disrupt jails, bridges, remote filesystems, VMs, or storage replication. Test the target release’s interface-specific service netif commands in a lab and prepare an explicit rollback to the previous address and route. Validate persistence with a planned reboot only after the runtime path is stable.
Choose a protocol based on its actual semantics
failover sends traffic through the active port, with the first member added as the master and later members available as standby. It is useful when the network expects one active link and another link should take over. The default receive policy generally accepts traffic only on the active member; net.link.lagg.failover_rx_all can relax that behavior for specific bridged setups, but changing it should follow the manual and a topology-specific test.
lacp negotiates aggregation with the peer using IEEE Link Aggregation Control Protocol. It can rebalance member use when connectivity changes, but both endpoints must agree on the LAG. Member speeds and duplex must be compatible. Monitor the partner and collecting/distributing states; carrier-up alone is not sufficient.
loadbalance is a static host configuration. It hashes selected Ethernet and IP headers to choose an outgoing member and does not negotiate an aggregation with the switch. The switch must be configured compatibly, or traffic can be misdelivered. roundrobin spreads outgoing packets in sequence and can cause reordering; it is not a default general-purpose throughput choice. broadcast replicates frames to all members and has different bandwidth and switch-learning implications. Use the exact mode documented for the intended topology, not a benchmark result from an isolated client.
Set honest throughput and resilience expectations
An aggregate distributes independent flows, not arbitrary fragments of one flow across every link. With common hashing, a single TCP connection is usually pinned to one member. Two 1-Gbit/s links therefore do not normally make one TCP flow run at 2 Gbit/s. Many independent flows can use different members, but hash selection, endpoint addresses/ports, and switch behavior determine distribution. A skewed workload can overload one member while another is mostly idle.
Measure per-member counters and aggregate throughput with multiple representative flows. Record client/server CPU, protocol, packet size, direction, and switch counters. Check whether the NIC’s RSS or flow-ID hash is valid for the actual workload; lagg(4) documents that hardware-provided hash behavior can produce poor distribution, and that local hash computation may be selected per interface. Do not change use_flowid or related sysctls without establishing a before/after flow distribution and checking driver support.
Redundancy also has a boundary. lagg can tolerate selected member-link failures; it cannot protect against a shared switch failure, a bad upstream route, an outage of the remote service, or a configuration error common to all links. Two cables into the same failed switch do not create switch redundancy. A meaningful design documents which fault domain each member covers and whether upstream routing, spanning tree, CARP, or another layer supplies additional resilience.
Observe member health and diagnose common failures
Collect ifconfig lagg0, each member’s ifconfig/media state, netstat -I counters, system messages, and the switch’s LACP or port-channel output at the same time. Look for a member disappearing, LACP partner changes, input/output errors, unexpected MTU differences, duplicate MAC learning, or a default route remaining on the wrong interface. A logical interface can be up while one member is down; conversely, all physical carriers can be up while the aggregate protocol has no usable partner.
LACP never reaches collecting/distributing. Check the switch port-channel, LACP mode, member speed/duplex, cabling, allowed VLANs, and whether the ports belong to the same expected partner. Do not change to static load balancing merely to make the host report UP unless the switch configuration is also deliberately changed.
Only one member carries traffic. Confirm the protocol and active state, then inspect the flow hash inputs and number of independent flows. A single connection pinned to one member is expected in many modes. Compare multiple flows and switch counters before concluding that the aggregate is broken.
Traffic drops during failover. Correlate link-state transitions, ARP/neighbor behavior, MAC-table movement, protocol convergence, and switch logs. Confirm that the peer learns the logical address on the new member and that the route/source address remains on lagg0. Test a cable pull on one member only, then restore it and verify the expected active state.
The aggregate works until reboot. Compare cloned_interfaces, member ifconfig_* values, and ifconfig_lagg0 against the tested runtime state. Check that another tool did not overwrite the clone list or reconfigure member addresses. Review startup logs and verify any dependent VLANs are created on the logical interface.
Perform a controlled failover and acceptance test
Before production, test protocol negotiation with the actual switch and representative traffic. Confirm both members are recognized, expected VLANs pass, address and route ownership are correct, and counters increment on the expected links. Then remove one member from service in a planned window, observe application latency and packet loss, and verify the aggregate’s state changes as designed. Restore the member and verify recovery rather than assuming it rejoined cleanly.
For LACP, validate partner system ID, actor/partner state, and collecting/distributing status from both host and switch. For failover mode, test the active-to-standby transition and monitor MAC relearning at the peer. For load-balance, test multiple flows in both directions and inspect distribution; a single-flow throughput benchmark cannot validate aggregate balance. Define a tolerated packet-loss interval and service-level acceptance threshold before testing.
Document the members, switch ports, protocol, VLAN/MTU policy, address owner, failure domains, rollback steps, and monitoring alerts. Alert on logical and member state separately, increasing physical errors, LACP state churn, and unexpected loss of all usable ports. The reliable outcome is not a lagg0 name in ifconfig; it is a negotiated or deliberately static topology whose distribution and failover behavior have been measured end to end.
Related:
- How to Configure Multi-Path TCP and Advanced Routing Tables on FreeBSD
- Diagnosing FreeBSD Ethernet Link Flapping at the Driver Level
Sources: