Skip to content
LinuxHow-To Published Updated 8 min readViews unavailable

Linux SO_TIMESTAMPING: Trace Packet Time Through the Network Stack

Trace Linux packet timing across scheduler, driver, and NIC points; correlate socket timestamp events with packet IDs and their actual clock domain.

Linux socket timestamping can show when packets cross several points in the networking path: when an outgoing packet enters a queueing discipline, when it reaches the driver, when the NIC reports a hardware timestamp, or when an incoming packet enters the kernel receive path. The API is useful for diagnosing queuing delay, measuring packet timing, and building protocols such as Precision Time Protocol (PTP). It does not make timestamps interchangeable or automatically synchronized.

The core design rule is to identify both the timestamp point and its clock domain. A software timestamp taken by the kernel and a raw NIC hardware timestamp can differ because of queueing, driver work, NIC clock offset, frequency drift, and path delay. Subtracting their numeric values without clock conversion can produce an impressive but meaningless latency figure.

Choose the timestamp point you need

SO_TIMESTAMPING accepts a bitmask, not a boolean. Generation flags request timestamp production; reporting flags determine which generated values are delivered to the socket. The distinction matters: asking the driver or NIC to generate a timestamp is separate from enabling delivery of that timestamp to userspace.

On transmit, SOF_TIMESTAMPING_TX_SCHED requests a timestamp before the packet scheduler. SOF_TIMESTAMPING_TX_SOFTWARE requests a software timestamp near the driver handoff, before the packet is passed to the interface. SOF_TIMESTAMPING_TX_HARDWARE asks for a NIC-generated transmit timestamp when the driver and device support it. If scheduler-to-driver time is the question, collect scheduler and software points and compare them in the same clock domain. The scheduler timestamp is particularly useful for identifying queuing delay.

On receive, SOF_TIMESTAMPING_RX_SOFTWARE requests a timestamp after the driver hands the packet to the kernel receive stack. SOF_TIMESTAMPING_RX_HARDWARE requests a receive timestamp from the network adapter. Hardware timestamping requires device-level support and configuration in addition to the socket flags. It may be unavailable for a particular interface, packet type, or driver path.

Software timestamps measure kernel processing milestones, not the instant bits crossed the physical medium. A hardware timestamp is generated at a device-defined point and usually uses that adapter’s PTP Hardware Clock (PHC). Treat each source as a measurement with documented semantics, not a generic “packet time.”

Enable transmit timestamps explicitly

This Linux-only helper requests software scheduler and driver timestamps, asks for a raw hardware timestamp where supported, and enables reporting of software and raw hardware results. OPT_ID supplies an identifier for matching asynchronous results, while OPT_TSONLY asks the kernel to return timestamp records without retaining the full original packet payload in the error-queue record.

#define _GNU_SOURCE
#include <linux/net_tstamp.h>
#include <sys/socket.h>

int enable_packet_timestamps(int fd) {
    int flags = SOF_TIMESTAMPING_TX_SCHED |
                SOF_TIMESTAMPING_TX_SOFTWARE |
                SOF_TIMESTAMPING_TX_HARDWARE |
                SOF_TIMESTAMPING_SOFTWARE |
                SOF_TIMESTAMPING_RAW_HARDWARE |
                SOF_TIMESTAMPING_OPT_ID |
                SOF_TIMESTAMPING_OPT_TSONLY;

    return setsockopt(fd, SOL_SOCKET, SO_TIMESTAMPING_NEW,
                      &flags, sizeof(flags));
}

The call returning success only configures the socket request. It does not prove the NIC produced a timestamp. Read the active hardware configuration, verify the driver’s timestamping capabilities, and check which timestamp slots are populated in returned records. Where per-packet sampling is needed, Linux also supports requesting transmit timestamp generation with a control message on an individual sendmsg() call; the socket still needs the appropriate reporting flags.

New applications should use SO_TIMESTAMPING_NEW and its 64-bit timestamp structure. The kernel documentation warns that the legacy SO_TIMESTAMPING_OLD layout yields incorrect timestamps after 2038 on 32-bit systems. Check the UAPI headers and libc/kernel compatibility for the deployment target, especially when building for older distributions or multiple architectures.

Read the transmit error queue

Transmit timestamp records are delivered asynchronously on the socket error queue. Consume them using recvmsg() with MSG_ERRQUEUE, and poll for POLLERR when waiting for a record; this queue read itself is nonblocking. Provide enough control-message space for both the extended error information and timestamp data. Check MSG_CTRUNC and reject incomplete ancillary data rather than interpreting truncated structures.

The timestamp ancillary message uses SOL_SOCKET and SCM_TIMESTAMPING. The associated IPv4 or IPv6 extended-error control message carries sock_extended_err; for timestamp notifications, the kernel documents ee_errno as ENOMSG. Use ee_info to determine which timestamp event was reported. SOF_TIMESTAMPING_OPT_ID provides an ID in ee_data, which is important because packet scheduling can reorder transmit events relative to the order of send() calls. Correlate using the ID, not by assuming the next timestamp belongs to the most recent send.

The new-format scm_timestamping64 record has three timestamp slots. ts[0] commonly holds a software timestamp, ts[2] holds raw hardware time, and ts[1] is deprecated. At least one slot is nonzero, but a requested timestamp may be absent if the source did not provide it. Do not treat a zero field as a valid time or assume every event includes both software and hardware values.

For receive timestamping, ancillary data accompanies normal packet data returned by recvmsg(); it is not read through the error queue. A single socket that handles both directions therefore needs separate parsing paths and enough control buffer capacity on each. Validate cmsg_level, cmsg_type, and cmsg_len, and handle MSG_CTRUNC before accessing the payload. Keep error-queue draining independent of ordinary packet reads so timestamp records do not accumulate and consume socket receive budget.

Configure and compare the NIC clock

Start with read-only capability and configuration queries:

ethtool -T "$IFACE"
ethtool --get-hwtimestamp-cfg "$IFACE"

ethtool -T reports timestamping capabilities and associated PHC information when supported by the tool and driver. The hardware timestamp configuration query reports the current device configuration where the netlink interface is implemented. The exact output varies by ethtool release and hardware. A failed query is not proof that software timestamping is unavailable; it may instead mean the driver or tool lacks that hardware interface.

Changing hardware timestamp filters can affect other applications using the same interface. The kernel documentation says only an administrator may change device timestamp configuration and warns that userspace must coordinate because processes can interfere with one another. Drivers may accept a more permissive configuration than requested, or reject a filter the device cannot support. Read back the actual configuration and restore the intended shared-device setting after a test.

Raw hardware timestamps use the NIC PHC time base. The kernel no longer uses the second timestamp slot for converted system time; applications that need to compare PHC and system clocks should use the PTP Hardware Clock interfaces and a clock synchronization/conversion strategy. CLOCK_REALTIME, CLOCK_MONOTONIC, and PHC time are not interchangeable merely because all are represented as seconds and nanoseconds. Record which clock produced every measurement and how it was synchronized.

Separate queueing, driver, and wire effects

Use TX_SCHED and TX_SOFTWARE to estimate time spent in the kernel’s transmit queue after the scheduling point. The scheduler-to-software difference can expose queueing and device/driver delays, but it is not automatically a pure “NIC latency” measurement. Multiple stacked devices can produce more than one scheduler timestamp along a path, and hardware timestamping can be implemented by an outer switch or another device rather than the host port.

SOF_TIMESTAMPING_TX_COMPLETION is a different measurement: it records when the kernel receives a transmit-completion report from hardware. A device may report multiple packets together, so completion time can reflect the report batch rather than the packet’s exact on-wire time. It should not be substituted for the hardware transmit timestamp when measuring wire timing.

On receive, compare software and hardware points only after confirming clock conversion and packet association. Hardware timestamp requests can be filtered or generalized by a driver, and a receive hardware timestamp may be attached to a packet with a different metadata path than expected. For PTP or precision measurement, check the PHC, the interface/port topology, timestamp filter, synchronization state, and packet timestamp identifier rather than relying on wall-clock display alone.

Keep sampling overhead bounded

Timestamping adds work in the networking stack, the device driver, and userspace control-message processing. Enabling receive timestamp generation can affect packets early in the receive path before their destination socket is known. For high packet rates, measure CPU use, queue depth, drops, socket buffer pressure, and application latency with timestamping disabled and enabled. Sample selected transmit operations using per-message control flags when full-rate telemetry is unnecessary.

Drain the error queue promptly. Without OPT_TSONLY, transmit records can carry the original packet data and consume socket receive memory until userspace reads them. If you need ancillary data such as packet-info metadata, check which reporting options can be combined; OPT_TSONLY intentionally omits the original payload needed for some control messages.

Do not treat timestamp precision as accuracy. Nanosecond representation does not mean the capture point is accurate to one nanosecond. The actual uncertainty depends on hardware, driver, clock synchronization, timestamp filter, packet path, and the semantics of the selected event. Document these uncertainties in any latency SLO or benchmark.

Troubleshoot missing or misleading records

If no transmit events arrive, verify that generation and reporting flags are both enabled, drain the error queue with MSG_ERRQUEUE, allocate adequate control-message space, and check ancillary truncation. Then inspect the socket option return code, interface capability, hardware timestamp configuration, driver, and PHC. A software event can be available while hardware timestamps are not.

If transmit timestamps arrive out of order, use OPT_ID and ee_data to match records to sends. If you see only one timestamp slot, check which event the extended error reports and whether the device generated the other requested point. If wall-clock subtraction produces large or negative latency, first check for mixed clock domains, PHC synchronization, and clock steps; do not “fix” the result by clamping negative values.

A useful test sends numbered packets at a controlled rate, records the userspace send time in an explicitly named clock domain, collects scheduler/software/hardware events, and stores the identifier and ee_info with each observation. Repeat at idle and under queue pressure. Compare distributions, not only averages, and preserve raw values so a later clock-correlation correction is possible.

Linux socket timestamping is most valuable when the measurement contract is explicit: which event is requested, which clock measures it, how it is correlated to the send or receive, and what device state is required. Establish those details before using the numbers to tune a queue or claim wire latency.

Related:

Sources:

Comments