Linux SCM_RIGHTS: Secure File-Descriptor Handoff over Unix Sockets
Pass Linux file descriptors with SCM_RIGHTS using bounded ancillary buffers, close-on-exec reception, peer checks, object validation, and explicit ownership.
SCM_RIGHTS lets cooperating processes transfer access to an already-open kernel object through an AF_UNIX socket. The receiver does not receive the sender’s integer descriptor number. The kernel installs a new descriptor in the receiver’s descriptor table that refers to the same open file description, much like dup() performed across process boundaries.
That is a capability transfer, not just a message containing a path. A received descriptor may grant access the receiving process could not obtain by opening the same pathname itself. It may also share a file offset and file status flags with the sender’s descriptor. Treat the sender, protocol, descriptor type, lifetime, and allowed operations as part of the security boundary.
Prefer a message-oriented local socket protocol
Linux supports descriptor passing with sendmsg() and recvmsg() over Unix-domain sockets. A connected SOCK_SEQPACKET channel is convenient when the application needs records: it preserves message boundaries and ordering. A SOCK_STREAM channel is also common, but it is a byte stream. Its ancillary data is associated with a position in that stream, not with an application-defined packet, so the receiver must combine descriptor handoff with a carefully framed protocol.
For Linux Unix-domain SOCK_STREAM, send at least one byte of ordinary payload in the same sendmsg() call as the ancillary data. A one-byte record marker works for a dedicated handoff message. Do not assume that one sendmsg() maps to one recvmsg() on a stream; a receiver may observe earlier stream bytes together with the byte associated with the descriptor. If the application needs simple request boundaries, SOCK_SEQPACKET can make the protocol easier to reason about.
Build one bounded SCM_RIGHTS message
Control messages use struct cmsghdr records. Use CMSG_SPACE() for the buffer size, CMSG_LEN() for cmsg_len, and the CMSG_* macros to find and access headers. Zero-initialize the control buffer and copy descriptor integers with memcpy(); CMSG_DATA() is not guaranteed to have alignment suitable for casting to every payload type.
The following sender expects an already-connected Unix SOCK_SEQPACKET socket. It transfers exactly one descriptor and one marker byte. It does not close the sender’s original descriptor.
#define _GNU_SOURCE
#include <errno.h>
#include <string.h>
#include <sys/socket.h>
int send_one_fd(int socket_fd, int fd_to_send) {
char marker = 'F';
struct iovec iov = {
.iov_base = &marker,
.iov_len = sizeof(marker),
};
union {
struct cmsghdr align;
unsigned char bytes[CMSG_SPACE(sizeof(fd_to_send))];
} control = {0};
struct msghdr msg = {0};
msg.msg_iov = &iov;
msg.msg_iovlen = 1;
msg.msg_control = control.bytes;
msg.msg_controllen = sizeof(control.bytes);
struct cmsghdr *cmsg = CMSG_FIRSTHDR(&msg);
if (cmsg == NULL) {
errno = EIO;
return -1;
}
cmsg->cmsg_level = SOL_SOCKET;
cmsg->cmsg_type = SCM_RIGHTS;
cmsg->cmsg_len = CMSG_LEN(sizeof(fd_to_send));
memcpy(CMSG_DATA(cmsg), &fd_to_send, sizeof(fd_to_send));
ssize_t sent;
do {
sent = sendmsg(socket_fd, &msg, MSG_NOSIGNAL);
} while (sent == -1 && errno == EINTR);
if (sent == -1)
return -1;
if (sent != (ssize_t)sizeof(marker)) {
errno = EIO;
return -1;
}
return 0;
}
The protocol must define what F means and which descriptor is expected. A successful sendmsg() means the kernel accepted the message for the socket; it does not prove that the peer received, validated, or adopted the descriptor. If the sender must know that the receiver accepted ownership, include a request identifier and an explicit acknowledgment in the higher-level protocol.
MSG_NOSIGNAL prevents this send operation from raising SIGPIPE when the peer has closed the connection; the call still reports the failure through its return value and errno. Applications targeting other Unix systems should use that platform’s documented socket-signal policy rather than assuming this Linux flag exists everywhere.
Receive atomically with close-on-exec and bounded cleanup
On Linux, MSG_CMSG_CLOEXEC asks recvmsg() to set close-on-exec on received descriptors as they are installed. This closes an important race in multithreaded programs where another thread may call execve() between receipt and a later fcntl(F_SETFD) call. The flag is Linux-specific; use a platform-specific equivalent or a documented fallback when portability is required.
This receiver accepts exactly one descriptor referring to a regular file, one marker byte, and no other ancillary message. It reserves room for a bounded descriptor set so it can close descriptors delivered with a malformed or over-sized message before rejecting it. A real service should select the bound from its protocol and resource policy.
#define _GNU_SOURCE
#include <errno.h>
#include <fcntl.h>
#include <stddef.h>
#include <string.h>
#include <sys/socket.h>
#include <sys/stat.h>
#include <unistd.h>
enum { MAX_RECEIVED_FDS = 16 };
static void close_all(int *fds, size_t count) {
for (size_t i = 0; i < count; i++)
close(fds[i]);
}
int receive_one_regular_file(int socket_fd) {
char marker = 0;
struct iovec iov = {
.iov_base = &marker,
.iov_len = sizeof(marker),
};
union {
struct cmsghdr align;
unsigned char bytes[CMSG_SPACE(sizeof(int) * MAX_RECEIVED_FDS)];
} control = {0};
struct msghdr msg = {0};
int received[MAX_RECEIVED_FDS];
size_t received_count = 0;
size_t rights_messages = 0;
int invalid = 0;
msg.msg_iov = &iov;
msg.msg_iovlen = 1;
msg.msg_control = control.bytes;
msg.msg_controllen = sizeof(control.bytes);
ssize_t n;
do {
n = recvmsg(socket_fd, &msg, MSG_CMSG_CLOEXEC);
} while (n == -1 && errno == EINTR);
if (n == -1)
return -1;
for (struct cmsghdr *cmsg = CMSG_FIRSTHDR(&msg);
cmsg != NULL;
cmsg = CMSG_NXTHDR(&msg, cmsg)) {
if (cmsg->cmsg_level != SOL_SOCKET || cmsg->cmsg_type != SCM_RIGHTS) {
invalid = 1;
continue;
}
rights_messages++;
if (cmsg->cmsg_len < CMSG_LEN(0)) {
invalid = 1;
continue;
}
size_t bytes = cmsg->cmsg_len - CMSG_LEN(0);
if (bytes % sizeof(int) != 0)
invalid = 1;
size_t count = bytes / sizeof(int);
for (size_t i = 0; i < count; i++) {
int received_fd;
memcpy(&received_fd, CMSG_DATA(cmsg) + i * sizeof(received_fd),
sizeof(received_fd));
if (received_count < MAX_RECEIVED_FDS) {
received[received_count++] = received_fd;
} else {
close(received_fd);
invalid = 1;
}
}
}
if (n != (ssize_t)sizeof(marker) || marker != 'F' ||
(msg.msg_flags & (MSG_CTRUNC | MSG_TRUNC)) != 0 ||
rights_messages != 1 || received_count != 1) {
invalid = 1;
}
if (invalid) {
close_all(received, received_count);
errno = EPROTO;
return -1;
}
struct stat st;
if (fstat(received[0], &st) == -1) {
int saved_errno = errno;
close(received[0]);
errno = saved_errno;
return -1;
}
if (!S_ISREG(st.st_mode)) {
close(received[0]);
errno = EINVAL;
return -1;
}
return received[0]; /* Ownership transfers to the caller. */
}
The example intentionally rejects unexpected ancillary records and more than one descriptor. MSG_CTRUNC means some control data was discarded because the buffer was too small; it is a protocol failure, not a reason to use whichever descriptor happened to fit. On Linux, descriptors omitted because of truncation or descriptor-table limits are automatically closed by the kernel, while the receiver must close every descriptor that was installed and delivered in the returned control data before rejecting the message.
The sample validates only that the received object is a regular file. That is not sufficient authorization for a privileged service. A production receiver may need to verify access mode with fcntl(F_GETFL), ownership and mode from fstat(), expected seals for a shared-memory object, a permitted device number, or application-specific metadata. Do not trust a filename, integer descriptor number, or sender-provided type string as proof of the object’s properties.
Authenticate the peer and scope transferred authority
For a connected Unix stream socket or socket pair, Linux SO_PEERCRED returns the peer credentials captured when the connection was established. SO_PASSCRED enables per-message SCM_CREDENTIALS ancillary data. These mechanisms can support a local authorization policy, but credential checks must match the threat model: a UID or PID is not a cryptographic identity, PIDs can be reused, privileged senders have additional credential-setting authority, and a credential result does not validate the descriptor itself.
Protect filesystem socket paths with a trusted parent directory and deliberate ownership and permissions. Abstract-namespace sockets do not have filesystem permission checks. Authenticate the peer before accepting authority-bearing descriptors, then apply a narrow allowlist for descriptor type and allowed use. If the service should not let the peer choose arbitrary files, accept only descriptors created through a trusted broker or require properties the receiver can independently verify.
Passing a descriptor shares a reference to an open file description. For regular files, reads and writes can share the current file offset. File status flags such as O_NONBLOCK are also associated with the open file description. FD_CLOEXEC, by contrast, is a descriptor flag, which is why the receiver should request MSG_CMSG_CLOEXEC and still document its close policy. Two processes can therefore interfere through a shared offset or flag even if they use different integer descriptor values.
Define ownership, limits, and failure behavior
The sender normally retains its original descriptor after sending; the receiver obtains another reference. The sender may close its copy when its own use is finished, but the underlying object remains available while any reference remains. The receiver owns every descriptor installed by recvmsg() and must close it on all rejection and shutdown paths. Keep the transfer state machine explicit: pending, delivered, accepted, rejected, and closed are different states.
Bound the number of descriptors per message and the number of in-flight handoffs. Descriptor installation is subject to RLIMIT_NOFILE and other kernel resource limits. A peer can otherwise consume file table capacity by sending many descriptors faster than the receiver validates and closes them. Use backpressure, per-peer quotas, timeouts, and metrics for accepted, rejected, truncated, and leaked-reference conditions.
A successful transfer does not make the object immutable, durable, or private. If the sender can still modify a shared file, the receiver must account for concurrent changes. For immutable memory-backed data, complete the write, apply appropriate seals, and only then transfer the descriptor. For persistent files, keep synchronization and crash-durability steps separate from the IPC handoff.
Test the protocol as an adversarial boundary
Test the exact socket type and Linux versions you deploy. Include a valid single-descriptor transfer, peer disconnect, EINTR, short or malformed payload, unexpected control message, multiple descriptors, a control buffer large enough and too small, MSG_CTRUNC, descriptor-table exhaustion, and a receiver that immediately executes another program. Verify that every rejected descriptor is closed and that no caller adopts a descriptor before its peer, object type, and policy checks succeed.
Instrument both ends with a request ID, peer identity, expected descriptor role, transfer outcome, and close result. Never log the raw descriptor integer as a cross-process identity; it is local to one process. Record whether acknowledgment was received, and distinguish kernel delivery from application acceptance. A descriptor handoff is complete only when the protocol’s ownership and cleanup rules say it is complete.
SCM_RIGHTS is a powerful local IPC primitive because it transfers a live kernel reference instead of asking the receiver to reopen a path. Use ancillary-data macros, set close-on-exec atomically, authenticate and authorize the peer, validate the object independently, and close every rejected reference. Those measures turn descriptor passing from an opaque side channel into a bounded capability protocol.
Related:
- Linux memfd File Seals: Immutable Shared Data Without a Pathname
- Linux pidfds: Race-Free Process Handles Beyond Numeric PIDs
Sources: