Linux tmpfs in Production: Capacity Limits, Swap, Inodes, and cgroups
Operate Linux tmpfs deliberately by sizing bytes and inodes, accounting for swap and cgroup limits, monitoring pressure, and protecting volatile data.
Linux tmpfs is useful for volatile files, shared memory, scratch data, and temporary application state, but its name can create the wrong operational model. It is a memory-backed filesystem, not a fixed RAM disk and not durable storage. It grows as files are written, can use swap when permitted, and loses its files when unmounted. Each mount can have a byte ceiling and an inode ceiling, while the host and a cgroup impose separate memory and swap constraints. A production limit is safe only when these layers are considered together.
This distinction matters for /dev/shm, build scratch space, runtime caches, container writable paths, and services that stream large temporary files. A mount that looks roomy in df can still compete with application memory or hit a cgroup limit. Conversely, a tmpfs size limit is only a maximum allocation for that mount; it does not reserve that amount of RAM and does not guarantee that the host can satisfy writes up to that ceiling.
What tmpfs stores, and what it does not promise
The kernel keeps tmpfs file contents in its internal caches and can swap unneeded pages to swap space when that is enabled. The filesystem does not write its files to a normal disk-backed filesystem, but pages may still be paged out to swap. Unmounting the filesystem discards its contents. A reboot, crash, mount replacement, or container teardown can therefore destroy data that an application failed to copy elsewhere.
Do not use tmpfs as the sole location for a database, queue, checkpoint, recovery journal, or any data that must survive a restart. A fast write is not a durable commit. Before redirecting an application directory to tmpfs, confirm the application’s restart behavior, whether the path is expected to persist, and which other copy is authoritative. A cache can usually be rebuilt; a spool, upload, or checkpoint may not be disposable just because it lives under a directory named tmp.
tmpfs is also different from ramfs. tmpfs has configurable size and inode limits and can use swap; ramfs lacks those safety limits and does not use swap. A tmpfs mount at /dev/shm is commonly used for POSIX shared memory and semaphores. That visible mount is distinct from the kernel’s internal shmem mechanisms, so system-wide shared-memory counters do not map one-to-one to a single tmpfs mount.
Set both byte and inode ceilings
The size= option sets the maximum bytes that an individual tmpfs instance may allocate. The nr_inodes= option separately caps how many inodes it may create. These limits protect against different exhaustion modes: a workload with many small files can exhaust inodes before reaching its byte limit, while a few large files can exhaust bytes with many inodes still free. Extended attributes also consume inode-limit accounting, so a high df -i count does not necessarily mean the mount contains the same number of ordinary files.
When neither size nor nr_blocks is supplied, current Linux man-page documentation describes the default byte limit as 50 percent of physical RAM. Do not use that default as a service budget. Defaults, visible memory, and container limits may differ across kernel builds and deployment environments. Several independent mounts can each have a generous ceiling, but all of them still compete for real memory, swap, and any applicable cgroup budget. Setting size=0 or nr_inodes=0 removes the corresponding limit; that is rarely an appropriate safety setting for a writable production mount.
Choose explicit limits from measured peak usage plus a documented margin. Include temporary bursts, concurrent jobs, cleanup lag, largest expected file, and file-count growth. Then verify that the combined workload remains within the host’s memory and swap policy, the service’s cgroup controls, and the mount’s own byte and inode ceilings. Do not simply set every tmpfs ceiling to a large fraction of host RAM: that creates an overcommit promise which multiple mounts and ordinary process memory may exceed together.
An illustrative /etc/fstab entry for a private, disposable application cache is:
tmpfs /run/example-cache tmpfs rw,nosuid,nodev,noexec,size=768M,nr_inodes=120k,mode=0700 0 0
The numbers are examples, not sizing advice. mode=0700 makes the mount root accessible only to its owner; provision ownership intentionally if the service runs as a non-root account. nosuid, nodev, and noexec may reduce risk for a data-only cache, but they can break software that expects executable files, device nodes, or set-ID behavior there. Validate application requirements and the exact mount options before rollout. For a shared scratch mount, access mode and sticky-directory behavior must be designed deliberately rather than copied from this private-cache example.
Find the actual mount and read its limits
Use the mount table and both df views to inspect a specific path. The first df invocation examines block usage; df -i examines inode usage.
findmnt --target /run/example-cache --output TARGET,SOURCE,FSTYPE,OPTIONS
df --human-readable /run/example-cache
df --inodes /run/example-cache
du --one-file-system --summarize --human-readable /run/example-cache
These commands answer different questions. df reports filesystem-wide allocated and available capacity. du walks visible files and can be slow or incomplete if permissions deny traversal or if files change during the scan. findmnt confirms that the path is actually backed by the expected tmpfs mount; if the mount failed or disappeared, a program may be writing into the underlying disk-backed directory instead. Do not interpret an empty directory on that underlying filesystem as proof that the temporary data was safely mounted elsewhere.
/proc/meminfo’s Shmem counter and the Shared figure from free include tmpfs pages, but they also include other shared-memory uses. They are host-level context, not a per-mount usage report. In cgroup v2, memory.stat distinguishes file and shmem accounting, and memory.current reports the cgroup’s total memory usage. Those values are useful for correlating filesystem growth with service pressure, but they do not replace df when the question is how close the tmpfs mount is to its own configured limit.
For a quick host and cgroup snapshot:
grep -E '^(MemAvailable|SwapFree|Shmem):' /proc/meminfo
cat /sys/fs/cgroup/example.service/memory.current
grep -E '^(file|shmem) ' /sys/fs/cgroup/example.service/memory.stat
cat /sys/fs/cgroup/example.service/memory.swap.current
The cgroup paths shown are examples. systemd, containers, and delegated cgroup hierarchies may place the service somewhere else; first resolve the process’s actual cgroup and confirm the cgroup v2 controllers are available. memory.current includes descendants, and memory statistics describe the cgroup rather than one mount. In a cgroup v1 deployment, the files and accounting model differ. Avoid monitoring scripts that assume a single hard-coded service path across all hosts.
Budget memory and swap independently
The tmpfs size= limit is not a reservation, and it is not the same as a cgroup’s memory.max or memory.high. A write may be constrained by the tmpfs ceiling, cgroup memory pressure, swap limits, or host-wide pressure, depending on the environment. A container can hit its memory budget while the tmpfs still has free capacity. A tmpfs can reach its mount cap while the container has unused memory. Monitoring only one of these numbers misses the other failure mode.
Under cgroup v2, inspect memory.current, the file and shmem fields in memory.stat, memory events, and swap usage together. shmem represents swap-backed cached filesystem data such as tmpfs, shared-memory segments, and shared anonymous mappings. memory.swap.current reports swap used by that cgroup and descendants. Kernel and system-manager versions affect the available controls, so deploy and test the correct interface for the host. Do not infer that tmpfs is excluded from service memory pressure because its files are not on an ordinary disk filesystem.
Swap behavior is an explicit design choice. With normal swap use, tmpfs pages can leave RAM under pressure, but that can add latency and can place data on swap storage. Since Linux 6.4, the noswap mount option can disable swap use for a tmpfs mount on kernels that support it. Verify the running kernel’s documentation and mount output before depending on that option; it is not available on every maintained distribution kernel, and remount behavior must respect the mount’s original setting. Disabling swap may improve predictability for a particular workload but increases pressure on RAM and does not turn tmpfs into durable or universally protected memory.
This distinction also matters for sensitive data. Ordinary tmpfs is not by itself a guarantee that bytes never reach nonvolatile storage: swap, hibernation, crash dumps, and other system-level mechanisms need their own policy. If the requirement is “this content must not be written to disk,” document and test the full host configuration rather than relying on the filesystem name or a single mount flag.
Change limits without risking the contents
Tmpfs can be remounted with a different size limit without discarding existing contents, but the new limit cannot be below current usage. Plan a reduction only after measuring the mount and reducing its content safely. Do not use a remount to paper over sustained growth; identify whether the workload has a leak, cleanup failure, changed concurrency, or a too-small capacity plan.
An approved increase to an already mounted tmpfs can be applied with a remount, for example:
sudo mount -o remount,size=1024M /run/example-cache
findmnt --target /run/example-cache --output TARGET,FSTYPE,OPTIONS
df --human-readable /run/example-cache
This changes the mount’s maximum, not the amount of physical memory reserved for it. Use the service’s change controls, verify the current source and filesystem type first, and confirm that the intended mount options remain active afterward. The kernel accepts resizing, but capacity governance still belongs to the operator.
Inode capacity can be exhausted independently. Track df -i, file creation rate, and cleanup behavior for workloads that create many tiny files, unpack archives, or generate per-request scratch files. Raising nr_inodes without reviewing memory pressure and application behavior may only postpone a leak. If many stale files accumulate, fix lifecycle cleanup and test it under failure and restart conditions instead of expanding limits indefinitely.
Diagnose pressure without confusing counters
When an application reports ENOSPC, first determine whether the full resource is bytes or inodes. Compare df and df -i for the exact path, then confirm findmnt shows the intended tmpfs mount. Inspect the process’s cgroup memory and swap files if its writes fail before the mount ceiling is reached. Check kernel logs and the service manager’s OOM and memory-pressure events. A host’s aggregate Shmem number alone cannot identify which mount or workload is responsible.
If df shows tmpfs capacity available but a workload is stalled, investigate memory pressure, cgroup throttling, swap activity, and the backing application’s write pattern. If df shows the mount full, find which directories and files account for usage before deleting anything. Removal is destructive for the application even on tmpfs; confirm file ownership and lifecycle. For a deleted-but-open file, freeing a directory entry may not reclaim the space until the writer closes the file, so inspect the owning process and restart behavior carefully.
Before production use, load-test the largest expected file, concurrent writers, inode-heavy cases, cleanup after interruption, and host/cgroup pressure scenarios. Alert separately on byte utilization, inode utilization, cgroup memory pressure, and swap use. Set thresholds early enough to allow a safe response, and record whether the service should reject work, spill to durable storage, or shed load when temporary capacity is exhausted.
Operational acceptance checklist
Treat a tmpfs-backed path as disposable only after the application owner confirms that its contents are reconstructible. Record the intended mount point, filesystem type, byte and inode limits, ownership, permissions, optional security flags, swap policy, service cgroup, and expected cleanup behavior. At boot and during deployment, verify that the mount exists before starting writers. At shutdown, determine whether data may be discarded or must first be flushed elsewhere.
The production claim should be concrete: what is allowed to disappear, what capacity and inode budget is guaranteed by policy, which controller can stop allocations first, and what alert fires before that boundary is reached? Tmpfs is a useful tool when those answers are explicit. It is an unsafe surprise when “in memory” is mistaken for unlimited, isolated, persistent, or always resident in RAM.
Related:
- Linux cgroup v2 memory.reclaim: Controlled Proactive Reclaim
- Linux zswap: Operate the Compressed Swap Cache Under Memory Pressure
Sources: