comm in Shell Pipelines: Compare Sets Only After Sorting Consistently
Use comm for line-set differences only when both inputs share a locale, ordering, record, duplicate policy, and stable input framing.
comm compares two sorted text files line by line and emits records unique to the first input, unique to the second, and common to both. It is a merge operation, not a general unordered set comparator. Both inputs must be sorted according to the same collation order that comm uses. If they are not, output can be misleading or the command may diagnose order only after encountering a problematic record.
The tool compares whole lines, not selected keys. It does not understand CSV fields or structured objects. A set comparison over package identifiers differs from a record join on one field. Normalize and validate input representation before sorting; do not assume visually equivalent Unicode strings compare equal.
Sort with the same locale and policy
Locale affects collation. If inputs were sorted under another LC_COLLATE, comm can see them as out of order. Use a consistent environment for sort and comm:
LC_ALL=C sort -- old.txt > old.sorted
LC_ALL=C sort -- new.txt > new.sorted
LC_ALL=C comm old.sorted new.sorted
This establishes byte-oriented ordering for reproducible machine identifiers. It is not appropriate for every human-language comparison. If the requirement is linguistic ordering, choose and pin the intended locale and use it consistently in every producer and consumer.
GNU comm provides a check-order option that fails when either input is not sorted. It is useful in verification jobs, but is a GNU extension; check target implementation before claiming portability. Validate sort completion before comm consumes each file. If sort fails from a read or disk error, incomplete output may still exist.
Understand columns and exit statuses
With no column-suppression options, comm emits three tab-separated columns: first-file-only, second-file-only, and common records. Options can suppress a column and therefore change output shape. Do not parse columns by blindly splitting every tab if records may contain tabs. For unambiguous machine records, constrain input or use structured representation.
Unlike cmp or diff, comm’s ordinary exit status is not a summary of whether the sets differed; normal completion returns zero. A caller must inspect output or compute counts if it needs a business decision. GNU order-check mode can return nonzero for unsorted input. Do not treat a zero exit code as equality; it only means comparison completed.
If duplicate records appear, comm behavior is not automatically equivalent to mathematical set semantics. Decide whether duplicates should be preserved, counted, or removed. sort -u is convenient for unique lines but discards multiplicity and can alter which representative survives when keys compare equal. Keep counts separately if duplicates are meaningful.
Filename and line framing assumptions
comm treats newline as a record terminator. A record containing newline cannot be represented as one item. Filenames are therefore not safe in ordinary line files unless names are constrained or encoded. Newline-terminated files are also an assumption; GNU documents that a missing final newline can be supplied during processing. Test end-of-file framing explicitly.
If using comm to compare generated manifests, define canonical escaping, case sensitivity, whitespace handling, and duplicate policy. Different serializers can represent the same logical value differently, while text comparison reports them as distinct. Lossy normalization before comm can collapse different logical values.
Safe staging and useful set workflows
For a release manifest, create both sorted files in a private staging area using exactly the same locale and normalization. Check each command’s exit status. Then use comm to generate additions and removals, review the result, and publish only if it matches expected counts and policy. A failed input read must not be interpreted as an empty set.
For large files, sort can spill to temporary storage and fail for lack of space. Place scratch files on a monitored filesystem. Do not overwrite source manifests before verification. Retain original inputs with digests if the comparison is used for audit or incident response.
Test equal files, disjoint sets, one-sided entries, duplicates, unsorted input, locale changes, empty input, tabs in records, missing final newline, and sort-stage failure. Verify column suppression behavior and do not depend on zero exit status to mean equality.
comm is effective when the data really is two ordered lists of line keys. Correctness depends on making order, representation, duplicates, and record boundaries identical on both sides.
Related:
- sort in Shell Pipelines: Locale, Keys, Stability, and Reproducible Output
- jq in Shell Pipelines: JSON, Arguments, Streams, and Exit Status
Sources: