tr in Shell Pipelines: Character Sets, Locales, and Byte-Safe Transformations
Use tr for character translation, squeezing, and deletion without assuming Unicode normalization, multibyte portability, or arbitrary-byte semantics.
tr reads standard input and translates, squeezes, or deletes characters according to two character sets. It is useful for simple byte- or character-level transformations, but it is not a Unicode normalization engine, a regular-expression processor, or a general string replacement language. Locale and implementation determine what a character set means, particularly for multibyte text.
Treat operands as sets and mappings with defined semantics. Translation maps positions in the first set to positions in the second; if lengths differ, standard and implementation rules determine how the final character is handled. Squeeze collapses repeated members of a selected set. Delete removes members. Combining these modes changes the data contract and should be tested with representative input.
Character sets are not regular expressions
Bracket expressions and character classes may be supported, but tr does not interpret an arbitrary regex. A range such as a-z can depend on locale ordering; in a non-C locale it may include a wider collation sequence than ASCII letters. If the requirement is strictly ASCII, set LC_ALL=C around the command and document why. If the input is human language, do not force the C locale without understanding consequences.
For example, normalize ASCII uppercase letters:
LC_ALL=C tr '[:upper:]' '[:lower:]' < input.txt
This uses the selected locale’s character classes. It does not perform full Unicode case folding, normalization, or locale-sensitive linguistic casing. Characters outside the class may pass through unchanged.
The complement option can be particularly surprising in multibyte locales. GNU documentation cautions that complement semantics are not always clear or portable for multibyte characters. Use it only with tested input and locale assumptions. For UTF-8 transformations that need Unicode properties, use a Unicode-aware tool or library.
Squeeze and delete are separate operations
The -s option squeezes runs of repeated characters listed in the last specified character set. It does not replace arbitrary repeated substrings. Delete removes all characters in its set and can join text that previously had a delimiter between tokens. When both are used, the precise order and set to squeeze are part of the interface; read the selected utility manual rather than infer behavior from option names.
Removing control characters from logs can destroy framing or conceal evidence. Before stripping bytes, determine whether they carry line endings, terminal escape sequences, NUL sentinels, or format separators. A lossy sanitizer should log what it removed and preserve original input when forensic traceability matters.
Preserve NUL and error behavior
Many text utilities are line-oriented and have implementation-specific behavior for NUL bytes. Do not assume tr is suitable for arbitrary binary data merely because a pipeline accepts bytes on standard input. Test exact versions and byte sequences or choose a binary-aware tool.
Quote set operands so the shell does not expand brackets or wildcard characters. If a set is built from external input, do not concatenate it into shell syntax without escaping and validation. A fixed expression is easier to audit. Keep input and output channels separate and avoid command substitution for data that may contain NUL or trailing newlines.
As a pipeline stage, tr can report read or write errors. A successful downstream process does not necessarily reveal that failure in a POSIX shell pipeline. Use a supported aggregate status policy or stage output and check each step. If the transformation canonicalizes identifiers before a security comparison, ensure both sides use the same locale, version, and normalization algorithm.
Test the exact transformation
Create fixtures with repeated characters, characters outside the set, leading and trailing bytes, locale-sensitive ranges, UTF-8 multibyte sequences, tabs, line endings, NUL bytes if allowed, and a final record without newline. Compare exact byte output under every supported locale. Test translate, delete, squeeze, complement, and combinations independently.
tr is a reliable small primitive when the data model is explicit. Document whether input is ASCII, locale text, or bytes; define output encoding; keep transformations intentionally lossy; and choose a Unicode-aware parser when correctness extends beyond the selected utility’s contract.
Related:
- cut in Shell Scripts: Delimiters, Character Sets, and Record Assumptions
- AWK Records and Fields: Parse Text Without Pretending It Is CSV
Sources: