Skip to content
Shell & TerminalFix Published Updated 9 min readViews unavailable

Fixing Locale and Encoding Failures in Shell Tools

Mojibake, invalid byte errors, and inconsistent sorting usually come from a locale mismatch - diagnose bytes and locale variables before converting anything.

When a shell tool rejects text as an invalid character sequence, displays mojibake, or sorts the same file differently on two systems, separate three layers before changing data: the file’s bytes, the process locale, and the terminal’s rendering. They are related, but they are not interchangeable. A UTF-8 locale does not convert an old file to UTF-8; a correct byte sequence does not guarantee the terminal font can display every glyph; and changing LANG does not repair text already written with the wrong encoding.

Start with evidence, not a conversion

Record the environment and identify a small failing sample. Do not paste confidential text into an online “encoding detector.” On Unix-like systems, these commands are useful first probes:

locale
locale -a
file --mime sample.txt
od -An -tx1c -N 96 sample.txt

locale prints the current process’s locale settings. locale -a lists names recognized by that implementation, although its output and available aliases vary by operating system. file guesses a file type or encoding from its contents; it is a heuristic, not proof. od shows bytes, which helps identify a byte-order mark, NUL bytes, ASCII text, or a suspected UTF-8 sequence. The --mime spelling for file is not universal; some systems provide a different option, so consult that host’s file(1) manual if it rejects the flag.

Capture the commands’ output from the environment that fails. A terminal, cron job, SSH command, CI runner, editor task, and system service can inherit different variables. Compare LANG, LC_ALL, the relevant LC_* values, and the exact command line. If possible, reduce the failure to a tiny file with a few representative characters and a known expected result. This avoids confusing a parsing bug, malformed input, and a display problem.

Understand what locale variables control

For POSIX locale selection, LC_ALL takes precedence over category-specific variables such as LC_CTYPE and LC_COLLATE; those categories in turn take precedence over LANG. An empty variable is not an explicit locale value. The selected categories affect different behavior: LC_CTYPE is relevant to character classification and multibyte handling, while LC_COLLATE affects collation and therefore many sorts and pattern ranges. Other categories affect number formatting, dates, monetary output, and messages.

GNU gettext has an additional LANGUAGE variable for message-language selection; it is a gettext-specific exception and does not replace locale settings for character encoding or collation. Avoid saying simply that one variable always controls “the locale” without stating the program and category. The process environment is input to locale-aware software; a program may set a locale itself, ignore environment variables in privileged contexts, or implement its own Unicode and sorting rules.

To inspect one command under an explicit locale without changing the parent shell, use a command-scoped assignment:

LC_ALL=C sort names.txt
LC_CTYPE=en_US.UTF-8 tool-that-needs-utf8 input.txt

These are examples, not portable locale names. Check availability first. C or POSIX is useful when the desired behavior is the standard baseline rather than language-specific collation. It is not a universal “fix Unicode” switch. A UTF-8 locale name such as en_US.UTF-8 may be unavailable or spelled differently; C.UTF-8 exists on some systems and not all. Set only the category you need when you want other locale behavior to remain in effect. LC_ALL overrides every category, so it is usually best scoped to one command rather than exported permanently in a login file.

Distinguish bytes, encoding, and display

UTF-8 is a byte encoding for Unicode scalar values. ASCII bytes retain their familiar values, while many other characters use multi-byte sequences. A program operating under a UTF-8 locale can validate or interpret sequences according to its library and policy, but the locale does not rewrite the file. A sequence that is invalid UTF-8 remains invalid input even if LANG is changed. Conversely, a legacy single-byte file can be valid in its source encoding and still appear as mojibake when a viewer interprets those bytes as UTF-8.

The same visible text may have different byte representations in different Unicode normalization forms. If a search or comparison behaves unexpectedly for accented text, inspect the actual byte sequence and the data producer before normalizing. Unicode defines normalization forms, but normalization is a content transformation and can affect identifiers, signatures, filenames, or application-specific equality; it should not be applied casually as a terminal repair.

file and similar detectors cannot reliably infer every encoding. ASCII-only text is valid in many ASCII-compatible encodings, so the bytes alone cannot reveal which encoding the producer intended. Some byte sequences are valid under more than one encoding but mean different characters. Ask the producer, check a file format specification, or use a known metadata contract. A successful conversion is not proof that the chosen source encoding was correct; conversion tools may map the same byte differently depending on the declared source.

For a small sample, inspect both bytes and a locale-aware rendering. Keep the original immutable while experimenting. Avoid pipelines that silently replace invalid input, drop bytes, transliterate characters, or normalize text before you have decided that behavior is acceptable. For production data, record source encoding and conversion policy alongside the import process.

Convert only from a known source encoding

Use iconv with both encodings named explicitly. Write to a new destination and check the exit status before publishing it:

#!/bin/sh
set -eu

source=legacy.txt
destination=legacy.utf8.txt

if [ -e "$destination" ]; then
    printf 'refusing to overwrite %s\n' "$destination" >&2
    exit 1
fi

# Set noclobber so a concurrent or unexpected existing file is not replaced.
set -C
if iconv -f WINDOWS-1252 -t UTF-8 "$source" >"$destination"; then
    printf 'converted %s to %s\n' "$source" "$destination"
else
    status=$?
    printf 'conversion failed with status %s; original kept unchanged; inspect partial output at %s before removing it\n' \
        "$status" "$destination" >&2
    exit "$status"
fi

Replace WINDOWS-1252 only after establishing the actual input encoding and confirming the local iconv supports that name. set -C is the shell’s noclobber mode for ordinary output redirection, but exact behavior and available options can differ in non-POSIX shells. The new output is kept separate from the source. If several processes may race to publish one output or the result is mission-critical, create a private temporary file using the platform’s documented mktemp form in the destination filesystem, validate it, then rename it with an explicit overwrite policy. Do not assume a cross-filesystem move is atomic.

Never add iconv -c, //IGNORE, or transliteration options merely to make a command finish. Such modes can discard or alter characters. A nonzero exit status should stop the pipeline and leave the original intact. Review diagnostics and compare record counts, line counts, checksums where appropriate, and representative decoded text. If the input contains mixed encodings, one global conversion cannot infer which region uses which encoding; repair or re-export it at the source when possible.

Make automation’s locale boundary explicit

Locale bugs often appear only outside an interactive terminal because jobs inherit a different environment. Cron implementations commonly provide a small environment and may default to /bin/sh; SSH may accept selected client locale variables only when server policy permits it; CI systems and containers can have different installed locales. Do not assume that a shell startup file ran, or that a locale set in an interactive profile reaches a daemon.

For a script whose contract requires UTF-8, validate the required locale at startup or pass the correct category to the exact command that needs it. Keep machine-readable protocols independent of localized output: prefer explicit delimiters and stable formats over parsing human messages, dates, decimal separators, or localized diagnostics. If reproducible ordering is required, document which tool and locale define that ordering; changing LC_COLLATE can change sort order without changing any file bytes.

For example, a cron entry can state an intended environment directly, but the locale name must exist on that host:

LANG=en_US.UTF-8
LC_ALL=en_US.UTF-8
15 2 * * * /usr/local/libexec/import-report

Cron’s parsing rules are implementation-specific. In Cronie-style crontabs, an unescaped percent sign in the command field has special meaning, so avoid embedding complicated command text there; keep logic in a script and inspect the installed crontab(5) manual. Exporting a locale in the crontab also cannot generate that locale or install missing character maps.

On Debian-family systems that use the locales package, locale-gen compiles selected locale definitions from /etc/locale.gen; other distributions have different mechanisms. Check the platform documentation and package state rather than editing locale files on an unfamiliar system. Then confirm the runtime can see the generated name with locale -a and that the actual service process receives the intended variables.

Diagnose the display layer last

If conversion and command output bytes are correct, investigate the renderer: terminal encoding configuration, font coverage, remote terminal settings, and the application that displays the data. A missing glyph can appear as a box even when UTF-8 bytes are valid. A mojibake display can result from interpreting correct bytes with the wrong encoding. Redirect a short sample to a file and inspect its bytes independently of the terminal. Compare the same file in a second known UTF-8-aware viewer; do not rewrite the source until the evidence points to bad bytes.

When connecting over SSH, environment forwarding is a separate mechanism from encoding conversion. The server may accept or reject locale variables by configuration. A forwarded LANG value can refer to a locale that is not installed remotely. Check the effective client configuration and the server’s accepted environment policy, then run locale on the remote side. Do not enable broad environment forwarding as a shortcut, especially across trust boundaries.

A small acceptance matrix

Before rolling a fix into a pipeline, test representative cases and record expected behavior:

  1. ASCII-only input under C and the intended UTF-8 locale.
  2. Valid non-ASCII text in the documented source encoding.
  3. Invalid or truncated byte sequences, which should fail or follow an explicit documented replacement policy.
  4. Sorting with characters whose order differs between C and the target locale.
  5. An unavailable locale name, which should produce a clear setup failure instead of silently assuming support.
  6. The same command from an interactive shell, cron or service context, and the actual CI/SSH environment used in production.

Store a representative fixture with the expected bytes, not just a screenshot. Assert conversion exit status and output checksum or decoded records where appropriate. The objective is not to force every tool into UTF-8 at all costs; it is to make the encoding contract explicit, detect violations, preserve original data, and ensure the display layer is not mistaken for the storage layer.

Related:

Sources:

Comments