Skip to content
Shell & TerminalDeep Dive Published Updated 7 min readViews unavailable

cut in Shell Scripts: Delimiters, Character Sets, and Record Assumptions

Use cut only for the record format it understands, distinguishing bytes, characters, fields, locale effects, and arbitrary filename boundaries.

cut selects byte positions, character positions, or delimiter-separated fields from each input line. It does not parse CSV quoting, JSON structure, or arbitrary records with embedded newlines. Its usefulness comes from a narrow, predictable model; production scripts become fragile when that model is mistaken for a general parser.

Before using cut, write down the producer’s record contract: one record per line, delimiter choice, whether delimiters may appear in values, whether empty fields matter, character encoding, and behavior for malformed rows. If any are unknown, validate or parse with a tool that understands the format.

Bytes, characters, and locale

The -b option selects byte offsets. Without -n, a selected byte range can split a multibyte character and yield invalid text downstream. POSIX defines cut -b -n to avoid splitting characters recognized by the active locale; verify support and locale behavior on the target implementation. The -c option selects characters according to locale and multibyte behavior. They are not interchangeable for UTF-8 text: the second character may span several bytes.

For a protocol defined over ASCII bytes, pinning the C locale and selecting byte ranges can make intent explicit. For localized text, select characters and test the exact locale and implementation. Neither option counts grapheme clusters as users perceive them; a visible accented letter or emoji can contain multiple code points. If display-column width is the real goal, use a terminal-aware library instead.

-n is the boundary-preserving option for byte ranges in a multibyte locale. It does not turn -b into a general Unicode-text operation: it adjusts which bytes are emitted so that the selected ranges do not split a recognized character, and an incomplete or malformed input sequence still needs an explicit policy. If a protocol defines bytes, run under a byte-oriented locale and use byte ranges. If the contract defines characters, use -c or a parser for that encoding. Test the exact implementation because GNU and BSD versions can differ in newer extensions while sharing the POSIX baseline.

Field selection is not schema validation

With -f, cut selects fields based on a delimiter. POSIX defines a single-character delimiter model for the standard utility. Consecutive delimiters can create empty fields, and trailing delimiter behavior should be tested on target implementations. A row with fewer fields may produce less output rather than an explicit schema error.

For tab-separated records with no quoting, a script can select stable columns:

cut -f 1,3 inventory.tsv

The field list selects field numbers, but output remains in the input’s order. Thus -f 3,1 does not reorder fields into “third, then first.” With the default tab delimiter, a run of tabs is not equivalent to a run of spaces, and cut does not trim whitespace. The standard field mode also passes through a line that has no delimiter unless -s is supplied. That behavior is easy to miss when a malformed header or a one-column line enters a stream: the result can look like valid output while having a different shape.

For a strict three-column feed, validate the record count separately before extracting fields. Do not count delimiter bytes if the format permits escapes or quoted delimiters. A line-oriented validator can still be wrong when the actual format permits a newline inside a quoted record; CSV, for example, needs a parser that understands quotes and record framing together. Validation should reject unexpected fields, missing required fields, and invalid encodings before any extracted value is used to choose a path, command, or privileged action.

GNU implementations accept additional extensions and option forms. Do not use an extension without pinning the tool. Avoid using comma as a delimiter for real CSV: quoted commas, escaped quotes, and embedded newlines invalidate the simple field model. Use a CSV library.

Use an upstream validator when exactly N fields are required. Counting delimiters is valid only if the grammar guarantees the delimiter cannot occur inside a field. Validate encoding and line endings separately when files arrive from mixed platforms; a carriage return before newline can remain attached to the last field.

Preserve argument and stream boundaries

cut reads paths passed as command operands or standard input. Quote shell variables so a path with spaces remains one argument. For filenames containing leading dashes, use option termination where supported or prefix a relative path with dot-slash; test older implementations if portability matters. Do not capture NUL-delimited path streams in shell variables, and do not confuse newline-delimited text fields with pathname records.

When cut is in a pipeline, a failure from the producer can be hidden by a successful final command in shells without pipefail. If output completeness matters, stage data and check producer and transformer statuses separately. A consumer may succeed on truncated input if truncation leaves valid-looking records.

Missing fields and output shape

Selecting fields can create output records with a different number or ordering of delimiters than the input. Document that shape for downstream commands. Output can be ambiguous if fields themselves contain tabs or newlines. If arbitrary field values must survive, use structured serialization rather than concatenating raw values into another delimiter-based stream.

cut does not tell the caller whether a selected field was empty, absent, or dropped because the line was short unless the script checks. For required values, validate each record before extraction. For optional values, distinguish an empty field from no field if that distinction matters. Avoid relying on terminal display to verify tabs and trailing spaces; compare bytes or use fixtures that expose framing.

The output is a projection, not a faithful serialization of the original row. When -f emits selected fields, it writes the field separator between the selected fields, even if intervening fields were omitted. This can preserve a simple TSV shape, but it cannot preserve arbitrary field values safely when the delimiter or newline is allowed inside a value. A consumer should receive the same format contract as the producer, including whether an empty selected field is meaningful and whether the number of output fields is fixed.

Shell quoting solves only the argument boundary between the shell and cut. It does not make the content safe to interpolate into a later shell command, SQL statement, or path. Keep extracted data as data: pass values through quoted arguments, use a language API for structured query parameters, and validate path components before filesystem operations. If the input can contain NUL bytes, cut is not a general binary-record parser; line-oriented text tools impose newline and text assumptions.

Test before publishing

Use fixtures with empty lines, leading and trailing delimiters, consecutive delimiters, too few fields, CRLF endings, multibyte characters, delimiter characters inside values, and a final line without newline. Confirm exact output bytes under the intended locale. If a pipeline produces a file, write to staging and publish only after input is fully validated and every command succeeds.

Include both selected and unselected multibyte characters around byte-range boundaries, and compare the result with and without -n. Exercise an invalid row before and after a valid row so the validator cannot accidentally accept partial output. If a pipeline writes a downstream artifact, send output to a temporary path on the same filesystem, check every command’s status, then rename it into place only after the whole input has been accepted. The check must account for shell pipeline semantics: POSIX shells do not all expose a pipefail option, so a portable script may need to stage intermediate output or check commands individually.

cut is a good choice for simple fixed records such as constrained TSV. It is not a format detector or validator. Correct use makes assumptions explicit and rejects input that violates them instead of silently manufacturing plausible but wrong columns.

Related:

Sources:

Comments