od for Shell Diagnostics: Inspect Bytes Without Guessing at Text
Use od to inspect byte streams, offsets, NULs, and encodings while keeping display formats separate from binary parsing and validation.
When a text file looks wrong, the visible characters are only one interpretation of the bytes. A terminal can conceal carriage returns, NUL bytes, a byte-order mark, invalid UTF-8, or a missing final newline. od turns a stream into a representation that exposes offsets and byte values. It is a diagnostic utility, not a decoder, format validator, or safe way to reconstruct arbitrary binary data through shell variables.
Use od to answer a narrow question: what bytes exist at a particular offset, where a delimiter occurs, or whether two files differ in framing? For protocol work, file-format debugging, and CI fixtures, byte-level evidence is often more reliable than cat, editor display, or a screenshot. The result is still a textual rendering chosen by an output type; do not confuse that rendering with the original representation.
Choose byte output deliberately
POSIX od supports address bases, skipping, byte counts, and type selections. A common GNU/BSD invocation prints byte values in hexadecimal and omits the offset column:
LC_ALL=C od -An -tx1 -v -- "$file"
-An suppresses addresses, -tx1 requests hexadecimal one-byte units, and -v prevents repeated lines from being collapsed into an asterisk. GNU and BSD implementations support these options, but always test the exact target utility and its treatment of --. If offsets matter, omit -An or specify an address base intentionally. Set the locale when using character output because character classification can be locale-sensitive.
Useful diagnostics include inspecting the beginning of a file with a byte count, checking a suspected delimiter, and viewing line endings. For example, od -An -tx1 -N 32 file displays up to the first 32 bytes on common systems. Compare the output with the expected protocol framing, not with an assumption that every file begins with a familiar signature. A matching prefix is evidence, not proof that the complete object is valid.
Text output is not Unicode decoding
The c output type can show printable characters and escapes, but “printable” depends on locale and implementation. Hex output avoids that classification when investigating raw byte values. UTF-8 characters can span multiple bytes; the bytes c3 a9 may represent one character under UTF-8, while the same bytes have other interpretations under other encodings. A byte dump does not normalize Unicode or tell you which encoding an upstream system intended.
For a suspected UTF-8 file, inspect bytes and then run an actual decoder or validator that rejects malformed sequences. Do not use iconv in a mode that silently substitutes invalid bytes if the goal is to prove validity. Check byte-order marks and line endings separately. When comparing text captured on Windows and Unix, distinguish CRLF (0d 0a) from LF (0a), and do not delete every carriage return unless the data contract says all of them are line-ending artifacts.
Binary layouts depend on type and endianness
Multi-byte numeric formats introduce byte order, width, alignment, and representation questions. An od option that interprets groups as unsigned short or integer uses the implementation’s documented sizes and native byte order unless an option says otherwise. A dump of native integers is not automatically a portable decoding of a network protocol. Network formats often define fixed-width big-endian values, while a host may be little-endian and use different C type widths.
Start with byte output and offsets, then decode with a parser that specifies width and endianness. Do not assume od -tx4 means the same semantic integer on every architecture. If using it for a quick inspection, state the machine and tool version. For reproducible tests, define input bytes exactly and assert on parser results rather than scraping locale-formatted od output.
Arbitrary filenames and stream boundaries
Passing a pathname as one quoted argument preserves spaces, tabs, and wildcard characters. A filename beginning with a hyphen can be interpreted as an option by utilities that lack or do not honor an end-of-options marker; prefix a relative name with ./ when appropriate. A pathname itself cannot contain NUL, but file contents can, and command substitution cannot preserve NUL bytes. Never capture an arbitrary binary dump into a shell variable and then expect byte-for-byte fidelity.
od reads bytes; it does not establish that a file is regular or safe to read. A FIFO can block waiting for a writer, a device can have side effects, and a huge file can produce a huge diagnostic. Bound inspection with a byte count, validate input type when appropriate, and consider time or resource limits for untrusted paths. Keep diagnostic output out of machine protocols unless the format is deliberately specified.
Use it in a repeatable investigation
Preserve the original artifact and record its checksum before running transformations. Capture the exact command, locale, utility version, offsets, and byte range. Compare a known-good and problematic file with cmp first; then use od to explain the difference. If the difference is beyond the sampled range, increase the range or use a purpose-built parser rather than drawing a conclusion from the prefix.
Test empty input, a single byte, repeated lines, NUL bytes, CRLF, malformed multi-byte text, and files large enough to trigger truncation or display elision. A byte dump is excellent forensic evidence when its range and representation are explicit. It becomes misleading when readers infer encoding, validity, architecture-independent numeric meaning, or full-file equality from a short visual sample.
Use bounded comparisons and preserve provenance
For incident handling, copy an artifact to a controlled evidence location before repeated inspection, record its size and digest, and avoid opening it with tools that may rewrite metadata. od is read-only for ordinary file operands, but reading can still update atime depending on filesystem policy. A forensic workflow should note that possibility and use a read-only mount or image when original evidence must remain unchanged.
When two byte streams differ, cmp can identify the first differing offset without producing a massive dump. Then request only a small window around that offset with od on each file. This avoids burying relevant evidence in a full dump and reduces the chance that terminal output is truncated. If a difference is caused by timestamps or random identifiers embedded in a binary format, use the format’s parser to map the offset to a field before declaring the files semantically different.
Do not use text tools such as sed, tr, or command substitution to transform arbitrary binary output. Shell variables cannot contain NUL bytes, and many utilities interpret encodings or newline-delimited records rather than opaque bytes. If byte-level extraction or patching is required, use a language runtime with byte arrays, explicit offsets, bounds checks, and a format-aware validator. Keep the original intact and verify transformed output against independent expected properties.
Reproducible diagnostics in CI
When an od dump is part of a test failure, cap output and attach the input digest, expected range, locale, and tool version. A regression test should compare exact bytes or parser-level values, not golden text whose spacing and address formatting could change. If using command output as an artifact, strip only presentation fields whose semantics are documented; do not normalize away bytes that the test is supposed to catch.
Test both standard input and file operands if the script can receive either. A pipeline may be interrupted, and the producing command’s failure must not disappear behind a successful od. In Bash, capture PIPESTATUS immediately if each stage’s result matters; any subsequent command can overwrite it. In portable shell, stage the input or use explicit temporary files and status checks. The byte view is only as trustworthy as the stream that reached the viewer.
When inspecting a partial network capture or truncated download, record the expected length and compare it with the actual object before interpreting offsets. A byte dump of a prefix can show a valid header even when the payload is incomplete. Conversely, a truncated dump may stop in the middle of a multi-byte field and make a correct stream look malformed. Capture the full relevant field boundary and note whether offsets are measured from the original object, a skipped region, or the beginning of a concatenated stream.
For incident reports, prefer a short reproducible command and a bounded excerpt over pasted terminal output with no context. Include the digest of the input, the exact range inspected, the rendering options, and the implementation. Review whether the excerpt itself contains credentials, personal data, or secret material before publishing it. od makes hidden bytes visible, which is valuable diagnostically but can also expose data that the normal text view obscures.
Related:
- cmp Exit Status in Shell: Equal Files, Different Files, and Errors
- Fixing Locale and Encoding Failures in Shell Tools
Sources: