Bats Tests: Isolate Shell Behavior and Assert the Right Status
Build trustworthy Bats suites with explicit version requirements, isolated fixtures, correct run semantics, reliable teardown, and useful failure output.
Bats, the Bash Automated Testing System, gives shell scripts named tests and TAP-oriented output, but it does not remove the need to understand Bash process behavior. Its special @test syntax is preprocessed by Bats, test setup has explicit lifetimes, and the run helper captures a command in a separate execution context. A suite that ignores those boundaries can accidentally test the harness instead of the program, leak state between cases, or turn a failed command into a green test.
Treat a Bats test as an executable specification with a controlled environment. Each test should arrange its own fixture, invoke the program with real argv boundaries, assert status and output separately, and leave no persistent system changes behind. Keep privileged, network, or destructive integration tests in a separately identified suite rather than hiding them in an ordinary fast test run.
Pin the Bats feature level
Bats features have evolved. A suite that uses a helper introduced in a newer release should state a minimum supported version so an older CI image cannot silently run a weaker test. For example, the upstream documentation provides bats_require_minimum_version; use it only when the installed Bats version supports that declaration, and set the declared floor to the features actually used by the suite.
bats_require_minimum_version 1.7.0
setup() {
export PATH="$BATS_TEST_DIRNAME/../bin:$PATH"
export APP_CONFIG="$BATS_TEST_TMPDIR/config.json"
}
The .bats file is not plain Bash source when it uses the @test "description" { ... } form. Bats preprocesses that notation before execution, so running bash -n suite.bats is not a substitute for bats suite.bats. The project also documents a function form marked with #@test that is valid Bash syntax and can be friendlier to some shell tooling. Choose a form consistently and ensure formatters, linters, and editors understand the exact syntax in use.
At CI startup, print the Bats version and invoke the suite with explicit paths. Do not rely on a developer’s global test runner defaults. If a feature, plugin, or assertion library is required, declare and pin that dependency just like other test infrastructure; a missing helper should cause a clear setup failure, not silently skip assertions.
Use a unique fixture for each test
The test-scoped $BATS_TEST_TMPDIR is intended for files needed only by the current test. Prefer it to writing into the repository, a developer’s home directory, or a shared /tmp filename. Isolated locations let tests run in parallel later without racing over the same fixture and keep a failed run inspectable until its normal cleanup policy removes the temporary tree.
@test "worker reports a missing configuration" {
run --separate-stderr worker --config "$BATS_TEST_TMPDIR/absent.conf"
[ "$status" -eq 2 ]
[ -z "$output" ]
[[ "$stderr" == *"configuration file not found"* ]]
}
The expected status and diagnostic in this example are part of the hypothetical program’s documented interface; replace them with the real contract under test. run saves the command’s status and output, then returns success to allow assertions to continue. Without an assertion on $status, a command that unexpectedly succeeds can leave the test green. Likewise, checking only a substring can miss a changed error code or an accidental success path.
run executes its arguments in a subshell. A variable assignment, current-directory change, or shell option change made by the command does not propagate back into the test process. Use a direct command when the test intentionally needs to observe a side effect in the current shell, but account for Bats’ own test isolation and errexit behavior. For filesystem state, assert on the resulting file or directory rather than expecting a shell variable to escape run.
Tests should isolate environmental state deliberately. Set a temporary HOME, XDG_CONFIG_HOME, locale, PATH, and application-specific environment when those variables influence behavior. Restore global state only when modifying it is unavoidable, and install cleanup before a mutation occurs. Prefer a temporary fake executable earlier in the test’s PATH to mocking shell builtins or rewriting the user’s real configuration.
Keep assertions attached to the process you ran
A common mistake is to write run command | filter. Bash parses the pipe outside run; the filter receives the output from run, not the command output that run captured. Use bats_pipe when a pipeline itself is under test, and escape the pipe so the helper receives it. The Bats documentation describes how bats_pipe propagates pipeline status; verify the version-specific syntax and choose which pipeline status matters to the assertion.
@test "producer failures are not hidden by a successful formatter" {
run bats_pipe producer-that-fails \| formatter-that-succeeds
[ "$status" -ne 0 ]
[[ "$output" == *"producer diagnostic"* ]]
}
This test should use a deterministic fixture rather than a flaky external service. A strong suite tests both the expected success and the failure mode that an operator needs to distinguish. If a pipeline can fail at either stage, test each stage independently, then test the aggregate policy. Do not assume that a terminal display is identical to the captured output; line arrays may omit empty lines unless configured to retain them, and combined stderr/stdout can obscure channel-specific behavior.
Use run --separate-stderr when the distinction between standard output and standard error is part of the interface. Assert exact output for stable machine-readable formats; use a narrowly chosen substring or regular expression for human messages that include paths or platform-specific details. Avoid asserting an entire trace if the test’s actual purpose is only to confirm an exit status.
Give setup and cleanup the right lifetime
setup and teardown execute for each test. setup_file and teardown_file wrap the tests in one file, while suite hooks are designed for an entire suite through their documented setup file mechanism. Select the narrowest lifetime that provides the needed fixture. A file-level database fixture can save time, but shared mutable state makes failures order-dependent; a test-level directory is safer for ordinary unit tests.
setup() {
export APP_HOME="$BATS_TEST_TMPDIR/home"
mkdir -p "$APP_HOME"
}
teardown() {
# Test-scoped temporary content is isolated; remove any external resource
# created by this test here, using its exact recorded identifier.
if [ -n "${CREATED_RESOURCE_ID:-}" ]; then
cleanup_test_resource "$CREATED_RESOURCE_ID"
fi
}
This sketch relies on Bats providing the test-specific directory. Any cleanup for resources outside that directory should use an exact identifier recorded by the test and must not broaden into a wildcard deletion. If teardown can fail, preserve both the test failure and cleanup diagnostic in a way that remains visible in the TAP result. Cleanup code should be idempotent where possible and should not mask the original failure with a misleading zero status.
Skipped tests still run setup and teardown. Do not put irreversible work in setup on the assumption that a test body will always execute. A test that can be skipped after allocating a remote resource must still clean it up. For a test with multiple setup stages, install each cleanup action as soon as the corresponding resource has been created.
Avoid false positives in the harness
Do not place test execution at file top level when it belongs inside a test or setup hook. Bats may evaluate files in multiple contexts, especially during filtering and discovery. Top-level side effects can run more than once or occur before the suite reports its plan. Keep declarations at the top level, but put mutable setup and assertions inside the appropriate hook or test case.
When a test launches background children, remember that Bats uses file descriptor 3 for TAP output. A long-lived child that inherits that descriptor can keep it open and make Bats wait even after the test function appears done. Close it for child processes that should not write test diagnostics, and wait for processes that the test intentionally started. Better yet, run short-lived helper processes in the foreground and capture their status.
A helper library can reduce duplication, but avoid implicit global names and hidden setup. Load libraries with an explicit path, test helper functions independently where practical, and be careful with Bash declare inside a sourced helper function: its variables can be local to the function that performed the source. Choose local or declare -g deliberately when sharing state, and document the API instead of relying on an accidental side effect.
Run tests like production code
Run the suite from a clean checkout using the project’s pinned Bats release. Execute it with a deterministic locale, known environment, and explicit working directory. Run the fast unit suite on each change, then reserve slower integration tests for a separate command or CI job. Publish machine-readable TAP output when downstream tooling consumes it, while keeping ordinary reporter output available to developers.
Exercise both the test harness and the program under test: a passing command, a failing command, empty stdout, stderr-only output, a path with spaces, an environment variable that is unset, and a cleanup failure. Confirm that a failing assertion makes Bats return nonzero and that filtered runs cannot accidentally become the required CI result. If using focus tags, ensure the normal CI invocation does not enable focus-only mode.
Finally, test the tests. Temporarily change an expected status or output in a disposable branch and confirm that the assertion fails with a useful location. A suite that passes only because its commands are skipped, mocked too broadly, or run from a developer’s working tree is not a reliable guard. Bats is small and effective when each test makes its process boundaries, fixture ownership, exit-status contract, and cleanup policy explicit.
Related:
- How to Keep a Shell Script ShellCheck-Clean from the Start
- How to Build a Reusable Shell Function Library
Sources: