Compression on macOS: Streaming Buffers, Finalization, and Safe Decoding
Process compressed data with Apple's Compression framework using bounded buffers, explicit finalization, algorithm metadata, and decompression limits.
Apple’s Compression framework offers buffer-oriented and stream-oriented APIs for encoding and decoding data. A streaming interface is useful when a file or network payload is larger than the memory budget available to hold both input and output at once. The stream keeps codec state between calls while the application supplies the next input and output blocks.
Compression is not a complete archive or file-format design. Your application still needs to identify the algorithm, frame the data, record lengths and versions where appropriate, and decide how to detect corruption. The Compression API transforms bytes according to a chosen algorithm; it does not define your business schema or guarantee that a receiver can infer the intended format from an arbitrary blob.
Choose a buffer or a stream
The single-step buffer functions suit bounded inputs whose decoded or encoded size fits a known destination. They return a byte count and can fail if the destination capacity is insufficient. Do not guess that compressed output will always be smaller than input: incompressible or very small data can expand because of codec overhead.
The stream interface is a state machine. Initialize a compression_stream with an encode or decode operation and algorithm, set the source and destination pointers and available byte counts, process, then advance to the next blocks using the updated pointer and size fields. Once no further source bytes will be supplied, signal finalization. Continue processing until the status is COMPRESSION_STATUS_END, then destroy the stream exactly once.
#include <compression.h>
#include <stdbool.h>
#include <stddef.h>
#include <stdint.h>
bool encodeZlibBounded(const uint8_t *source, size_t sourceLength,
uint8_t *destination, size_t destinationCapacity,
size_t *destinationLength) {
if (!source || !destination || !destinationLength) return false;
compression_stream stream = {0};
compression_status status = compression_stream_init(
&stream, COMPRESSION_STREAM_ENCODE, COMPRESSION_ZLIB);
if (status != COMPRESSION_STATUS_OK) return false;
stream.src_ptr = source;
stream.src_size = sourceLength;
stream.dst_ptr = destination;
stream.dst_size = destinationCapacity;
do {
size_t oldSourceSize = stream.src_size;
size_t oldDestinationSize = stream.dst_size;
status = compression_stream_process(&stream, COMPRESSION_STREAM_FINALIZE);
if (status == COMPRESSION_STATUS_ERROR ||
(status == COMPRESSION_STATUS_OK &&
(stream.dst_size == 0 ||
(stream.src_size == oldSourceSize &&
stream.dst_size == oldDestinationSize)))) {
compression_stream_destroy(&stream);
return false;
}
} while (status == COMPRESSION_STATUS_OK);
*destinationLength = destinationCapacity - stream.dst_size;
bool complete = status == COMPRESSION_STATUS_END;
compression_stream_destroy(&stream);
return complete;
}
This deliberately bounded example has one source block and one destination buffer. If output capacity is exhausted before the stream reaches END, it fails and the caller must discard the partial bytes or retry with a larger staging destination. A production file pipeline should instead refill source blocks and drain destination blocks without losing the stream’s updated pointers and remaining sizes.
Drive an incremental pipeline
For a large file, read a fixed-size source block, point the stream at it, and repeatedly provide a fixed-size output block. Write produced bytes to a staging destination before reusing that output memory. Preserve any source bytes the stream has not consumed before replacing the source block. The exact sequence is easy to get wrong, so test the loop against inputs that end at each boundary around the chosen buffer size.
Do not issue another read merely because the previous call consumed some input. First inspect the stream’s remaining src_size; unconsumed bytes still belong to the current source block. Likewise, do not assume one process call produces a complete frame. Maintain a written count for every output block and handle short writes to the underlying file or socket separately from compression progress.
Finalization is a protocol transition. COMPRESSION_STREAM_FINALIZE means no additional input will arrive for this stream. Do not use it on every block of a multi-block input; doing so tells the encoder that the current block is the end of the whole stream. After the final source chunk is supplied, continue calling process until the stream reports END or an error. Always destroy an initialized stream, including on cancellation and I/O failure.
When decompressing, treat input as untrusted. A small compressed payload can expand into a much larger output. Set a maximum decoded byte count, maximum processing duration, and any nesting or record limits imposed by your format. Abort when a limit is exceeded and do not publish partial decoded output as a valid document. Integrity checks should be part of the container or application protocol, not inferred from successful decompression alone.
Algorithm and format metadata
Choose an algorithm supported by both endpoints and record it in a versioned container when the receiver cannot otherwise know which decoder to use. Do not assume all codecs emit interchangeable framing or that a compressed blob is a ZIP archive. The algorithm constant is a codec choice, not a filename extension policy.
For long-lived files, specify framing, checksum policy, uncompressed size limits, and compatibility behavior in the file-format specification. If the receiver needs random access or independently recoverable blocks, one continuous stream may not be the right format. Chunking can limit memory and bound damage from a corrupt section, but each chunk then needs its own metadata and integrity rules.
If a network transfer is interrupted, decide whether to restart from the beginning or use independently framed chunks with explicit offsets. Compression state is not automatically resumable after process exit. Persisting a few compressed bytes without the corresponding stream state is not enough to continue an arbitrary in-progress encoding session safely.
Ownership and failure handling
Pointers stored in the compression_stream must remain valid for the duration of each processing call. Keep input and output buffers alive until the call returns, and do not concurrently mutate them. The stream structure owns codec state only after successful initialization; destroy it only in that state and do not copy the initialized structure as if it were a value object.
Separate codec state from file state. The codec may consume input while a destination write later fails. Write to a temporary file or object, then validate and publish only after finalization, flush, and the container checks succeed. On cancellation, close files, destroy the stream, and remove or quarantine partial output. Preserve a prior valid destination until the replacement passes validation.
If the source changes during processing, either snapshot it or verify a revision before publishing. A background compression job should not silently write a result for an obsolete document version. Include algorithm, source revision, and byte counts in diagnostics, but avoid logging raw payload data.
Performance without hidden memory growth
Reuse fixed-size buffers where possible. Measure compression ratio, throughput, CPU, and peak resident memory using representative inputs. A faster codec can produce larger files; a higher compression level may spend substantially more CPU for modest savings. Choose settings based on a measured workload rather than assuming one algorithm is universally best.
Bound the queue of compression jobs as well as the buffers inside one job. Launching a task per file can still use unbounded memory if every task retains its buffers and output. Apply backpressure to the producer, cap concurrent streams, and ensure cancellation reaches both the stream owner and pending file/network operations.
File and network boundaries
For a file-to-file transform, separate input reading, codec processing, and output writing into independently measurable stages. A slow disk can be the bottleneck even when the codec is fast. If the writer applies backpressure, stop feeding new input until output drains rather than buffering the entire compressed result in memory. Treat short writes and file close failures as errors even when the codec reached END.
For a network stream, frame the compressed payload so the receiver knows where one message ends and what algorithm and schema it must use. TCP provides an ordered byte stream, not application message boundaries. Do not allow a decoder to consume bytes belonging to the next record simply because the current compressed stream has not finalized. A length prefix must be validated before allocation and should have a strict maximum.
If a format includes an uncompressed-length field, treat it as an untrusted claim. Check it against the configured maximum before reserving output, then verify the actual decoded byte count matches the format contract. A mismatch is corruption. A compression ratio cap alone is insufficient for security because an attacker can choose payload shapes that consume excessive CPU or memory within a superficially plausible ratio.
Keep encryption and compression order explicit. If a container encrypts compressed records, authenticate the metadata that identifies the algorithm and length along with the payload. Do not assume Compression supplies encryption or integrity protection. If confidentiality matters, use a cryptographic format with a documented authentication contract and validate before exposing decoded content.
Acceptance matrix
Test empty input, one byte, data smaller than a buffer, exact buffer boundaries, multiple blocks, incompressible input, destination exhaustion, truncated compressed input, corrupted bytes, expansion-limit rejection, source read failure, partial destination write, cancellation, and restart after a failed attempt. Verify output length, decode round-trip, algorithm ID, checksum or container metadata, and that the old destination remains intact after a failed replacement.
Compression stream statuses should be observable in diagnostics. Distinguish codec initialization failure, process error, insufficient capacity, source exhaustion, destination write error, size-limit rejection, and successful end-of-stream. The goal is not merely to produce bytes but to know whether the complete, validated output was safely committed.
The Compression framework supplies efficient block and stream transforms. Your application owns framing, algorithm negotiation, backpressure, decoded-size limits, durability, and validation. Keep those contracts explicit and a large-data path can remain bounded, interruptible, and recoverable.
Related:
- FileHandle Streaming I/O on macOS: Bounded Reads, EOF, and Ownership
- Image I/O on macOS: Incremental Decoding, Thumbnails, and Metadata
Sources: