Skip to content
Haiku OSDeep Dive Published Updated 7 min readViews unavailable

Haiku BFlattenable: Type Codes, Buffer Bounds, and Stable Wire Formats

Design safe Haiku BFlattenable types with explicit type codes, bounded parsing, versioned payloads, and compatibility distinct from BArchivable.

BFlattenable is an interface for turning a value object into a byte sequence and reconstructing it later. It lets Haiku APIs such as BMessage carry custom typed data without every caller inventing its own copy routine. The interface is deliberately small: the object reports a type code, whether its flattened size is fixed, the required size, and methods to flatten and unflatten bytes.

Flattening is a data-format contract, not automatic object persistence. The interface does not define your field meanings, schema version, byte order, ownership graph, or trust policy. A successful call only means the object accepted a buffer under its implementation. Applications still need bounds checks, compatibility rules, validation, and tests for malformed input.

Define a unique type and exact size semantics

Implement all required virtual methods. TypeCode() identifies the flattened representation. Use a type code reserved for your own format unless your bytes are intentionally compatible with an existing Haiku type. Reusing a familiar four-character code because it looks descriptive is unsafe: a reader may then interpret your payload according to a different contract.

IsFixedSize() must match actual serialized behavior. FlattenedSize() returns the number of bytes needed for the current value. For a fixed-size record, that value is constant by format. For variable-size data, calculate from validated field lengths and guard against integer overflow before returning ssize_t. Callers may allocate based on this value, so an incorrect size can produce truncation or memory errors.

status_t StoreValue(BMessage& message, const BFlattenable& value)
{
    ssize_t size = value.FlattenedSize();
    if (size < 0)
        return B_BAD_VALUE;

    return message.AddFlat("payload", &value);
}

BMessage::AddFlat() asks the object for its flattenable representation. Do not assume this makes the format portable across unrelated architectures or future versions. The type code tells a consumer which representation to request; it is not a cryptographic identity or schema registry.

Make Flatten() capacity-safe

The caller provides a buffer and a size. Flatten() must reject null storage when bytes are required and must not write more than the supplied capacity. Check the provided size against the exact number needed before writing. On failure, return an error without leaving the destination in a state that could be mistaken for a complete value.

Avoid serializing a raw C++ struct with memcpy. Padding bytes, alignment, pointer fields, compiler ABI, and host byte order are not stable wire-format details. Instead, define field order and widths explicitly. For integers, specify byte order; for strings or arrays, use bounded lengths and define whether a terminator is included. If the data can cross a version boundary, encode a version field and decide how older and newer readers behave.

For variable-sized content, compute a checked total before allocating or writing: fixed header, each length field, then each payload. Reject negative lengths and sums that exceed the representable size or application maximum. A malicious message should not be able to force an unbounded allocation simply by advertising a huge field length.

Parse transactionally in Unflatten()

Unflatten(typeCode, buffer, size) must verify that the code is accepted, the pointer and length are valid, and every field fits within the remaining input. The default AllowsTypeCode() implementation compares the supplied code with TypeCode(). Override it only when you genuinely accept a documented compatible representation, and make the conversion explicit.

Parse into temporary local state first. Validate enum ranges, nested lengths, checksums if your own format uses them, and required fields. Only after the entire payload is valid should you replace the object’s current state. This avoids leaving a half-updated object after a parse fails near the end. Return the appropriate error rather than silently accepting trailing garbage unless the schema explicitly reserves an extension area.

status_t Unflatten(type_code code, const void* data, ssize_t size) override
{
    if (data == NULL || size < kHeaderSize)
        return B_BAD_VALUE;
    if (!AllowsTypeCode(code))
        return B_BAD_TYPE;

    // Decode into temporary fields, validate all lengths and values,
    // then commit them to the object only after the full input is valid.
    return DecodeAndCommit(data, size);
}

DecodeAndCommit() above names an application helper, not a Haiku API. In the actual implementation, perform checked cursor arithmetic before each read and never advance beyond the supplied size. If data is untrusted, fuzz the parser with truncated headers, oversized lengths, invalid types, duplicate fields, and random byte sequences.

Choose between flattenable data and archivable objects

Haiku documentation distinguishes BFlattenable as a compact byte representation from BArchivable, which reconstructs richer live objects through BMessage archives. Use flattenable data for a value with a clear byte-level representation, such as a structured identifier, geometry value, or custom message payload. Use archiving when restoration needs class construction, named fields, or relationships among object state.

Neither interface automatically makes data safe to execute or trustworthy. A deserialized value can still request an out-of-range operation, refer to a missing file, or carry invalid configuration. Keep parsing separate from applying the result. Validate policy at the point where the value is used, especially when messages can arrive from another application or persistent storage.

Compatibility strategy

Treat the flattened bytes as a versioned API once they are written to disk, sent across process boundaries, or consumed by independently updated code. Preserve old readers with explicit migration branches or reject unsupported versions with a recoverable error. Avoid changing the meaning of a field while keeping the same type code and version; such changes often parse successfully but create subtle semantic corruption.

Document fixed versus variable size, maximum payload, numeric widths, byte order, string encoding, and extension behavior. If a format includes optional trailing fields, define how old readers handle them and how new readers distinguish a truncated payload from an older valid payload. Maintain golden byte fixtures so changes can be compared across compiler and system upgrades.

For example, a versioned point payload might be a fixed header containing a version and two explicitly encoded integer coordinates. That contract should specify exact byte order and width instead of relying on sizeof(Point), even if the current struct happens to contain only two integers. If a later version adds a label, keep the old version readable and put the string length in a bounded field. This makes compatibility review possible from a byte fixture without needing the original compiler or class layout.

Keep a written table of byte offsets beside the implementation and compare it to golden fixtures during review. The table should identify reserved bytes and whether they must be zero, ignored, or preserved. If an older reader might encounter a newer payload, choose deliberately between rejecting the version and parsing a compatible prefix; accepting unknown bytes accidentally is not forward compatibility. Fuzz and fixture tests should cover both the parser and the size calculation so they cannot drift apart.

Operational checks

Test an empty or minimum value, maximum allowed payload, exact buffer size, one-byte-short buffer, null pointers, unsupported type codes, unknown version, truncated fields, integer-overflow lengths, and valid historical fixtures. Verify both directions: flatten, persist or send, then unflatten into a fresh instance. Confirm a failed unflatten leaves the previous object unchanged.

Measure allocation behavior if this format is used in a media callback or other latency-sensitive path. Preallocate bounded buffers when necessary, and avoid expensive parsing under a global or window lock. For diagnostics, report the type code, version, parse offset, and status without dumping sensitive payload bytes.

Acceptance criteria

Accept a BFlattenable type when type codes are unique and documented, size methods agree with actual bytes, output honors the caller capacity, parsing is transactional and bounded, and wire-format compatibility is tested independently of the current process. Verify malformed input and persisted older fixtures before release.

BFlattenable supplies a reusable serialization hook. The class author remains responsible for a stable format, memory safety, migration, and application-level trust decisions.

Related:

Sources:

Comments