Skip to content
macOSDeep Dive Published Updated 8 min readViews unavailable

Core ML on macOS: Model Loading, Prediction, and Compute Boundaries

Ship Core ML models with validated feature contracts, deliberate compute settings, serialized prediction ownership, measurable latency, and safe failure handling.

Core ML runs a model; it does not define what a prediction means to your product. A model file has a feature contract, preprocessing assumptions, supported operations, and expected output types. A production integration validates those contracts at the boundary, gives the model a clear owner, and treats inference as a fallible operation whose latency and output quality must be measured.

For a model added to an Xcode target, Xcode can generate a Swift interface from the model description. That wrapper is usually the safest starting point because it exposes named inputs and outputs. The lower-level MLModel interface is useful when models are downloaded and compiled at runtime, when a common host supports models with varying schemas, or when the app needs to inspect metadata and feature descriptions dynamically.

Choose the loading path deliberately

An .mlmodel source asset is not necessarily the file your app opens at runtime. Xcode compiles model assets into a device-optimized representation for inclusion in the app bundle. A runtime-downloaded model follows a different flow: obtain the model from a trusted product distribution path, compile it with Core ML, keep the returned compiled URL, and load that compiled model. A successful download alone does not make a file loadable as an MLModel.

Use an explicit configuration so the deployment policy is visible in code. computeUnits describes the set of processing-unit configurations the model is allowed to use. .all allows Core ML to choose from supported processing units; it does not promise that every prediction runs on a Neural Engine or GPU. Availability depends on the model, OS, and hardware. .cpuOnly can be appropriate for a workload that must avoid GPU contention, but it can also change latency and power use. Benchmark the actual model on representative supported Macs rather than inferring performance from the configuration value.

import CoreML
import Foundation

func loadCompiledModel(at url: URL) async throws -> MLModel {
    let configuration = MLModelConfiguration()
    configuration.computeUnits = .all
    return try await MLModel.load(
        contentsOf: url,
        configuration: configuration
    )
}

func predictPrice(model: MLModel, size: Double) throws -> Double {
    let input = try MLDictionaryFeatureProvider(dictionary: [
        "size": MLFeatureValue(double: size)
    ])
    let output = try model.prediction(from: input)
    guard let value = output.featureValue(for: "price"),
          value.type == .double else {
        throw PredictionError.missingOrInvalidOutput
    }
    return value.doubleValue
}

enum PredictionError: Error {
    case missingOrInvalidOutput
}

The feature names and numeric types above are illustrative. They must match the model’s actual interface. A model expecting an image, multi-array tensor, or differently named scalar will not accept the example input. Do not silently substitute zero or an empty tensor when a required input is absent; return an explicit validation error that can be diagnosed and tested.

Treat the model description as a runtime contract

Before constructing dynamic inputs, inspect model.modelDescription.inputDescriptionsByName and outputDescriptionsByName. A feature description reports its name, type, optionality, and type-specific constraints. Validate that the model version your application loaded has the contract the caller expects. This is particularly important when a server can update a model independently of the application binary.

Feature names are identifiers, not localized labels. Keep a versioned schema alongside the model release: expected names, shapes, value types, units, normalization, allowed ranges, and output interpretation. Core ML can validate structural requirements, but it cannot know that a floating-point value is in meters rather than centimeters or that an image was normalized using the correct channel order. Those are application-level contracts.

Where possible, add a narrow adapter that converts domain inputs into model features and maps outputs into domain results. This keeps a model-specific wrapper out of views and persistence code. If replacing a model changes the feature schema, the adapter becomes the deliberate compatibility boundary instead of a scatter of string keys and casts throughout the application.

Generated model wrappers are convenient when a model is bundled with the app and its schema is known at build time. Direct MLFeatureProvider implementations can avoid unnecessary copies when the source data is produced asynchronously or is already held in a suitable representation. Use the dynamic interface only when it solves a real integration need; it transfers more responsibility for names, value types, and validation to the application.

Own model use and concurrency explicitly

Apple documents that an MLModel instance should be used on one thread or dispatch queue at a time. Serialize calls through one owner, or create separate model instances for independently concurrent workers. Do not assume that a Swift Task automatically makes a shared model safe to call concurrently. A model service can own loading, the model instance, input validation, timing, and cancellation policy, then expose a smaller API to the rest of the app.

Loading can be expensive and can fail because the compiled model is incompatible, unavailable, corrupt, or uses operations the runtime cannot execute. Load away from latency-sensitive UI work. Cache a successfully loaded instance according to a deliberate lifetime policy, but keep the source model identifier and configuration with it. If loading fails, preserve a usable non-ML path when the feature permits one; do not block the app’s initial launch on an optional prediction feature.

For a batch, use the batch APIs only when the model and the workload support them. Batching can reduce repeated overhead, but it may increase peak memory and time-to-first-result. A UI that needs one result quickly may prefer a single request while a background analysis pipeline can process bounded batches. Limit concurrent requests and input size; an image decoder or tensor allocation can consume substantial memory before inference begins.

Validate output semantics, not just output presence

A model returns values according to its trained objective. A class score is not automatically a calibrated probability. A regression value is not automatically a safe recommendation. A text label may be drawn from a model vocabulary that changes across revisions. Document the transformation from output tensor or feature values to user-facing behavior, including thresholds and unsupported or low-confidence results.

Separate three outcomes in code: input rejected before inference, inference failed, and inference completed with a valid but uncertain output. This distinction improves telemetry and prevents an empty result from being mislabeled as a framework failure. Use deterministic fixtures for representative inputs, invalid shapes, boundary values, empty data, and model-version changes. For image models, pin the orientation, resizing, crop, color-space, and normalization pipeline in tests.

Record model version, app version, OS version, configuration, input dimensions, duration, and error category when diagnosing the feature. Avoid logging raw images, text, or user data merely to understand inference failures. For privacy-sensitive data, keep processing on-device where appropriate, while remembering that application code can still upload inputs or outputs independently of Core ML.

Measure on supported hardware

Create a benchmark matrix for the oldest supported Mac, current Apple silicon systems, and Intel Macs if they remain in scope. Measure cold model load, warm prediction latency, throughput, peak memory, energy impact, and UI responsiveness. Separate preprocessing and image decode time from prediction time. Compare compute-unit configurations using the same data and model build, and report distributions rather than a single best run.

Include a scalar or non-ML baseline when one is practical. Core ML is not inherently the right solution for a small deterministic rule. Monitor accuracy on a representative labeled evaluation set whenever the model, preprocessing, or OS changes. A performance optimization is not acceptable if it silently degrades quality beyond the product’s documented tolerance.

Core ML provides optimized on-device execution and a typed model boundary. Reliable products still own feature contracts, scheduling, output meaning, error recovery, privacy, and evidence that the chosen model works on real supported machines.

Deploy model updates without losing a working version

If the app installs model assets independently of its binary, represent each model release with an application-owned identifier, expected schema version, and activation state. Download and compile a candidate without replacing the active model. Validate that its inputs and outputs match the adapter contract, run a small deterministic smoke corpus, and only then switch new requests to it. Keep the previous known-good model available until the new one passes the product’s health checks.

Concurrent requests should capture the model generation they started with. A model update must not free an instance still in use or cause a delayed result to be labeled with the new version. Retire the old model after outstanding operations complete. If the app restarts while activation is in progress, recover from a clear persisted state such as candidate, validated, active, or rollback-required.

This is application deployment policy, not a Core ML guarantee. Core ML can load a compiled model and report structural errors; it cannot decide whether the model is accurate enough for your users or whether a remote distribution response is an approved release. Keep the rollout, integrity checks, and rollback criteria explicit and test them with simulated failures.

Related:

Sources:

Comments