Skip to content
macOSDeep Dive Published Updated 8 min readViews unavailable

MetricKit on macOS: Field Reports, Diagnostics, and API Migration

Use MetricKit to analyze real-world macOS performance with delayed reports, diagnostic triage, privacy-aware uploads, and a versioned API migration plan.

MetricKit provides performance metrics and diagnostic reports collected from real use of an app. It complements local profiling; it is not a live monitoring stream, an event-by-event crash callback, or a replacement for reproducing a problem in Instruments. Its value is a field view across conditions the development team may not be able to recreate.

The API surface is changing. Apple’s current documentation introduces MetricManager, MetricReport, and DiagnosticReport as the asynchronous-sequence replacement for the older MXMetricManager subscriber API on macOS 27 and later. The new symbols are marked beta in the documentation. The older MX APIs are deprecated, while the documented legacy API supplies metric reports on macOS 26 and later and diagnostic reports on macOS 12 and later. Select an implementation from the app’s deployment targets and SDK status; do not ship a beta dependency merely because it appears in the newest API page.

Understand the reporting cadence

Metric reports summarize approximately the previous 24 hours and are delivered at most once per day per metric source. A device may deliver an undelivered report together with a later one, and separate metric sources can result in multiple payloads. This cadence is unsuitable for a dashboard that promises minute-by-minute user telemetry.

Diagnostic reports describe events such as crashes, hangs, CPU exceptions, or disk-write problems. The legacy API documents immediate delivery when diagnostics are available on supported OS versions, but the app still needs to be running and subscribed to receive them. Do not build an operational alert that assumes each failure wakes the app or produces a callback at the moment it occurs.

Use the reports to answer aggregate questions: Did launch time regress for a release? Are hangs concentrated in a particular workflow? Did a code path increase disk writes or memory pressure? Pair field metrics with app-defined release and feature context only when the report schema supports that analysis and the added labels do not identify individual users.

Subscribe with a long-lived owner

For deployments using the legacy subscriber API, keep the subscriber object alive for the intended collection period and register it with MXMetricManager. Make report handlers short. Persist or enqueue a bounded representation for later processing rather than running expensive analysis or network uploads inside the callback. Remove the subscriber only when the feature genuinely stops collecting reports.

import MetricKit

final class LegacyMetricSubscriber: NSObject, MXMetricManagerSubscriber {
    func start() {
        MXMetricManager.shared.add(self)
    }

    func stop() {
        MXMetricManager.shared.remove(self)
    }

    func didReceive(_ payloads: [MXMetricPayload]) {
        for payload in payloads {
            enqueueMetricPayload(payload)
        }
    }

    func didReceive(_ payloads: [MXDiagnosticPayload]) {
        for payload in payloads {
            enqueueDiagnosticPayload(payload)
        }
    }
}

private func enqueueMetricPayload(_ payload: MXMetricPayload) {}
private func enqueueDiagnosticPayload(_ payload: MXDiagnosticPayload) {}

This example uses the deprecated API intentionally for deployments that still need its documented runtime coverage. The callbacks are delivery boundaries, not a promise that every metric field is nonnil. Process each payload by checking which metric objects it contains and by preserving its reporting interval and app-version metadata.

For the newer MetricManager API, retain one manager at the owner lifetime you intend and iterate its metricReports and diagnosticReports asynchronous sequences from long-lived tasks. Apple’s documentation warns that creating multiple managers for the same domains and consuming the same sequence concurrently does not produce a full copy for every consumer. Centralize the manager and fan reports out to application-owned processing after receipt. Treat beta availability as a release caveat until the API is stable for the target OS.

Make reports useful without overclaiming

Metric payloads contain measurements such as CPU time, memory, launch behavior, network transfers, disk I/O, and custom signpost intervals. A daily metric is aggregated evidence, not a trace of an individual user action. Keep its time interval, app version, OS version, and device metadata when interpreting a trend. If a report spans multiple application versions, avoid attributing the entire window to the latest build.

Diagnostics can include call-stack information and event-specific details. Symbolicate against the matching dSYM and build artifacts before grouping by function or release. A missing symbol is not proof that the operating system or user caused the failure; it can mean the archive or UUID does not match. Preserve the exact binary identity used for each shipped release.

Custom signpost metrics help relate app-defined work to system measurements. Use stable, low-cardinality names for operations such as document open, search, or model load. Do not put a document title, account ID, URL, or other user data in a signpost name. Define a begin/end interval around the work whose duration matters and verify that all code paths end the interval. MetricKit can aggregate those measurements; it does not infer the business meaning of an event for you.

State-contextualized metrics, where available, allow analysis by app-defined states or domains. Keep state labels small and stable, such as a feature mode or broad workflow phase. Avoid one state per document, customer, or session. A state dimension that changes meaning between releases makes longitudinal comparisons unreliable.

Persist and upload with backpressure

Metric delivery is delayed and can include older reports. Make ingestion idempotent using the payload’s reporting period and stable app/build dimensions rather than assuming one callback is one unique report. If upload fails, queue a bounded number of encoded reports and retry with backoff. When the queue reaches its retention limit, use a documented oldest-first or priority policy; do not allow diagnostic uploads to consume unbounded disk.

Reports may contain sensitive operational details. Minimize fields, define retention, protect transport, and give users the disclosures required by the product and distribution rules. Avoid treating a system-provided device identifier or stack trace as a user identity. A local report can be analyzed without uploading it, and a server pipeline should receive only data needed to answer a defined reliability question.

Track ingestion health separately from app performance. Record report received time, payload interval, parser version, upload result, and deduplication outcome. A drop in reports can mean reduced app usage, OS behavior, an implementation regression, or a reporting delay; it does not necessarily mean users experienced no failures.

Test the collection pipeline

MetricKit supports simulated payloads in development workflows so the parser and persistence path can be tested without waiting for the normal reporting interval. Simulated data validates code paths and schema handling, not actual field performance or delivery cadence. The framework’s documentation also calls for testing report callbacks on physical devices for the legacy workflow.

Test missing metrics, multiple payloads per period, reports spanning app versions, unknown future report values, malformed archived data, symbolication gaps, duplicate delivery, upload failure, storage pressure, app upgrade, and API migration. Keep parser switches forward-compatible where the API uses extensible enumerations and preserve unknown cases for later analysis rather than crashing.

Define acceptance metrics before launch: cold and warm startup duration, hang rate, peak memory distribution, disk-write volume, parser success rate, deduplication rate, upload delay, and report retention. Compare release cohorts only when the app version, OS, time window, and population definitions are sound.

MetricKit turns real-device behavior into delayed, structured evidence. A trustworthy integration respects reporting cadence, treats diagnostics as asynchronous inputs, preserves symbolication provenance, limits data collection, and migrates away from deprecated APIs only when the supported OS and stable SDK make that boundary explicit.

Plan the migration as a compatibility layer

Keep the rest of the application independent of MetricKit’s concrete report types. Define a small internal representation for the measurements your product actually analyzes, such as a metric kind, numeric value or histogram, interval, app build, and source. Then write one adapter for the legacy subscriber callbacks and a separate adapter for the asynchronous MetricManager sequence when its runtime and SDK status are suitable. This limits migration work to parsing and delivery instead of spreading old MX types through analytics code.

Do not assume the two APIs expose identical schemas or delivery behavior. Map only fields with a verified semantic equivalent; version the internal payload and leave unsupported fields absent. In migration tests, feed the same fixture through both adapters and compare the normalized values that the product expects. Keep unknown enum cases, new metric results, and changed optionality from crashing the collector.

The newer documentation may evolve while symbols are marked beta. Confirm availability annotations in the SDK used to build the release, compile with the minimum deployment target, and run on the actual stable OS versions in scope. Do not infer that a source-code example compiles merely because the online page can display it. Mark any experimental branch separately in product release evidence.

Separate field signals from root-cause evidence

A rise in average launch duration is a signal, not an explanation. Segment by supported app version, OS version, feature state, and hardware class before drawing a conclusion. Keep sample counts and distribution shape alongside averages; a small number of slow launches can be hidden by a mean. For hang diagnostics, symbolicate and inspect the relevant thread state before assigning a regression to a particular subsystem.

Use a release canary process: define thresholds, wait for enough reports to make a comparison meaningful, and decide whether to pause rollout or investigate. Because daily reports are delayed and may include multiple versions, a release gate cannot depend solely on an immediate MetricKit callback. Pair it with synchronous build tests, local performance tests, crash monitoring appropriate to the product, and user-reported diagnostics.

Related:

Sources:

Comments