Windows Service Control Handlers: Stop, Preshutdown, and Honest Status Reporting
Implement Windows service control handlers that return promptly, report pending checkpoints, honor shutdown budgets, and avoid blocking the Service Control Manager.
The Service Control Manager (SCM) delivers stop, shutdown, preshutdown, and other controls to a service through its registered control handler. That callback runs on a control-dispatch path, not on an arbitrary background worker. If it blocks on database flushes, network calls, or a lock held by a worker that is waiting for the handler, it can delay control delivery and make the service appear hung. A robust service treats the handler as a short state transition: record intent, report the new status when it changes, signal a worker, and return.
Register once and advertise only supported controls
ServiceMain registers a handler using RegisterServiceCtrlHandlerEx when extended control codes or a context pointer are needed. The service reports accepted controls in dwControlsAccepted through SetServiceStatus; the SCM uses that state to determine which controls can be sent. Do not claim to accept STOP, PAUSE, SHUTDOWN, or PRESHUTDOWN unless the implementation handles that path. All services must handle SERVICE_CONTROL_INTERROGATE, but it does not mean the service should invent a state transition merely because the SCM queried it.
Status reporting is a contract. During start or stop, use SERVICE_START_PENDING or SERVICE_STOP_PENDING with a checkpoint that advances while progress occurs and a realistic wait hint. Once work is done, report the stable state and clear pending-only fields. Do not send a status update from a worker after the service has already reported SERVICE_STOPPED; make one state owner responsible for transitions so races cannot publish contradictory states.
Keep the handler bounded
Microsoft documents a 30-second upper bound for the control handler to return. Lengthy stop work belongs on a service-owned thread or task that can be joined by the service’s main lifecycle path. The callback should request shutdown and return, while the worker closes listeners, drains requests, flushes only required state, and signals completion. Never hold the service’s main mutex while waiting for that worker if the worker needs the same mutex to exit.
DWORD WINAPI HandlerEx(DWORD control, DWORD, LPVOID, LPVOID context) {
auto* state = static_cast<ServiceState*>(context);
if (control == SERVICE_CONTROL_STOP) {
state->request_stop(); // Short, synchronized state change.
state->signal_stop_worker(); // Do not join or perform network I/O here.
}
return NO_ERROR;
}
The sketch leaves out registration, synchronization primitives, status reporting, and error handling. request_stop must be safe if controls race with normal work, and signal_stop_worker must not wait for the worker. The lifecycle owner - not the handler callback - should publish SERVICE_STOP_PENDING, advance checkpoints when meaningful progress is made, and eventually report SERVICE_STOPPED.
When a service accepts STOP and receives it, it must stop by transitioning to SERVICE_STOP_PENDING or SERVICE_STOPPED. For supported Windows versions, do not expect the SCM to send another ordinary control after STOP. Therefore, the stop path should be idempotent and complete even if an application-specific “prepare” control was never delivered.
Preshutdown and shutdown are not unlimited maintenance windows
SERVICE_CONTROL_PRESHUTDOWN is sent to services that advertised SERVICE_ACCEPT_PRESHUTDOWN. It occurs before ordinary shutdown notifications and can delay system shutdown while the SCM waits for the service to stop or the configured preshutdown timeout to expire. Use it only when there is a specific cleanup task that cannot be handled during normal operation or ordinary shutdown. Adding preshutdown support to perform routine cleanup can make every reboot slower and create failure risk under low battery or power-loss conditions.
The later SERVICE_CONTROL_SHUTDOWN notification has a constrained window. A default shutdown budget is short, and the system can continue after its configured limit even if a service remains busy. Do not increase WaitToKillServiceTimeout to conceal slow cleanup. Persist important state during normal operation; shutdown should flush only the delta that cannot reasonably be saved earlier. Avoid waiting on a remote server during shutdown, because network loss can consume the entire budget.
Do not assume dependent services remain available during the shutdown phase. The SCM does not use normal dependency ordering for the shutdown notifications by default; a service may discover that something it normally calls has already stopped. Design the shutdown action to work from local durable state or fail quickly with a clear, recoverable diagnostic.
Checkpoints are progress evidence, not a way to extend time forever
While a service is pending, the checkpoint tells the SCM that work is progressing; the wait hint estimates the interval before the next progress report. Advance the checkpoint only when a real phase or milestone has completed, not on a timer that masks a deadlock. A service should set a bounded timeout for every external operation and report which phase is still outstanding. If it cannot finish within the allowed shutdown period, preserve restart-safe state and stop blocking the system where possible.
Track the service lifecycle with explicit states such as accepting requests, draining, flushing, closing dependencies, and stopped. Reject new requests once draining begins. For long-running request handlers, use a deadline and cancellation policy, then decide whether to finish, abort idempotently, or write a recovery marker. A useful shutdown metric is time spent in each phase, not merely total process uptime.
Startup deserves the same rigor as shutdown. While a service is SERVICE_START_PENDING, report real checkpoints if initialization has distinct phases, such as opening a local database, loading a policy snapshot, and binding a listener. Do not advertise that the service accepts a control or request before the associated subsystem is ready. If initialization fails after partially acquiring resources, unwind those resources in reverse order, publish SERVICE_STOPPED with an appropriate Win32 exit code, and avoid leaving a worker thread running behind a failed service process state.
HandlerEx can receive controls that require extra registration or that a service has explicitly advertised. Device notifications, session changes, and user-defined controls should be handled as separate inputs to the state machine. Validate the control code, check whether it is valid for the current state, and make duplicate or late notifications harmless. For example, a stop arriving while the service is already draining should not start a second flush or enqueue an unbounded number of teardown tasks. A single atomic transition from Running to StopRequested is often a useful gate.
Keep SCM-facing work and application protocol work distinct. A successful return from ControlService says the control request was sent, not that the service has completed cleanup. Administrators should poll QueryServiceStatusEx for state and wait for the service to reach STOPPED within the caller’s own deadline. An installer should not kill the process immediately when it sees STOP_PENDING; it should report which phase is slow and apply its documented recovery policy. Likewise, a service should not fake completion by reporting STOPPED while background workers still hold resources or can mutate the service’s durable state.
Recovery configuration in the SCM can restart a failed service, but it does not repair an inconsistent shutdown protocol. Make startup idempotent: if the previous process was terminated during a write, detect incomplete work from a journal or marker and recover before accepting new requests. Ensure the stop path releases named mutexes, ports, files, and child processes in a repeatable order. Use service failure actions for process-level recovery only after the application can safely start from every interruption point.
Failure modes to test
Exercise STOP while the service is idle, under high request load, while a dependency is unavailable, during an in-progress durable write, and while a worker holds locks. Send duplicate stop requests through the administrative path and ensure they do not start a second teardown. Test PRESHUTDOWN and SHUTDOWN independently, since accepting one does not imply accepting the other. Simulate an operation exceeding its wait hint and confirm checkpoints reflect actual milestones. Restart after forced termination and verify the service can reconcile any incomplete work.
Log the control code, prior state, new state, checkpoint, phase duration, and failure result. Avoid logging secrets or user payloads during shutdown. Provide a diagnostic command that can report the current phase and outstanding worker count without trying to acquire a lock held by the stuck operation.
Automate a lifecycle test harness that installs or starts the service in an isolated test VM, waits for SERVICE_RUNNING, exercises a representative request, then sends stop and verifies SERVICE_STOPPED plus absence of child workers and leaked handles. Repeat with a dependency paused, an injected disk-full condition, and a worker that is deliberately delayed. For shutdown behavior, record the time between the control and final state and compare it with the advertised wait hints and configured timeout. Keep these tests separate from production service installation; they should not change registry timeout settings on a user machine.
Review every handler callback for forbidden blocking: synchronous network access, unbounded waits, locks with unclear ownership, UI prompts, and process joins. A stop command should remain usable even if the service is already degraded. During incident response, the operator needs predictable control behavior more than a lengthy shutdown attempt that never reaches a checkpoint.
The service control handler is a control-plane callback, not a place to perform the entire shutdown synchronously. Return promptly, make progress visible, keep work bounded, and treat system shutdown as an unreliable dependency boundary rather than as a guaranteed cleanup opportunity.
Related:
Sources: