Skip to content
SRE & DevOpsDeep Dive Published Updated 7 min readViews unavailable

AWS Secrets Manager Rotation: Idempotent Stages and Safe Credential Cutover

Operate Secrets Manager rotation through AWSPENDING tests and AWSCURRENT promotion, handling retries, dual-user strategies, network access, and audit evidence.

AWS Secrets Manager rotation is a coordinated update between a stored secret and the service that accepts it. A schedule alone does not make rotation safe. The rotation function must create a candidate value, apply it to the target, verify that the target accepts it, and only then promote that value to AWSCURRENT. Retries, partial failures, multiple application instances, cached credentials, network isolation, and database user strategy all affect whether clients continue to work during that transition.

For Lambda-based rotation, Secrets Manager invokes a function through four logical steps: createSecret, setSecret, testSecret, and finishSecret. Staging labels identify which version is pending and current. The client request token is a stable identifier for a rotation attempt and supports idempotent retries. A robust rotation implementation can receive a step again after an error without creating inconsistent credentials or accidentally changing a different target resource.

Follow the four-step state machine

createSecret checks whether the version for the request token already exists, generates or obtains a new value when needed, and stores it with the AWSPENDING label. Reusing the same token and staging state on retries makes this step idempotent. Do not generate a second unrelated password on every retry; that can leave the stored pending value different from the value the target service received.

setSecret changes the database or external service to accept the pending value. This step must verify that the current and pending versions refer to the same intended resource before making a privileged change. For a database, validate the endpoint and username; for a remote API, validate the target account or tenant. Secrets Manager warns that a rotation Lambda is a privileged deputy with access to both the secret and target, so arbitrary user-controlled secret fields must not redirect the function to another resource.

testSecret authenticates to the target using the pending value and performs a safe operation. The test should prove the new credential works without changing business data. finishSecret moves AWSCURRENT to the new version and leaves the previous one as AWSPREVIOUS. Avoid manually removing AWSPENDING in a separate call; a leftover or inconsistent stage can make later attempts interpret a previous rotation as still in progress.

{
  "Step": "setSecret",
  "SecretId": "arn:aws:secretsmanager:REGION:ACCOUNT_ID:secret:service-credential",
  "ClientRequestToken": "ROTATION_VERSION_ID",
  "RotationToken": "OPTIONAL_CROSS_ACCOUNT_TOKEN"
}

This shows the shape of the invocation event, not a deployable credential or a full Lambda handler. Values are identifiers and placeholders; never log secret values or the raw request when it contains sensitive fields. For cross-account or assumed-role rotation, handle the rotation token according to the current AWS contract so Secrets Manager can validate the source role. Use the official rotation template as a starting point for supported secret types rather than inventing a custom state machine without retry tests.

Design idempotency and concurrent-client behavior

Rotation may be retried if a step fails. Each step should inspect staging labels and current target state before performing a change. An idempotent setSecret can detect that the target already accepts the pending credential and return success rather than resetting it unpredictably. A retry after setSecret but before finishSecret is a normal partial state, not necessarily a reason to revoke the new value.

Single-user rotation changes one account’s password. Some clients with pooled connections may continue to use an old authenticated session, while new connections with a cached old password fail. Dual-user rotation alternates between two accounts and can reduce downtime because the active user is not modified in place, but it introduces account synchronization, privilege parity, and username lifecycle concerns. Choose the strategy based on the target engine and application connection behavior, not just the rotation Lambda template.

Applications often cache secrets for performance. Rotation does not force every process to reload a value when AWSCURRENT moves. Set a cache lifetime compatible with the rotation schedule and define what a client does when authentication fails. Ensure the target can tolerate an overlap window in which both old and new connections exist, or drain and recycle connection pools during cutover. Test this against the actual application client libraries.

Grant the rotation function only necessary access

The Lambda execution role needs permission to read and update the selected secret, use its KMS key where applicable, and reach the target service. Scope permissions to specific secret ARNs, key ARNs, and network paths. The Secrets Manager service principal also needs permission to invoke the rotation function through its resource policy. AWS recommends using source-account conditions to reduce confused-deputy risk; determine whether an additional source-ARN condition is appropriate for a function dedicated to one secret.

The rotation function needs network connectivity to both Secrets Manager and the database or service. If it runs in a VPC, verify subnets, security groups, DNS, routes, and VPC endpoints or NAT. A Lambda placed in private subnets can successfully call the target database but fail to reach Secrets Manager, or the reverse, depending on endpoint configuration. Validate TLS and certificate behavior for the target. Do not expose connection strings or pending values in logs while debugging a network path.

Separate who can configure rotation from the function’s runtime role. Permissions such as creating an IAM role and attaching policies can allow privilege escalation, so only a controlled deployment identity should create and modify rotation infrastructure. Protect the function package, dependency lock data, layers, and resource policy as production security assets. Patch bundled libraries and external binaries because the function code is part of the secret trust boundary.

Configure schedules and understand retry windows

Secrets Manager can run rotation on a schedule and retry the full rotation sequence when a step fails during the open rotation window. A scheduled rotation is not guaranteed to finish at an exact wall-clock moment. Rotation windows use UTC, begin on the hour, and must not extend into the next window. Keep the window long enough for network and service recovery. Monitor rotation status and CloudTrail events; alert on repeated failures and on a pending version that remains unresolved beyond the normal window.

Use describe-secret and the VersionIdsToStages mapping to inspect state without retrieving the secret contents. CloudWatch logs can help diagnose a function step, but redact request details and use safe identifiers. AWS records RotationFailed as a Secrets Manager service event through CloudTrail; an EventBridge rule can route that event to an alert or incident workflow. Investigate the failing stage, target authentication, KMS permission, function timeouts, Lambda execution role, and network path. Avoid immediately disabling rotation for all secrets when one target is failing; first determine whether the issue is isolated.

Test the full transition before production

Create a disposable secret and target with the same secret JSON shape, database driver, VPC path, and application cache behavior. Run each step repeatedly with the same client token to prove idempotency. Inject failures after createSecret, after target update, during testSecret, and before finishSecret. Confirm that retries converge to one valid current version and that an operator can identify the previous known-good credential without exposing it.

Test the consumer side too. Start a representative application version, rotate its credential, and observe whether new connections use the updated secret before old authentication stops working. Validate read/write behavior with a least-privilege test action. For a database, confirm connection pools recover and the old user or password is retired only when policy and application behavior permit. Document the recovery path if the target accepted the new password but the secret stage did not advance.

Troubleshoot without exposing the credential

If no rotation activity occurs, check the rotation schedule, secret configuration, Lambda invocation permission, and CloudTrail. If createSecret fails, inspect KMS access, version-token idempotency, and password-character constraints. If setSecret fails, verify that AWSCURRENT and AWSPENDING identify the same resource and that the function can reach and authenticate to the intended target. If testSecret fails, distinguish a bad candidate value from a permission or network issue. If finishSecret is stuck, inspect the staging labels and preserve the pending value until state is understood.

Do not fetch and print the secret to compare strings. Compare version IDs, labels, resource identifiers, and authentication outcome. Rotate or revoke any credential that may have entered logs or shell output. After recovery, run the rotation again in a test secret, confirm monitoring detects the partial state, and update the runbook with the actual failure boundary.

Rotation is a state transition between a secret store and a real credential consumer. Reliable operation requires idempotent step logic, target validation, least-privilege runtime roles, functional network connectivity, client cache planning, and observable staging labels. A green schedule without a verified application login is not proof that rotation works.

Related:

Sources:

Comments