Skip to content
SRE & DevOpsDeep Dive Published Updated 8 min readViews unavailable

AWS App Runner: From Source or Container Image to a Managed Web Service

How App Runner builds, deploys, scales, secures, and observes HTTP services-and where its simplified platform stops being the right abstraction.

AWS App Runner is a managed application platform for HTTP services. It can deploy from supported source repositories or from a container image, create a TLS endpoint, run health checks, and scale instances without asking the team to model ECS clusters, services, load balancers, or node capacity directly.

Deployment model

A service points to source or an image plus runtime configuration. Each deployment applies the selected source or image and service configuration; record the source commit or immutable image digest together with the deployment operation so the running release can be identified later. Automatic deployments can follow repository or image changes, but the trigger behavior depends on the configured source and deployment settings. Convenience does not replace release discipline: review who can change the watched source, prefer immutable image digests for promotion, and verify application health before treating a deployment as successful.

Runtime, scaling, and cost

App Runner’s autoscaling configuration uses maximum concurrency and instance-count bounds. Maximum concurrency is the request-concurrency threshold used to scale an instance; minimum size establishes provisioned capacity, while maximum size limits active instances. Minimum capacity can include provisioned instances that are not actively handling requests, so budget for that baseline rather than assuming only busy instances affect cost. App Runner bills memory for all provisioned instances and CPU for the active subset; a deployment can temporarily double provisioned instances to maintain capacity for old and new code. These settings protect latency and downstream dependencies only when they match the application: a database that accepts fifty connections will not become safe because the frontend can scale to hundreds of instances. Model connection pools, rate limits, queue backpressure, and rollout overlap explicitly.

Networking, identity, and secrets

The service can use an instance role for AWS API access. Secrets and parameters should be referenced from managed stores rather than baked into images or source. App Runner has separate controls for incoming and outgoing traffic: its default public endpoint can be changed to a Private endpoint backed by a VPC interface endpoint and VPC Ingress Connection, while a VPC Connector controls outbound access from the service to resources in a VPC. These are distinct network paths, not one setting with two directions. Check the current endpoint constraints and quotas before selecting App Runner for a complex service-mesh or multi-network topology.

Observability and fit

Application logs and service metrics belong in CloudWatch; add correlation IDs and tracing inside the application. App Runner is strongest for straightforward web services and APIs that benefit from a narrow operational surface. ECS, EKS, or Lambda can be better when a workload needs background workers, detailed networking, specialized compute, event-native invocation, or control over deployment primitives.

Build source and image source have different trust paths

For a source-code service, App Runner receives repository access through a connection and builds with the selected managed runtime and configuration. The release therefore depends on source revision, build instructions, runtime version, dependency resolution, and the permissions of the connection. Pin dependency locks, review runtime support, and capture build logs. For an image-based service, the build occurs elsewhere; App Runner needs an access role to retrieve a private ECR image, and the publisher must preserve the digest that passed testing.

Automatic deployment is convenient but broadens the release trigger. Decide whether every source commit or image update is production-worthy, especially when a mutable tag is used. A safer pipeline can publish an immutable image, verify it, and explicitly update the service. Whatever path is chosen, record the deployment operation, source version, image digest where applicable, configuration revision, and health result so rollback is based on evidence rather than the latest repository state.

Request lifecycle and scaling boundaries

App Runner routes HTTP traffic to running instances and scales according to its autoscaling configuration. Maximum concurrency is the request threshold used to drive scaling, while minimum and maximum size constrain provisioned baseline and active capacity. These settings must reflect application concurrency, memory consumption, startup latency, and dependency capacity. A high concurrency value can hide saturation inside an instance; an aggressive maximum can overwhelm databases or external APIs. Set the threshold from load-test evidence: measure tail latency, queueing, CPU and memory, and dependency-pool utilization at realistic request mixes, then verify that scale-out begins before the service-level objective is breached. Recheck the calculation after changing runtime, instance size, or connection-pool defaults.

Health checks can use TCP or HTTP. An HTTP endpoint should be inexpensive and should distinguish whether the process is ready to serve without making every transient dependency failure restart the service. Decide whether readiness means that the process can accept requests or that a critical dependency is available; an endpoint that fans out to every dependency can turn a partial outage into a fleet-wide restart loop. Handle termination signals, stop accepting new work, and bound request duration. App Runner is designed around web requests; durable asynchronous work should be acknowledged into a queue and processed by a platform with an explicit worker lifecycle rather than kept alive only in an HTTP request. Test graceful shutdown while requests are active and confirm that clients can safely retry interrupted requests.

Network and identity review

An instance role grants the running application AWS API access. Keep it separate from the ECR access role used by the service to retrieve an image. Reference Secrets Manager secrets or Systems Manager Parameter Store parameters through supported runtime configuration and grant only those resources. App Runner resolves referenced values during deployment; changing a secret or parameter does not automatically update the environment of a running service. Plan a redeployment after rotation, or have application code retrieve the value through the SDK with a deliberate refresh and cache policy. Do not assume that changing the secret store alone has rotated credentials for existing processes. Also remember that plain-text environment variables are not encrypted in the service configuration and can appear in application logs, so reserve them for non-sensitive settings.

A VPC Connector affects outbound traffic from the service to resources in a VPC; it is not the control for inbound access. For private ingress, enable App Runner’s Private endpoint and connect it to a VPC interface endpoint through the App Runner VPC Ingress Connection. The service then accepts incoming requests through that VPC path rather than the public internet. Plan endpoint security groups and subnets across Availability Zones. For this private incoming endpoint, AWS documents that VPC endpoint policies are not supported; use endpoint security groups to control the network path. AWS also warns that WAF source-IP rules do not receive original client IP data for private App Runner services, so use security-group rules for source CIDR controls. For outbound traffic, confirm connector subnet capacity, security groups, DNS, routes, and NAT or service endpoints for external calls. Separately test both request directions; a successful private inbound path does not prove egress to a database, AWS API, or public endpoint is configured correctly.

Operability and exit criteria

Monitor request count, latency, HTTP status classes, active instances, CPU and memory, deployment failures, and application-specific service-level indicators. Logs need request IDs and trace context because platform metrics cannot explain a logically incorrect response. Set alarms on sustained errors and saturation, not on every scale event. Test a bad deployment, a slow-starting revision, loss of a dependency, and a rollback before the service becomes critical. Preserve the deployment identifier, source revision, image digest, and relevant configuration in the incident timeline. A useful synthetic check exercises a real user path while a shallow health check remains cheap enough for frequent readiness probes. Keep dashboards separated by service and environment, and alert on user-visible symptoms as well as platform resource metrics.

App Runner’s abstraction is valuable only while its contract matches the workload. Re-evaluate when the service needs sidecars, non-HTTP protocols, long-running consumers, custom scheduling, privileged access, specialized accelerators, advanced traffic policy, or a service mesh. An explicit exit criterion prevents teams from forcing unsupported patterns into a platform merely to avoid migrating to ECS or EKS.

Document quotas and regional dependencies in the service runbook. Before a traffic event, verify the current service, build, connection, networking, and autoscaling quotas in the target Region, request increases where supported, and load-test the complete dependency chain rather than only the App Runner endpoint. Record a clear exit trigger for the platform: requirements for unsupported protocols, host-level controls, specialized hardware, or custom traffic policy should prompt an explicit move to ECS, EKS, or another fit-for-purpose service instead of accumulating fragile workarounds.

Related:

Sources:

Comments