Skip to content
SRE & DevOpsDeep Dive Published Updated 6 min readViews unavailable

AWS Fargate: Serverless Container Compute for ECS and EKS

What Fargate manages, how task-level isolation and resource sizing work, and when its operational simplicity outweighs reduced host control.

AWS Fargate is a compute engine for containers used by Amazon ECS and Amazon EKS. It is not a separate orchestrator: ECS services/tasks or Kubernetes pods still express desired work, while Fargate supplies isolated compute without exposing an EC2 fleet for the customer to patch or scale.

The unit of capacity changes

With EC2-backed containers, teams buy instances and then pack workloads onto them. With Fargate, they request supported CPU, memory, ephemeral storage, platform, and architecture combinations per task or pod. This removes host bin-packing and idle-node management, but poor requests can still waste money or cause throttling and out-of-memory failures. Measure real workload demand before standardizing sizes.

Security and networking

Fargate tasks use awsvpc networking and receive their own network identity. Security groups, private subnets, VPC endpoints, and least-privilege task roles remain customer responsibilities. The isolation boundary is stronger than placing unrelated containers on one customer-managed host, yet application vulnerabilities, permissive IAM, exposed listeners, and compromised images remain fully relevant.

Operations and observability

There is no host to SSH into. Diagnostics therefore depend on container logs, ECS or Kubernetes events, metrics, distributed traces, and optional runtime debugging mechanisms. Images must start cleanly, handle termination signals, expose meaningful health checks, and externalize persistent state. Platform-version changes and supported feature matrices should be reviewed before relying on a kernel capability or storage behavior.

When to choose it

Fargate works well for variable services, scheduled tasks, and teams that do not need privileged containers, host networking, custom kernel modules, or specialized instance tuning. Stable, large fleets can be cheaper or more flexible on carefully managed EC2 capacity, especially with reservations or Spot. Compare the complete operating cost and risk, not only compute price.

ECS tasks and EKS pods are not identical products

On ECS, the task definition selects a supported Fargate CPU and memory combination, runtime platform, ephemeral storage, networking, roles, and containers that share the task boundary. On EKS, a Fargate profile selects eligible pods by namespace and optional labels, and the platform maps each pod to Fargate capacity. Kubernetes daemon sets, privileged pods, and some host-oriented assumptions do not translate to that model. Check the current ECS and EKS feature matrices separately instead of assuming a Fargate feature in one orchestrator exists in the other.

Platform versions and architecture choices matter. For ECS on Fargate, Linux ARM64 requires Fargate platform version 1.4.0 or later; AWS documents an Availability Zone limitation in us-east-1, so verify current Region and Availability Zone requirements before promotion. EKS Fargate has a separate feature matrix, so an ECS architecture option must not be assumed to apply to Pods on EKS Fargate. Publish a multi-platform manifest only when every referenced image was built and tested for each advertised platform. Image size and initialization affect startup latency, while task resource allocation and runtime duration are inputs to Fargate cost; check the current pricing model rather than assuming image-transfer time is billed as active compute.

Resource sizing is the capacity plan

Fargate enforces task-level CPU and memory combinations. Container-level reservations and limits must fit inside the task. CPU starvation, memory leaks, large image extraction, temporary files, and log buffering can still cause failures even though there is no node to manage. Load-test representative requests, observe peak resident memory and throttling, and include sidecars and startup spikes. Set an explicit ephemeral-storage budget and export durable state to an external service.

Autoscaling does not resize a running task. ECS Service Auto Scaling or Kubernetes controllers add and remove replicas based on metrics and policy. Scale-out latency includes placement, image pull, and application readiness, so keep images small and choose a minimum capacity that matches the latency objective. Scale-in must respect connection draining and termination handling. Downstream concurrency, not just frontend CPU, should cap the maximum replica count.

Network path and least privilege

Every ECS Fargate task uses awsvpc networking and receives an ENI. Select subnets with enough addresses and routes, and give the task only the security-group flows it needs. A task in a public subnet does not automatically receive a public IP unless configured; a private task needs NAT or appropriate VPC endpoints to reach required services. ECR image pulls can also depend on S3 and ECR endpoints according to the documented path.

Keep the ECS task execution role separate from the task role. The execution role is used by the ECS/Fargate agent for operations such as pulling images and sending logs; the task role is made available to application code. On EKS Fargate, the Pod execution role authorizes infrastructure actions such as image pulls and log routing, but containers in the Pod cannot assume that role. AWS directs workloads on EKS Fargate to use IAM Roles for Service Accounts (IRSA) for application access to AWS services. These role boundaries are not interchangeable, and a successful image pull does not prove that the application has the intended AWS permissions. Secrets should come from a managed source through least-privilege references, and sensitive values must not be written to logs or baked into image layers.

For ECS, make those identity and architecture choices visible in a reviewed task definition. This compact example uses placeholders for account-specific role ARNs and image URI; it is illustrative configuration, not a deployable account manifest:

{
  "family": "orders-api",
  "requiresCompatibilities": ["FARGATE"],
  "networkMode": "awsvpc",
  "cpu": "512",
  "memory": "1024",
  "executionRoleArn": "arn:aws:iam::123456789012:role/ecsTaskExecutionRole",
  "taskRoleArn": "arn:aws:iam::123456789012:role/ordersApiTaskRole",
  "runtimePlatform": {
    "operatingSystemFamily": "LINUX",
    "cpuArchitecture": "ARM64"
  },
  "containerDefinitions": [
    {
      "name": "api",
      "image": "123456789012.dkr.ecr.us-east-1.amazonaws.com/orders@sha256:REPLACE_WITH_DIGEST",
      "essential": true
    }
  ]
}

Before using a definition like this, verify the CPU/memory pair against the current Fargate task-size table, use a real digest, and confirm that the task execution role has only the image-pull and logging permissions it needs. The application role should be narrower still and should not inherit deployment or infrastructure-administration permissions. When moving from x86-64 to ARM64, rebuild native dependencies, scan the image, and run representative integration and performance tests on the target architecture; matching manifest metadata alone is not evidence that native code is compatible.

Diagnosis without host access

Start with orchestrator evidence: ECS stopped-task reasons and service events, or Kubernetes pod status and events. Then inspect container exit codes, health checks, logs, CPU and memory metrics, ENI and DNS behavior, IAM denials, and image architecture. ECS Exec can provide controlled interactive access when it is deliberately enabled and authorized, but it should not replace logs, metrics, traces, or a reproducible debug image.

Maintain runbooks for image-pull failure, OutOfMemoryError, failed health checks, subnet exhaustion, dependency timeout, and stuck draining. A serverless compute plane removes host patching; it does not remove the need to know which artifact ran, which identity it used, which network path failed, and how to reproduce the task configuration.

For cost review, separate steady provisioned demand from burst demand and include public IPv4 charges where applicable, NAT processing, logs, data transfer, load balancers, and idle minimum replicas. Compare those measured costs with an EC2 capacity model that includes patching, autoscaling, unused headroom, and on-call labor; a raw vCPU price comparison omits the operational trade Fargate is designed to make.

Related:

Sources:

Comments