Skip to content
SRE & DevOpsDeep Dive Published Updated 9 min readViews unavailable

Go GOMAXPROCS in Kubernetes: CPU Limits, Cgroups, and Tail Latency

Tune and verify Go concurrency in Kubernetes by relating GOMAXPROCS to CPU quotas, Go module compatibility defaults, throttling, and changing limits.

Go applications can see the host’s logical CPUs even when Kubernetes has placed them in a container with a much smaller CPU limit. Before Go 1.25, the runtime’s default GOMAXPROCS generally followed the logical CPU count, so a process on a large node could try to run many goroutines concurrently while its Linux cgroup allowed much less average CPU time. The kernel then enforced the container’s CPU bandwidth limit by throttling it. That can turn a seemingly modest workload into bursty latency, especially when garbage collection or request fan-out creates short parallel spikes.

Go 1.25 introduced container-aware defaults on Linux: if no explicit override disables them, the runtime considers the cgroup CPU quota, CPU affinity, and logical CPU count when choosing GOMAXPROCS, and it can update the default when those inputs change. The Go runtime also rounds fractional CPU throughput limits up to an integer and normally does not select a value below two unless the logical CPU count or CPU affinity itself is below two. These are useful defaults, not a guarantee that the application will never be throttled. The Go team’s container-aware GOMAXPROCS article and current runtime documentation describe the behavior and its boundaries.

Keep CPU request, CPU limit, and GOMAXPROCS distinct

Kubernetes uses a CPU request primarily for scheduling and for a relative share of CPU under contention. A CPU limit is a throughput ceiling enforced by the Linux kernel through cgroups. GOMAXPROCS is an application-runtime parallelism setting: it limits how many goroutines can execute Go code simultaneously. These values are related, but they are not interchangeable.

For example, a Pod can request 500m CPU and have no CPU limit. It may receive more than its request when the node has idle CPU, so using the request as a hard GOMAXPROCS cap could discard useful available capacity. Conversely, a container can have a 500m CPU limit; its total CPU time is then constrained over the cgroup period even if it runs more than one goroutine at a time. The number of CPU threads is not the same thing as the total CPU time available over a period.

The Go runtime’s automatic default is based on the container CPU limit, not the Kubernetes request. This is why a Go service can behave differently after adding a limit even if its request stays the same. It is also why there is no universal rule that every production container should have a CPU limit: limits can make resource ceilings more predictable, but a too-tight limit can throttle bursts that are important to tail latency. Test the workload’s latency and throughput goals along with quota usage before deciding.

Verify the Go version and compatibility defaults

Using a new Go compiler alone may not activate every new runtime behavior for an older module. Go’s runtime compatibility defaults are influenced by the go directive in go.mod. Current runtime documentation notes that GODEBUG=containermaxprocs=0 and GODEBUG=updatemaxprocs=0 are the defaults for language version 1.24 and earlier. Those settings preserve the earlier logical-CPU behavior and disable periodic updates unless the program opts in or changes its compatibility baseline.

Review all of the following when diagnosing a container:

  1. The Go toolchain that built the deployed binary, not merely the local developer’s Go version.
  2. The go language version in the module’s go.mod file.
  3. The runtime GODEBUG value, including values compiled into the binary through build settings.
  4. Whether the process receives an explicit GOMAXPROCS environment variable.
  5. Whether application startup code or a library calls runtime.GOMAXPROCS with a positive value.
  6. The Pod’s actual CPU request and limit, and whether the node applies a CPU affinity mask.

An explicit GOMAXPROCS value or call to runtime.GOMAXPROCS(n) disables the runtime’s automatic default and updates. A startup helper that computes a fixed value from the cgroup can therefore become stale if Kubernetes later changes a Pod’s CPU limit. Remove or update such an override only after testing the actual binary and deployment configuration; the environment variable, Go source, runtime GODEBUG, and module directive can each affect the result.

Log what the running process selected

Expose the runtime setting in startup diagnostics or a metrics endpoint. The following minimal example reports the currently selected parallelism and the logical CPU count visible to the process:

package main

import (
	"fmt"
	"runtime"
)

func main() {
	fmt.Printf("GOMAXPROCS=%d logical_cpus=%d go=%s\n",
		runtime.GOMAXPROCS(0),
		runtime.NumCPU(),
		runtime.Version(),
	)
	// Start the application after recording runtime diagnostics.
}

Passing zero to runtime.GOMAXPROCS reads the current setting without changing it. That makes it appropriate for diagnostics; it does not prove why that value was chosen. Interpret it with the binary’s Go version and compatibility settings, GODEBUG, process environment, cgroup quota, and CPU affinity. runtime.NumCPU is not a substitute for reading container limits and does not report the amount of CPU time the process can consume under quota.

A workload template can declare container resources like this. This is a configuration excerpt, not a complete Pod manifest:

# Under spec.template.spec.containers[] (or spec.containers[] for a Pod)
resources:
  requests:
    cpu: "500m"
    memory: "512Mi"
  limits:
    cpu: "1500m"
    memory: "768Mi"

This example gives the Go runtime a Linux CPU quota corresponding to 1.5 CPUs, assuming the runtime sees the container’s cgroup. The current runtime documentation says non-integer quotas are rounded up when selecting an integer default; it also normally keeps GOMAXPROCS at least two when the process has at least two logical CPUs and an affinity mask of at least two. The kernel still enforces the 1.5-CPU throughput limit, so the default can permit some parallel bursts that are later throttled. Do not assume that a quota of 1500m means the runtime selects exactly 1 or 1.5 processors.

Diagnose throttling rather than guessing from CPU usage

A flat CPU utilization graph at the configured limit does not show whether the service’s user-visible latency is acceptable. Compare request latency and queue depth with container CPU throttling metrics, process GOMAXPROCS, garbage-collection activity, and the effective cgroup quota. The exact metric names and label sources depend on the node exporter or kubelet metrics pipeline; verify the metrics that your cluster actually exposes instead of copying a dashboard query without checking its source.

CPU usage below the limit does not by itself disprove a concurrency problem. A short CPU burst may be throttled within the kernel’s quota period while a one-minute average remains low. Likewise, setting GOMAXPROCS too low can reduce throughput even when the container has burstable capacity and the node is idle. Use representative load tests and p50, p95, and p99 latency to compare behavior across realistic request sizes, concurrency, and CPU-limit changes.

Check for cgroup v1 versus v2 only when investigating the low-level evidence. On cgroup v2, CPU quota and period are represented through the cpu.max interface; on cgroup v1, corresponding CFS quota and period files are used. Container runtimes and cgroup namespaces can expose different paths, so do not hard-code a host path into application logic. The runtime is designed to discover its cgroup from the process environment; operators should validate the process’ effective constraints rather than assuming a particular host mount layout.

When Kubernetes resizes a CPU limit in place, Go 1.25+ can periodically reconsider its default, provided the module compatibility defaults and runtime overrides allow it. A custom setting can disable that automatic adaptation. Test a real limit change in a staging cluster and log the setting both before and after the change. The fact that the Pod spec now contains a different limit does not prove that the process is using a new GOMAXPROCS value; confirm the applied container resources and runtime diagnostic independently.

A controlled rollout procedure

  1. Record the deployed binary’s Go version, go.mod directive, GOMAXPROCS, GODEBUG, CPU request and limit, and any CPU manager or affinity configuration.
  2. Establish a baseline under representative load: throughput, latency percentiles, restarts, CPU usage, and throttling evidence.
  3. Reproduce the same request mix in a non-production environment with the production CPU limit and cgroup configuration.
  4. Compare the runtime default with any explicit value only as a controlled experiment. Avoid changing the CPU limit and GOMAXPROCS simultaneously, since that obscures which change affected results.
  5. Test both steady load and short bursts. Measure garbage-collection-heavy requests and the service’s normal concurrency pattern.
  6. If the limit can change in place, verify runtime behavior after a resize; if the application overrides the default, decide whether to remove the override or update it deliberately.
  7. Roll out gradually and monitor latency, throttling, saturation, errors, and resource changes together. Roll back through the same source of truth that controls the Pod template or autoscaling policy.

Do not use a Pod’s CPU request as if it were a hard cap, and do not increase GOMAXPROCS just because a CPU chart looks low. First determine whether the bottleneck is throttling, too little concurrency, downstream blocking, lock contention, garbage collection, or a non-CPU resource. Each has a different fix.

Production checklist

  • The deployed Go version and go.mod language version enable the behavior you expect.
  • No environment variable, GODEBUG, startup helper, or library silently pins a stale GOMAXPROCS value.
  • The Pod’s CPU request and limit match the intended scheduling and throughput policy; the request is not mistaken for a quota.
  • The process reports its actual runtime parallelism and the setting is correlated with the deployed container version.
  • Performance tests measure latency percentiles and throttling under both sustained and bursty load.
  • A dynamic CPU-limit change is tested in staging before relying on the runtime to adapt automatically.
  • Operators can revert the Pod template or resource policy and can distinguish container throttling from application-level saturation.

Container-aware GOMAXPROCS makes Go’s default better aligned with CPU limits, but it does not decide the service’s resource policy or eliminate the effects of a constrained quota. Production tuning still depends on measured latency, available burst capacity, the exact compatibility defaults, and whether the application allows the runtime to adapt.

Related:

Sources:

Comments