GitHub Actions Ephemeral Runners: One-Job Isolation and Runner Groups
Build safer self-hosted Actions capacity with one-job ephemeral runners, scoped runner groups, external log retention, and autoscaling failure controls.
Self-hosted GitHub Actions runners execute repository-controlled workflows on infrastructure your organization owns. That control is useful when jobs need private networks, specialized hardware, or predictable capacity, but it also means workflow code can interact with the runner’s filesystem, process environment, network, and credentials. A runner that handles untrusted pull requests and then returns to a trusted deployment queue is a cross-job trust boundary, not merely a cost-saving worker.
Ephemeral runners reduce the persistence window by accepting one job and then being removed or discarded. They are not automatically secure: the VM or container still needs isolation, the workflow still needs least-privilege credentials, and external systems may remain reachable. The core objective is to make every job start from a known clean image, preserve required evidence outside the runner, and ensure teardown happens even when the workflow fails or the runner loses connectivity.
Model the runner as a job-scoped principal
A GitHub-hosted runner is managed by GitHub. A self-hosted runner is registered to an organization, repository, or enterprise and executes on your system. A persistent self-hosted runner typically retains host-level state between jobs: workspaces, caches, tool installations, Docker layers, credentials accidentally written to disk, and potentially malicious background processes. Deleting only the job workspace does not prove the host has returned to a safe baseline.
An ephemeral runner is registered with the one-job option and is expected to stop accepting additional jobs after completing one. For autoscaling, GitHub recommends an ephemeral architecture so an instance is not reused for a subsequent job. A controller such as Actions Runner Controller (ARC) can scale runner scale sets on Kubernetes; other supported designs can create short-lived VMs or containers. Choose one lifecycle owner and ensure the instance is destroyed even when the runner process exits unexpectedly.
One job per runner is a boundary, not a complete containment guarantee. A workflow can still exfiltrate credentials or reach internal services during that job. Use network egress controls, minimal instance roles, short-lived job tokens, secret scoping, and separate runner pools for different trust levels. Public repositories that accept contributions from forks should not run untrusted code on a runner that can reach production networks or has durable cloud credentials.
Separate runner groups by trust and capability
Runner groups let administrators restrict which repositories or organizations can use a runner pool. Use groups to express real security boundaries: public contribution tests, internal build jobs, privileged release jobs, and hardware-specific workloads should not share a pool merely because they use the same operating system image. Repository allowlists should be reviewed like access-control policy, and changes to group membership should be auditable.
Labels route jobs to compatible runners, but labels are not a substitute for access control. A workflow author may request a label; if the runner group is available to that repository, the job can be scheduled there. Keep privileged runners in a group with narrowly selected repositories, and protect workflow files with branch rules and code review. Do not expose an administrator-controlled group to a public repository because its workflows request a convenient label.
The trust classification should include the trigger, not just the repository’s visibility. Fork pull requests, issue comments, dispatch events, release workflows, and deployment workflows can have different secrets and permissions. A private repository may still run code from an external contributor, while a public repository’s trusted default-branch push has a different provenance. Build a small policy matrix that identifies event, source ref, runner group, GITHUB_TOKEN permissions, cloud identity, network routes, and allowed secrets.
Provision from a clean, versioned image
Bake the runner image or launch template from reviewed infrastructure code. Pin operating-system packages, toolchain versions, the runner binary version or update policy, and the bootstrap process. Avoid copying a developer workstation image with cached cloud credentials or broad SSH keys. Store registration tokens only in the control plane that needs them, use short-lived registration material, and do not bake it into an image or user data log.
A robust lifecycle is: create an isolated instance; fetch a short-lived runner registration token; register one ephemeral runner with the intended group and labels; execute one job; stream or upload required logs and artifacts; revoke or expire runner registration; terminate the compute; and report lifecycle state to the autoscaler. The controller should have narrowly scoped permissions to create and destroy instances and manage runner registrations. An instance should not have permission to modify the autoscaler or access sibling runner state.
apiVersion: actions.github.com/v1alpha1
kind: AutoscalingRunnerSet
metadata:
name: ci-ephemeral
namespace: actions-runner-system
spec:
githubConfigUrl: "https://github.com/actions/runner-images"
runnerGroup: "untrusted-ci"
minRunners: 0
maxRunners: 20
template:
spec:
ephemeral: true
labels:
- linux-x64
- untrusted-ci
containers:
- name: runner
image: ghcr.io/actions/actions-runner:REPLACE_WITH_APPROVED_VERSION
command: ["/home/runner/run.sh"]
This illustrates intent and is not a universal ARC manifest: API versions, chart values, generated Custom Resource schemas, and container images depend on the installed ARC release. The GitHub URL is a real public repository solely to keep the example syntactically concrete; replace it with a repository or organization you administer. Validate against the exact ARC CRD and official deployment guide before applying. Use a digest-pinned image in production, store GitHub credentials through the supported secret integration, and set a namespace and node policy that isolate runner Pods. The word ephemeral in a manifest is meaningful only if the controller and runner application implement the expected lifecycle.
Preserve logs without preserving the machine
Ephemeral hosts disappear by design, so runner diagnostic and job logs needed for incident response must be forwarded externally. Preserve controller events, runner registration and removal records, job identifiers, instance IDs, image digests, workflow commit SHAs, and timestamps in a centralized system. Keep workflow artifacts separate from runner caches: artifacts are explicit job outputs, while a cache is a reusable optimization that can cross run boundaries.
Do not archive whole workspaces indiscriminately. Logs and artifacts can contain tokens, environment dumps, request bodies, customer data, or generated credentials. Define retention and access controls, redact secrets before upload, and avoid shell tracing around secret-bearing commands. Keep a tamper-resistant audit path for runner creation and teardown, and alert on runners that remain registered or instances that outlive their job beyond a bounded cleanup window.
GitHub’s Actions UI records workflow and job logs, but an ephemeral host can fail before the final log upload. Use infrastructure-level logging for bootstrap and shutdown, and capture the runner’s service logs with a mechanism that does not depend on a successful workflow step. Ensure the external sink is reachable with a narrowly scoped identity and that failure to log does not silently convert a security control into a no-op.
Scale without allowing noisy-neighbor failures
Autoscaling uses queued-job demand and your capacity policy to create runners. Set maximum capacity based on cloud quotas, cost limits, subnet address availability, registry throughput, and downstream service capacity. Sudden runner scale-out can overload artifact storage, package mirrors, Kubernetes clusters, or a private database even when GitHub’s queue looks healthy. Use a bounded scale-up rate, sensible minimum capacity for latency-sensitive trusted jobs, and separate pools if jobs have different startup or security requirements.
Scale-down must not terminate an instance in the middle of a job. Use the controller’s supported drain and termination lifecycle, and distinguish idle runners from busy runners. A job timeout that exceeds the instance lease, spot interruption behavior, and controller resync intervals can all produce abandoned work. Test cancellation, runner process crash, host termination, and controller restart. A complete design treats these as ordinary lifecycle events rather than rare anomalies.
Security checks before rollout
Before enabling a pool, prove that an untrusted test job cannot read a trusted job’s filesystem, credentials, Docker socket, network namespace, or cloud metadata. Confirm that no later job is scheduled on the same host. Test repository access with a repository outside the allowlist and verify that GitHub refuses to dispatch to the group. Check whether Docker-in-Docker or privileged Kubernetes Pods create a path from workflow code to the host kernel or neighboring workloads.
Audit every environment variable and token source. The runner may receive a per-job token, but instance profiles, mounted Kubernetes service-account credentials, cloud CLI profiles, package registry credentials, and bootstrap secrets can be more powerful. Use workload federation for cloud access when available and scope it to the exact repository, ref, environment, and workflow identity. Never assume that ephemeral: true revokes external credentials leaked during the job.
Finally, pin and review every action by a full commit SHA where your policy requires it, guard workflow definitions with code ownership, and keep the runner software and base image patched. Separate build and deployment pools if a build process executes untrusted code. A well-designed ephemeral runner has a narrow authorization envelope, predictable single-job lifetime, no reusable state, and an auditable deletion path.
Related:
- How to Set Up a CI/CD Pipeline with GitHub Actions
- GitHub Actions Reusable Workflows: Contracts, Permissions, and Safe Composition
Sources: