Skip to content
SRE & DevOpsDeep Dive Published Updated 10 min readViews unavailable

GitLab CI Production Pipelines: DAG Dependencies, Artifact Contracts, and Deployment Locks

Design GitLab CI DAGs that avoid duplicate pipelines, transfer explicit artifacts, and serialize production deployments without hiding release races.

GitLab CI pipelines are often introduced as a sequence of stages, but production workflows have several different contracts hidden inside that sequence: which events are allowed to create a pipeline, which jobs must finish before another can start, which exact files cross a job boundary, and how concurrent deployments interact with the same environment. A pipeline can be green while one of those contracts is wrong. It might have built the same change twice, tested a different artifact from the one later deployed, or allowed two jobs to mutate production concurrently.

The goal is not to replace every stage with a complex directed acyclic graph (DAG). Stages remain a useful visual grouping and a simple default ordering. Use needs when a real producer-consumer dependency can safely start before an entire prior stage finishes. Keep the graph small enough that a reviewer can answer, from the YAML, what evidence authorizes a deployment and which artifact is being deployed.

Decide which events create a pipeline

GitLab evaluates workflow before it evaluates individual jobs. A job-level rule cannot restore a pipeline that the workflow rejected. This makes workflow: rules the right layer for deciding whether a branch push, merge request, tag, schedule, API request, or downstream trigger should create a pipeline at all.

A common source of wasted work is allowing both a branch pipeline and a merge request pipeline to run for the same change. The following policy runs merge request pipelines when a merge request is open, otherwise runs branch push pipelines. It deliberately does not admit tags, schedules, web/API pipelines, or downstream pipelines; add those sources explicitly if the repository needs them.

workflow:
  rules:
    - if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
    - if: '$CI_COMMIT_BRANCH && $CI_OPEN_MERGE_REQUESTS'
      when: never
    - if: '$CI_PIPELINE_SOURCE == "push" && $CI_COMMIT_BRANCH'

Order matters: rules are evaluated in order, and the first matching rule determines the result. The merge-request event is accepted first. A push to a branch that already has an open merge request is then suppressed; a push to a branch without one is accepted by the last rule. This is a policy choice, not a universal template. If a repository uses scheduled maintenance, release tags, parent-child pipelines, or multi-project triggers, name those sources deliberately and test that the workflow admits them.

Avoid adding a final broad when: always just to make a missing pipeline appear. It can admit event types the team did not intend and can recreate duplicate branch/MR work. When a pipeline is unexpectedly absent, inspect its event source and workflow decision before weakening the policy.

Draw the dependency graph around real outputs

Without needs, a job in a later stage ordinarily waits for the earlier stage to finish. With needs, it waits for named jobs and can run as soon as those dependencies complete. This can shorten feedback in a monorepo or a pipeline with independent components, but it also makes the named edges the actual correctness boundary. A deployment that needs only a package job has not necessarily waited for the integration tests or policy checks that should authorize it.

The example below is intentionally generic. Replace every ./ci/... command with a reviewed script in the repository. It builds one package, tests that package, and allows production deployment only after the required integration test succeeds. The security and policy gate is a required dependency rather than an optional job that could silently disappear from the graph.

stages:
  - verify
  - package
  - test
  - deploy

lint:
  stage: verify
  needs: []
  script:
    - ./ci/lint.sh

package:
  stage: package
  needs:
    - job: lint
  script:
    - ./ci/build-release.sh
    - ./ci/write-build-manifest.sh "$CI_COMMIT_SHA" release/manifest.json
  artifacts:
    name: "release-${CI_COMMIT_SHORT_SHA}"
    paths:
      - release/
    expire_in: 7 days
    when: on_success

integration:
  stage: test
  needs:
    - job: package
      artifacts: true
  script:
    - ./ci/test-release.sh release/

policy:
  stage: test
  needs:
    - job: package
      artifacts: false
  script:
    - ./ci/check-release-policy.sh "$CI_COMMIT_SHA"

deploy_production:
  stage: deploy
  needs:
    - job: integration
      artifacts: false
    - job: policy
      artifacts: false
    - job: package
      artifacts: true
  resource_group: production
  environment:
    name: production
  rules:
    - if: '$CI_PIPELINE_SOURCE == "push" && $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH'
      when: manual
    - when: never
  script:
    - ./ci/verify-release-manifest.sh release/manifest.json "$CI_COMMIT_SHA"
    - ./ci/deploy-production.sh release/
    - ./ci/smoke-production.sh

The needs: [] declaration lets linting start immediately because it has no build output dependency. The package waits for lint. Integration and policy work can then run independently, and deployment waits for both. The graph encodes a release gate; it is not merely a speed optimization. If a check is optional, document why and model its absence explicitly. needs:optional is appropriate only when the needed job may genuinely not exist in that pipeline and the downstream job is still safe without it. It is not a shortcut for making a required check green when its configuration is broken.

Treat artifacts as explicit job interfaces

Jobs do not share a working directory. Artifacts are one mechanism for handing outputs from a producer to a later job. If a job uses needs, declare whether it should download each producer’s artifacts. In the example, integration and deployment receive the package; the policy job needs only the producer’s success and downloads no files. That is easier to review than implicitly downloading every earlier artifact.

Use a narrow paths list and fail the producer if a required output is missing. A successful build command that wrote files to the wrong directory should not publish an empty or misleading artifact as if it were a release. Keep names stable enough for consumers, and include a commit identity in the manifest so the consumer can reject a mismatch between the tested source and the source selected for deployment.

expire_in is a retention setting, not an archival guarantee. The instance or project may impose a maximum, and artifact availability also depends on the configured storage and retention policies. Temporary job artifacts are useful for handoffs and evidence; they should not be the only durable copy of a release that must remain deployable months later. For container delivery, a registry image referenced by digest is often a more appropriate promotion unit than rebuilding an image in the deployment job. For other packages, publish to a package or artifact repository with a retention and immutability policy suited to the release process.

The manifest should identify what was built, not merely repeat a mutable tag. At minimum, validate the expected commit and artifact identity before deployment. Depending on the release system, record a digest, version, build timestamp, and relevant test evidence. Do not treat a filename such as latest.tar as proof of provenance.

Serialize changes to a shared environment

Two pipelines can pass their tests at nearly the same time and both attempt to deploy to production. resource_group provides mutual exclusion for jobs using the same resource group, so a second deployment job waits rather than running concurrently with the first. It does not prevent a human, another CI system, or a separate resource-group name from changing the same environment. The lock must represent the real shared target and be used consistently by every deployment path.

Serialization also does not decide which queued deployment should run first. GitLab resource groups have a configurable process mode, and the default behavior should not be mistaken for a release-order guarantee. Review the configured mode and choose deliberately. If a newer commit may supersede an older queued deploy, the deployment operation must be safe when an older job is skipped or when a newer version is applied first. Idempotence, commit eligibility checks, and a clear rollback path remain application and release-system responsibilities.

The environment declaration gives GitLab deployment metadata for the target, but the example’s when: manual rule is not a complete authorization model. Configure protected environments and appropriate approvals in the project or group, scope deployment credentials to the environment, and restrict which refs can reach the job. The YAML should make the intended branch condition visible; project permissions must enforce who can approve and execute the action.

Keep the lock for the entire operation that must not overlap. If the script launches a separate asynchronous deployment and exits immediately, GitLab can release the resource group while the external rollout is still in progress. Wait for the actual rollout result or use a deployment controller that owns the concurrency boundary. Conversely, avoid holding a production lock during builds and tests; those stages can run in parallel before the job acquires the production resource.

Validate the pipeline as a contract

Before enabling the pipeline on the production branch, test its configuration against representative event types: a branch push without a merge request, a branch push with an open merge request, a merge request event, a tag, a schedule, and any downstream trigger used by the repository. Confirm that each event creates exactly the intended pipeline. Pipeline configuration validation catches syntax and many graph errors, but it cannot prove that a deployment is safe or that a script checks the correct target.

Then test failure behavior, not just the happy path:

  • Make lint fail and verify the package cannot be published as a successful release.
  • Make the package command produce no output and verify the artifact step fails rather than publishing an empty package.
  • Corrupt the manifest commit and confirm deployment refuses to proceed.
  • Fail the integration or policy job and verify production remains unchanged.
  • Start two deployment pipelines and observe that their production mutations do not overlap.
  • Queue an older and a newer commit and confirm the selected resource-group process mode matches the intended release policy.
  • Interrupt a deployment and verify the next run can determine actual environment state before applying another change.
  • Expire or remove a temporary artifact in a test project and confirm the recovery path does not rebuild an unreviewed release silently.

For each production run, retain the pipeline ID, commit SHA, package or image digest, approval record, deployment result, and post-deploy health evidence in a place operators can retrieve. A green job status is one signal; it is not a complete audit trail or proof that the running environment matches the tested artifact.

Production review checklist

  • Are pipeline-creation rules explicit for pushes, merge requests, tags, schedules, and downstream events?
  • Does each needs edge represent a genuine dependency, and do deploy jobs wait for every required test and policy gate?
  • Are needs:optional edges truly optional under a documented policy?
  • Does every artifact have a narrow path, explicit consumer, sensible retention, and identity that can be verified?
  • Does the deployment consume the tested output instead of rebuilding from a mutable reference?
  • Do all writers to one environment use the same resource-group lock, and is queue order intentional?
  • Do environment permissions and credential scopes enforce the intended approval boundary?
  • Have failure, concurrency, interruption, and rollback behavior been exercised outside production?

GitLab CI becomes easier to trust when its graph makes the release contract visible: one intended pipeline per event, explicit producer-consumer edges, artifacts that carry verifiable identity, and a lock around the actual shared environment. Faster scheduling is valuable only when those boundaries remain correct.

Related:

Sources:

Comments