Skip to content
SRE & DevOpsDeep Dive Published Updated 8 min readViews unavailable

Terraform Data Sources: Plan-Time Reads, Deferred Queries, and Unknown Values

Understand when Terraform reads data sources, why queries move to apply, and how unknown results affect plans, dependencies, tests, and reproducibility.

Terraform data sources let a configuration query information from a provider without asking Terraform to create, update, or destroy that remote object. The result can feed resource arguments, outputs, or other expressions. Because a data source is read-only from Terraform’s perspective, it can look harmless. Its timing still matters: a read can happen during planning or be deferred until apply, and that difference changes what the plan can show.

A query whose arguments are known can usually run during planning, after Terraform’s refresh work, so the returned attributes can participate in the plan. If an argument depends on a value that will not be known until apply, Terraform cannot issue the query yet. It marks the data-source result unknown and defers the read. Resources that use that result may then have unknown arguments, delayed planning decisions, or more conservative replacement behavior.

A plan-known lookup

A data source is associated with a provider-defined query and returns provider-defined attributes. For example, a configuration can find an approved machine image from a known owner and name pattern, then pass its identifier to an instance resource:

data "aws_ami" "application" {
  most_recent = true
  owners      = [var.image_owner]

  filter {
    name   = "name"
    values = [var.image_name_pattern]
  }
}

resource "aws_instance" "application" {
  ami           = data.aws_ami.application.id
  instance_type = var.instance_type
}

If the owner and pattern are known before planning and the provider can perform the lookup, Terraform can read the data source during plan and display the selected identifier. The reference from the instance to the data source also gives Terraform a dependency relationship. The instance consumes the result; there is no need to add a separate depends_on edge just to restate that reference.

The query should still be constrained enough to return the intended object. A lookup such as most recent image is a moving selection: a later plan may legitimately choose a different image even when the Terraform source has not changed. Add owner, status, naming, architecture, or release-channel filters appropriate to the provider and application. If deterministic deployments are required, use a reviewed, fixed identifier or a controlled release manifest instead of relying on a broad moving query.

Why Terraform may defer a read

Terraform may defer a data-source read to apply when at least one query argument depends on a resource attribute or other value that cannot be predicted during planning. Common causes include a managed resource changing in the same plan, a computed value flowing into the query, or a dependency that forces the read to wait. The plan marks the data-source result as known after apply because the provider has not returned it yet.

This can be correct and unavoidable. If a new network identifier is created in the same operation and a provider query requires that identifier, Terraform cannot query until creation has produced it. The downstream resource may then remain partially unknown during plan. Do not interpret an unknown value as a failed lookup or as proof that Terraform will make no changes; it means the plan lacks that value until a later phase.

A direct data-source reference to a resource Terraform already manages is usually unnecessary. If a resource argument needs an attribute, reference the managed resource directly. Reading the object back through a provider API adds a query and can defer planning without adding useful information. Use data sources for external objects or provider-specific lookup operations, not as an indirect way to address something already in the dependency graph.

Explicit depends_on changes read timing

An explicit depends_on on a data source is sometimes required when the query depends on a hidden side effect that Terraform cannot infer from its arguments. It tells Terraform to complete all actions on the dependency first, including reads. Because it adds an ordering constraint, a data-source read that could otherwise happen during planning may be postponed until apply.

Before adding the edge, ask whether an input can reference the value it actually needs. A normal expression reference communicates the precise data flow and generally allows Terraform to make a more accurate plan. An explicit dependency says that the query relies on the whole dependency object’s actions, which can be broader than the particular value needed.

The same caution applies to module-level dependencies. If a data source is inside a module and the entire module depends on another module, Terraform may defer more work than intended. Make dependencies as narrow as possible, document why they are hidden, and inspect plan output for values that changed from known to unknown.

Plan determinism and external change

A data source can return a different result on a later plan because the external system changed. A new image may have been published, a DNS record may have moved, an inventory entry may have changed, or a service may have rotated an identifier. Terraform configuration can remain identical while the provider query returns new data.

This is expected for a live query, but it has operational consequences. The plan should show the data source result and downstream actions so a reviewer can decide whether the change is intended. If the query is deferred until apply, the final value may not be available in the original plan in the same way as a plan-time lookup. Choose identifiers and release inputs that support the level of repeatability your environment requires.

Saved plans help preserve the reviewed proposal produced by a planning operation, but they do not turn a mutable lookup into an immutable source of truth. A saved plan may include the data known at plan time; values that are deferred still have apply-time behavior. Treat a stale or regenerated plan as a new change proposal and re-review it when external inputs or configuration have changed.

Avoid depending on incidental ordering between unrelated data sources and resources. If a query returns a “latest” object, pin or record the chosen identifier when the deployment requires the same object across environments. If the query is intended to follow current external state, make that behavior explicit and accept that the plan can change without a source-code edit.

Diagnose unknown data-source values

When a plan says a data source will be read during apply, inspect every argument and custom condition on that data block. Identify references to computed resource attributes, objects with planned changes, or explicit dependencies. Then trace the unknown through resources that consume the data. The goal is to understand whether the query genuinely needs an apply-time value or whether a broad dependency or unnecessary lookup caused deferral.

Do not remove a necessary dependency merely to make the plan look more concrete. If the query truly requires the result of a resource action, the deferred read reflects the actual dependency. Instead, consider whether the resource can consume a known input directly, whether the external data should be produced in a separate stage, or whether the architecture needs a distinct deployment boundary.

When the read fails at apply, separate provider authentication and API errors from Terraform graph timing. Preserve the provider diagnostic, query arguments, plan, and state version. Do not expose sensitive query arguments in logs. Retry only when the provider operation is safe and the failure is transient; changing a filter or target to get a passing result can select the wrong infrastructure object.

Testing data-source-dependent modules

Unit-style tests can use a mock provider to supply deterministic values for data sources and assert how the module uses them. This is useful for checking name construction, filtering inputs, conditional resources, and outputs without requiring a cloud account. Keep the mock values realistic enough for the expressions and conditions being tested.

A mocked query cannot establish that the real provider API returns a result, that the caller has permission, or that a region contains the expected object. Use a controlled integration plan when provider-specific behavior matters. Keep the query constrained to a dedicated test account or known test object, and avoid integration tests that silently select a current production resource.

Test both paths when a data result is optional: the case where a valid result exists and the case where the module should reject or omit a dependent resource. If the query is intentionally deferred, inspect the planned unknowns and validate the apply behavior in an isolated environment. A test should explain whether it is verifying Terraform expression logic or a real remote lookup.

Data sources are not synchronization primitives

A successful read says that the provider returned data at a point in the Terraform operation. It does not guarantee that a service is healthy, that a configuration has propagated to every replica, or that a subsequent write will observe the same value. A data source is not a retry policy, a readiness probe, or a transaction.

If an application must wait for a service to become ready, use a health check and rollout mechanism that observes readiness. If a remote system has eventual consistency, use a bounded retry in the correct orchestration layer or a provider-supported waiter. Adding depends_on may order Terraform operations, but it does not make a read poll until an external condition is true unless the provider’s data source explicitly implements that behavior.

Keep provider queries read-only in design and narrow in scope. A data source should report remote information, not become an obscure side effect. If the same system needs to be managed, declare a resource with Terraform ownership or coordinate the external owner explicitly so two controllers do not fight over the same object.

Data sources enrich plans with external facts, but their value depends on when those facts are available and how stable they are. Constrain queries, prefer direct references for Terraform-managed values, keep dependencies narrow, and review unknowns rather than treating them as blanks. That makes a data-source-based plan honest about what Terraform knows before apply.

Related:

Sources:

Comments