Skip to content
WSLDeep Dive Published Updated 8 min readViews unavailable

Apache Airflow in WSL: A Local DAG and Scheduler Lab

Run Airflow in WSL for DAG development, inspect scheduling and backfills, and keep SQLite, Linux paths, and workstation uptime limits explicit.

Apache Airflow is useful in WSL when a developer needs to author and debug DAGs against a Linux Python environment without standing up a shared orchestration platform. The right scope is a disposable local scheduler and web interface for development, not a promise that scheduled workloads will run continuously. Windows shutdown or an explicit distribution termination stops the guest; during sleep, the host is not continuously executing ordinary scheduled work. A systemd service cannot override that host and VM lifecycle.

Airflow models workflows as DAGs, tasks, and dependencies. For time-based schedules, a DAG run is associated with a logical date and data interval, while the scheduler decides which eligible task instances can run; event-driven DAGs have different triggering semantics. Local tests can validate imports, dependency wiring, templating, task code, and a bounded backfill. They cannot validate distributed executor behavior, production secrets integration, independent metadata-database recovery, or continuous scheduling through host outages.

Keep the Python environment and project on the Linux side

Install and run Airflow inside the WSL distribution, not through a Windows Python interpreter that happens to be reachable from a Linux shell. Put the project, virtual environment, DAG folder, logs, and local metadata database on the distro filesystem, for example below $HOME. Microsoft recommends keeping Linux command-line workloads in the Linux filesystem for performance and more predictable Linux file semantics. Cross the /mnt/c boundary only when intentionally testing Windows file access.

Start with the Airflow installation page for the selected release. Airflow has several components and dependencies, so use its documented local-start mechanism and version constraints rather than copying a random pip install command. The current installation guide documents pipx run apache-airflow standalone and the analogous uvx form as a development/testing quick start. It starts a minimal local environment with an automatically generated administrator password and SQLite metadata database. Store the generated password safely; do not paste it into a committed file or shared terminal transcript.

mkdir -p "$HOME/airflow-lab/dags" "$HOME/airflow-lab/logs"
cd "$HOME/airflow-lab"
export AIRFLOW_HOME="$HOME/airflow-lab"
pwd
df -T .
pipx run apache-airflow standalone

The one-command start is intentionally convenient, not a production recipe. SQLite is documented for experimentation and development, not a production Airflow metadata database. For repeatable work, pin the chosen Airflow release and Python version in a project-specific environment, record provider packages, and use the installation guide’s constraints file for dependency resolution. Do not install a new Airflow version into a general-purpose Python environment that also serves unrelated projects.

Understand what the local scheduler proves

The scheduler parses DAG files and creates task instances when a DAG run’s dependencies and timing allow it. A DAG’s schedule is not simply a wall-clock timer for arbitrary code: the logical date and data interval matter. A daily DAG typically processes a completed interval after that interval ends, rather than processing data “for today” at midnight. When a task is retried, the task code may execute more than once, so external side effects should be idempotent or guarded by an application-level deduplication key.

Before writing a DAG, state the input interval, output partition, retry behavior, and what a rerun is allowed to overwrite. A local run that writes to a temporary table or a uniquely named file is safer than a test that silently mutates a shared database. Use a short, explicit DAG start date and a small time window; do not create hundreds of historic runs by giving a test DAG an old start date and an unrestricted catchup policy.

Backfills are particularly useful for validating time partition logic. Airflow’s current backfill interface supports a date range and reprocessing behavior, and exposes a dry-run option. First inspect which logical dates the range selects. Then constrain concurrency and use test-only data. Reprocessing “failed” runs and reprocessing “completed” runs have different consequences; understand the mode before launching it. A successful local backfill proves neither that a production executor has capacity nor that a downstream system can safely absorb duplicate writes.

For the current Airflow 3.3 CLI, a preview has this shape; replace the DAG ID and dates with a small test interval and confirm the options with airflow backfill create --help for the installed release:

airflow backfill create --dag-id example_daily --from-date 2026-09-01 --to-date 2026-09-03 --dry-run

Design a safe local acceptance exercise

Use a simple DAG with one deterministic task that emits its logical interval and a downstream task that validates the expected partition key. Keep the test output below the lab directory. Check that the DAG appears in the UI, that the scheduler creates the intended run, that dependencies execute in order, and that a deliberate task failure is marked accurately. Then retry it and verify that the second attempt has the expected effect. Avoid embedding credentials in DAG source; local developer credentials are still credentials.

Use Airflow CLI help from the exact installed version before relying on commands in scripts. Inspect the DAG list and task states, then read the scheduler and task logs rather than treating a green UI badge as sufficient. If using the CLI to test a backfill, perform a dry run first, select a narrow interval, and set a small maximum number of active runs. Clean up only the test DAG’s outputs and only after confirming their location.

Also test the workflow under modest concurrency. A laptop can run out of memory or file descriptors when several Python tasks start together, even if each task succeeds alone. Keep task-level parallelism low for the first pass, and add concurrency only when the application requirement calls for it. Airflow pools and per-DAG run limits can shape scheduling, but a local exercise should record the configured executor and limits rather than assume they match another environment. Measure task start delay and duration separately from the time spent parsing the DAG.

Database and service lifecycle boundaries

The standalone SQLite metadata file stores Airflow’s own operational state, not your task’s business data. Treat it as disposable unless you have a documented backup. Do not copy it while the process is writing and call the copy a consistent backup. If the test requires a PostgreSQL backend, use an isolated development database and follow the Airflow guide for schema migration; stop Airflow components before manually running database migrations. Never point a local experiment at a production metadata database.

You can enable systemd per distribution and run services while that distro is alive, but this does not guarantee that WSL remains running. A Windows Task Scheduler entry can start a distribution for a host-controlled workflow, but that is a different orchestration boundary and must be tested under the relevant Windows account and logon conditions. Do not confuse “systemd says active” with “the machine will be awake at the scheduled time.”

Keep listeners private. The web UI is an administrative interface capable of launching code through tasks. Do not assume the standalone quick start is loopback-only: Airflow’s API server host setting defaults to 0.0.0.0. Inspect the effective setting with airflow config get-value api host, and bind or firewall it deliberately before using the lab on a shared or untrusted network. If you change the bind address, verify Windows-browser access under the active WSL networking mode rather than widening the listener as a proxy workaround. NAT localhost forwarding and mirrored networking have different rules.

Troubleshoot by layer

If the DAG is missing, inspect the configured DAG folder, Python import errors, file permissions, and scheduler logs. If a task stays queued, inspect executor configuration and available local process capacity. If imports work in an interactive shell but fail in a task, compare the executable, Python environment, working directory, and environment variables used by the task process. If a backfill produces unexpected dates, review the DAG timetable and data interval semantics before changing the system clock.

When disk use grows, separate task logs, DAG outputs, and the metadata database. Set retention for disposable artifacts and keep enough logs to diagnose a failed run. Avoid deleting Airflow’s metadata database to fix a scheduler symptom until you understand what state the deletion removes.

Acceptance criteria

A WSL Airflow lab is ready when the installed version and Python environment are recorded, project state is on the Linux filesystem, the local UI and scheduler start, one deterministic DAG runs with its expected interval and dependency order, a controlled failure/retry behaves as designed, and a dry-run backfill selects only the intended dates. The team should also know what happens when WSL stops and which state is intentionally disposable.

This is a productive place to develop DAG code and learn scheduler semantics. It is not a production Airflow architecture, durable scheduling service, or test of worker/metadata-database high availability.

Related:

Sources:

Comments