Lesson 8 · Scheduling & execution

Task lifecycle, retries & trigger rules

How a task instance moves through states — and decides whether to run at all.

Your win: name the task-instance states, explain retries, and use trigger rules to control when a task runs given its upstreams — the last piece of the execution model.

The task-instance lifecycle

none → scheduled → queued → running → success ↘ failed → up_for_retry → (running again) upstream failed → upstream_failed branch not taken → skipped

A task instance (Lesson 2) is what carries state. The scheduler moves it scheduled → queued; the executor runs it (running); it ends success or failed.1

Retries

Set retries and retry_delay (usually in default_args). A failed task with retries left goes up_for_retry, waits, then runs again — automatic resilience against transient failures. Alerts fire via on_failure_callback.

Anchor Our templates pull retries and max_active_runs from rules.yaml into default_args (common-spark-job.py:22-24), set a dagrun_timeout of 90 minutes, and wire on_failure_callback=[slack_alert({{failed_mentions}})] — which is why failed_mentions is a required field (the generator refuses without it).

Trigger rules: when does a task run?

A task's trigger rule decides whether it runs based on its upstream tasks' states.2

RuleRuns when…
all_success (default)all upstreams succeeded (a skipped upstream cascades → skip)
all_doneall upstreams finished, whatever their outcome (great for cleanup / notify)
one_failedat least one upstream failed (e.g. an error handler)
none_failed_min_one_successno failures and at least one success (joins after branching)
The classic gotcha A "join" task after a branch shows up skipped under the default all_success, because a skipped upstream cascades. Use none_failed_min_one_success (or all_done) so the join still runs. This exact scenario is a favourite interview question.
Backend use case Backend use case: understanding task-instance state is what makes Slack alerts, retries, timeouts, and stuck-run debugging make sense in this repo’s generated DAGs.
Common mistake Common mistake: explaining task states from memory without tying them back to the scheduler/executor behavior that causes those states.
Read this next

Airflow — Tasks (states & trigger rules) + DAG Runs

The state diagram, retries, and the full trigger-rule list with examples.

core-concepts — Tasks
core-concepts — Trigger Rules

In plain English Plain English: a task instance moves through states as Airflow decides it is ready, queues it, runs it, retries it, or marks it done.

Check yourself (from memory)

Q1. The default trigger rule is…

all_success — and a skipped upstream cascades, which is the join-task gotcha.

Q2. A task that failed and will retry is in state…

up_for_retry while it waits retry_delay, then runs again. upstream_failed is when a parent failed.

Q3. trigger_rule="all_done" makes a task run…

Once all upstreams finish, regardless of outcome — ideal for cleanup or "always notify" tasks.
Sketch the task-instance lifecycle, and name the key trigger rules.
recall, then click to reveal
A task instance flows: none → scheduled → queued → running → success | failed. A failed task with retries left goes up_for_retry and re-runs (retries + retry_delay in default_args); a failed parent makes it upstream_failed; an untaken branch is skipped. TRIGGER RULES decide when a task runs given upstream states: default all_success (all ok — and skipped cascades); all_done (any outcome — cleanup/notify); one_failed (error handler); none_failed_min_one_success (joins after branching). Alerts via on_failure_callback (our slack_alert(failed_mentions)).
🎓 That's Part 2 — scheduling & execution You can reason about data intervals (when runs fire), catchup vs backfill, the executors, and the task lifecycle + trigger rules — the interview-critical core. Part 3 gets concrete: this repo's Airflow — how a rule.yaml becomes a DAG that runs a Spark job.
Ready for Part 3 (this repo's Airflow), or a mixed quiz across Lessons 5–8 first? Ask me.

1. Airflow — Tasks.

2. Airflow — Trigger Rules.