Airflow Summit 2026 is coming August 31 - September 2 in Austin, TX. Register now to secure your spot!

Metrics Configuration

Airflow can be set up to send metrics to StatsD or OpenTelemetry.

Setup - StatsD

To use StatsD you must first install the required packages:

pip install 'apache-airflow[statsd]'

then add the following lines to your configuration file e.g. airflow.cfg

[metrics]
statsd_on = True
statsd_host = localhost
statsd_port = 8125
statsd_prefix = airflow

If you want to use a custom StatsD client instead of the default one provided by Airflow, the following key must be added to the configuration file alongside the module path of your custom StatsD client. This module must be available on your PYTHONPATH.

[metrics]
statsd_custom_client_path = x.y.customclient

See Modules Management for details on how Python and Airflow manage modules.

Setup - OpenTelemetry

To use OpenTelemetry you must first install the required packages:

pip install 'apache-airflow[otel]'

An OpenTelemetry Collector (or compatible service) is required for connectivity to a metrics backend. Add the Collector details to your configuration file e.g. airflow.cfg

[metrics]
otel_on = True
otel_host = localhost
otel_port = 8889
otel_prefix = airflow
otel_interval_milliseconds = 30000  # The interval between exports, defaults to 60000
otel_service = Airflow
otel_ssl_active = False

Note

The following config keys have been deprecated and will be removed in the future

[metrics]
otel_host = localhost
otel_port = 8889
otel_interval_milliseconds = 30000
otel_debugging_on = False
otel_service = Airflow
otel_ssl_active = False

The OpenTelemetry SDK should be configured using standard OpenTelemetry environment variables such as OTEL_EXPORTER_OTLP_ENDPOINT, OTEL_EXPORTER_OTLP_PROTOCOL, etc.

See the OpenTelemetry exporter protocol specification and SDK environment variable documentation for more information.

Enable Https

To establish an HTTPS connection to the OpenTelemetry collector You need to configure the SSL certificate and key within the OpenTelemetry collector’s config.yml file.

receivers:
  otlp:
    protocols:
      http:
        endpoint: 0.0.0.0:4318
        tls:
          cert_file: "/path/to/cert/cert.crt"
          key_file: "/path/to/key/key.pem"

Histogram Metrics and Backend Requirements

Airflow’s timing metrics (timing() / timer()) are emitted as OpenTelemetry histograms aggregated with exponential bucket histograms, so bucket boundaries adapt automatically to the observed range and you do not have to hand-tune explicit buckets for metrics that span very different scales (milliseconds to hours).

To ingest these correctly end-to-end, the metrics backend you connect to must support OpenTelemetry exponential histograms and (for Prometheus) their conversion to native histograms:

  • OpenTelemetry Collector — use opentelemetry-collector-contrib version 0.115.0 or above. Older versions do not translate OTLP exponential histograms into Prometheus native histograms.

  • Prometheus — native histograms must be enabled explicitly, and how you do that depends on the Prometheus version:

    • 2.40 to 3.8 — start Prometheus with the --enable-feature=native-histograms flag.

    • 3.8 and above — set scrape_native_histograms: true in the scrape configuration (this option was added in 3.8, and from 3.9 the feature flag is a no-op so the config setting is required):

      global:
          scrape_native_histograms: true
      

If the backend does not support native histograms, exponential-histogram data points may be dropped or rendered incorrectly. A reference stack (Collector, Prometheus, and Grafana) wired up for local development is available via breeze start-airflow --integration otel; see the contributor docs for details.

Allow/Block Lists

If you want to avoid sending all the available metrics, you can configure an allow list or block list to send or block only certain metrics. Each list is a comma-separated set of regular expressions matched anywhere in the metric name (anchor a pattern with ^ to match a prefix). If both lists are set, the block list is ignored:

[metrics]
metrics_allow_list = scheduler,executor,dagrun,pool,triggerer,celery
[metrics]
metrics_block_list = scheduler,executor,dagrun,pool,triggerer,celery

Rename Metrics

If you want to redirect metrics to a different name, you can configure the stat_name_handler option in [metrics] section. It should point to a function that validates the stat name, applies changes to the stat name if necessary, and returns the transformed stat name. The function may look as follows:

def my_custom_stat_name_handler(stat_name: str) -> str:
    return stat_name.lower()[:32]

Custom Metrics

You can emit your own metrics from inside a task, plugin, or custom operator through the same stats client Airflow uses internally. In Airflow 3 the recommended import path is airflow.sdk.observability:

from airflow.sdk.observability import stats

stats.incr("my_service.processed")
stats.decr("my_service.in_flight")
stats.gauge("my_service.queue_depth", 42)
stats.timing("my_service.batch_ms", 1234)

with stats.timer("my_service.batch"):
    ...

Added in version 3.3.0: The module-level stats functions (stats.incr(), stats.gauge(), and so on).

On earlier versions, use the Stats class instead: from airflow.sdk.observability.stats import Stats, then Stats.incr(...).

incr, decr, gauge, timing and timer also accept an optional tags mapping for dimensional metrics on backends that support them:

stats.incr("my_service.requests", tags={"endpoint": "checkout"})

incr and decr also accept count and rate, and gauge accepts rate and delta, following the StatsD data types.

Note

Tag support depends on the backend. The classic StatsD protocol has no concept of tags.

  • OpenTelemetry (otel_on) sends tags as native attributes.

  • StatsD (statsd_on) drops the tags mapping by default. To turn tags into labels, enable a tagged wire format, either statsd_influxdb_enabled = True (InfluxDB name,key=value) or statsd_datadog_enabled = True (DogStatsD |#key:value). The Prometheus statsd_exporter reads the tags from either format and turns them into labels. These flags only change how tags are written on the wire. You can also embed the values in the metric name and map those name segments back to labels with statsd_exporter mapping rules.

Note

Metric names must be 250 characters or fewer and may only contain the characters a-z, A-Z, 0-9, _, ., - and /. An invalid name is logged and the metric is not emitted.

Note

These metrics are silently dropped unless a backend is enabled (see Setup - StatsD or Setup - OpenTelemetry).

Note

If your custom metrics do not appear, check [metrics] metrics_allow_list and [metrics] metrics_block_list (see Allow/Block Lists). When metrics_allow_list is set, only metrics matching it are emitted, so a custom metric that is not listed is silently dropped.

Other Configuration Options

Note

For a detailed listing of configuration options regarding metrics, see the configuration reference documentation - [metrics].

Metric Descriptions

Counters

Name

Legacy Name

Description

{job_name}_start

-

Number of started {job_name} job, ex. SchedulerJob, LocalTaskJob

{job_name}_end

-

Number of ended {job_name} job, ex. SchedulerJob, LocalTaskJob

{job_name}_heartbeat_failure

-

Number of failed Heartbeats for a {job_name} job, ex. SchedulerJob, LocalTaskJob

operator_failures

operator_failures_{operator_name}

Operator {operator_name} failures.

operator_successes

operator_successes_{operator_name}

Operator {operator_name} successes.

resumable_job.fresh_submit

-

Number of times a ResumableJobMixin operator submitted a fresh job with no prior run ID stored (first run). Metric with operator tagging.

resumable_job.already_succeeded

-

Number of times a ResumableJobMixin operator found a stored run ID whose job had already completed successfully and skipped resubmission. Metric with operator tagging.

resumable_job.terminal_resubmit

-

Number of times a ResumableJobMixin operator found a stored run ID whose job was in a terminal (failed) state and submitted a fresh job. Metric with operator tagging.

resumable_job.reconnect_attempt

-

Number of times a ResumableJobMixin operator found a stored run ID on retry and attempted to reconnect. Metric with operator tagging.

resumable_job.reconnect_success

-

Number of times a ResumableJobMixin operator successfully reconnected to an active job on retry. Metric with operator tagging.

ti_failures

-

Overall task instances failures. Metric with dag_id and task_id tagging.

ti_successes

-

Overall task instances successes. Metric with dag_id and task_id tagging.

previously_succeeded

-

Number of previously succeeded task instances. Metric with dag_id and task_id tagging.

task_instances_without_heartbeats_killed

-

Task instances without heartbeats killed. Metric with dag_id and task_id tagging.

scheduler_heartbeat

-

Scheduler heartbeats

dag_processor_heartbeat

-

Standalone Dag processor heartbeats

dag_processing.processes

-

Relative number of currently running Dag parsing processes (ie this delta is negative when, since the last metric was sent, processes have completed). Metric with file_path and action tagging.

dag_processing.processor_timeouts

-

Number of file processors that have been killed due to taking too long. Metric with file_path tagging.

dag_processing.other_callback_count

-

Number of non-SLA callbacks received

dag_processing.callback_only_count

-

Number of DAG file processing runs that processed callbacks only, without full DAG parsing

dag_processing.file_path_queue_update_count

-

Number of times we’ve scanned the filesystem and queued all existing Dags

scheduler.tasks.killed_externally

-

Number of tasks killed externally. Metric with dag_id and task_id tagging.

scheduler.orphaned_tasks.cleared

-

Number of Orphaned tasks cleared by the Scheduler

scheduler.orphaned_tasks.adopted

-

Number of Orphaned tasks adopted by the Scheduler

scheduler.loop_exceptions

-

Number of times the scheduler loop exited with an unhandled exception. Metric with exception_class tagging.

scheduler.executor_events.processed

-

Number of executor events processed per process_executor_events call.

scheduler.executor_events.failed

-

Number of times process_executor_events raised an exception. Metric with exception_class tagging.

scheduler.zombies.detected

-

Number of zombie task instances detected by the Scheduler. Metric with reason tagging (heartbeat_timeout).

scheduler.critical_section_busy

-

Count of times a scheduler process tried to get a lock on the critical section (needed to send tasks to the executor) and found it locked by another process.

ti.start

ti.start.{dag_id}.{task_id}

Number of started task in a given Dag. Similar to {job_name}_start but for task. Metric with dag_id and task_id tagging.

ti.finish

ti.finish.{dag_id}.{task_id}.{state}

Number of completed task in a given Dag. Similar to {job_name}_end but for task. Metric with dag_id and task_id tagging.

dag.callback_exceptions

-

Number of exceptions raised from Dag callbacks. When this happens, it means Dag callback is not working. Metric with dag_id tagging

dag.serialization_writes

dag.serialization_writes.{dag_id}.{bundle_name}

Number of times a Dag was serialized and written to the metadata DB. Metric with dag_id and bundle_name tagging.

celery.task_timeout_error

-

Number of AirflowTaskTimeout errors raised when publishing Task to Celery Broker.

celery.execute_command.failure

-

Number of non-zero exit code from Celery task.

task_removed_from_dag

task_removed_from_dag.{dag_id}

Number of tasks removed for a given Dag (i.e. task no longer exists in Dag). Metric with dag_id and run_type tagging.

task_restored_to_dag

task_restored_to_dag.{dag_id}

Number of tasks restored for a given Dag (i.e. task instance which was previously in REMOVED state in the DB is added to Dag file). Metric with dag_id and run_type tagging.

task_instance_created

task_instance_created_{task_type}

Number of tasks instances created for a given Operator. Metric with dag_id and run_type tagging.

triggerer_heartbeat

-

Triggerer heartbeats

triggers.blocked_main_thread

-

Number of triggers that blocked the main thread (likely due to not being fully asynchronous)

triggers.failed

-

Number of triggers that errored before they could fire an event

triggers.succeeded

-

Number of triggers that have fired at least one event

asset.updates

-

Number of updated assets

asset.triggered_dagruns

-

Number of Dag runs triggered by an asset update

deadline_alerts.deadline_created

-

Number of deadline alerts created for a Dag run

deadline_alerts.deadline_missed

-

Number of deadline alerts that fired because a Dag run missed its deadline

deadline_alerts.deadline_not_missed

-

Number of deadline records deleted because the Dag run finished before the deadline

ol.emit.failed

-

Number of failed OpenLineage event emit attempts

api_server.dag_bag.cache_hit

-

Number of cache hits when retrieving SerializedDAG from DBDagBag in the API server

api_server.dag_bag.cache_miss

-

Number of cache misses when retrieving SerializedDAG from DBDagBag in the API server

api_server.dag_bag.cache_clear

-

Number of times the DBDagBag cache was cleared in the API server

connection_test.success

-

Number of worker-dispatched connection tests that completed successfully.

connection_test.failed

-

Number of worker-dispatched connection tests that completed with a failure.

connection_test.reaped

-

Number of stale connection tests marked failed by the scheduler reaper. Metric with prior_state tagging.

edge_worker.heartbeat_count

edge_worker.heartbeat_count.{worker_name}

Number heartbeats in an edge worker.

edge_worker.ti.start

edge_worker.ti.start.{queue}.{dag_id}.{task_id}

Number of task instances started on an edge worker.

edge_worker.ti.finish

edge_worker.ti.finish.{queue}.{state}.{dag_id}.{task_id}

Number of task instances finished on an edge worker.

kubernetes_executor.pod_creation_status

-

Number of Kubernetes create_namespaced_pod calls from the Kubernetes Executor, tagged by HTTP response status (200 on success, the ApiException status code on failure, error for non-API exceptions).

kubernetes_executor.pod_deletion_status

-

Number of Kubernetes delete_namespaced_pod calls from the Kubernetes Executor, tagged by HTTP response status (200 on success, the ApiException status code on failure).

kubernetes_executor.pod_patching_status

-

Number of Kubernetes patch_namespaced_pod calls from the Kubernetes Executor, tagged by HTTP response status (200 on success, the ApiException status code on failure).

Gauges

Name

Legacy Name

Description

asset.orphaned

-

Number of assets marked as orphans because they are no longer referenced in Dag schedule parameters or task outlets

dagbag_size

-

Number of Dags found when the scheduler ran a scan based on its configuration

api_server.dag_bag.cache_size

-

Current number of SerializedDAG objects cached in the API server’s DBDagBag

connection_test.active

-

Number of connection tests currently in flight (queued + running), sampled by the scheduler each tick.

connection_test.pending

-

Number of connection tests waiting in the pending backlog (queue depth), sampled by the scheduler each tick.

dag_processing.import_errors

-

Number of errors from trying to parse Dag files

dag_processing.total_parse_time

-

Seconds taken to scan and import dag_processing.file_path_queue_size Dag files

dag_processing.file_path_queue_size

-

Number of Dag files to be considered for the next scan

dag_processing.last_run.seconds_ago

dag_processing.last_run.seconds_ago.{file_name}

Seconds since a DAG file was last processed. Metric with file_path, bundle_name and file_name tagging.

scheduler.tasks.starving

-

Number of tasks that cannot be scheduled because of no open slot in pool

scheduler.executor_events.batch_size

-

Number of executor events in the batch processed per process_executor_events call.

scheduler.tasks.executable

-

Number of tasks that are ready for execution (set to queued) with respect to pool limits, Dag concurrency, executor state, and priority.

scheduler.dagruns.running

-

Number of DAGs whose latest DagRun is currently in the RUNNING state

executor.open_slots

executor.open_slots.{executor_class_name}

Number of open slots on executor. Legacy metric only emitted when multiple executors are configured.

executor.queued_tasks

executor.queued_tasks.{executor_class_name}

Number of queued tasks on executor. Legacy metric only emitted when multiple executors are configured.

executor.running_tasks

executor.running_tasks.{executor_class_name}

Number of running tasks on executor. Legacy metric only emitted when multiple executors are configured.

pool.open_slots

pool.open_slots.{pool_name}

Number of open slots in the pool.

pool.queued_slots

pool.queued_slots.{pool_name}

Number of queued slots in the pool.

pool.running_slots

pool.running_slots.{pool_name}

Number of running slots in the pool.

pool.deferred_slots

pool.deferred_slots.{pool_name}

Number of deferred slots in the pool.

pool.scheduled_slots

pool.scheduled_slots.{pool_name}

Number of scheduled slots in the pool.

pool.starving_tasks

pool.starving_tasks.{pool_name}

Number of starving tasks in the pool.

triggers.running

triggers.running.{hostname}

Number of triggers currently running for a triggerer (described by hostname).

triggerer.capacity_left

triggerer.capacity_left.{hostname}

Capacity left on a triggerer to run triggers (described by hostname).

ti.scheduled

ti.scheduled.{queue}.{dag_id}.{task_id}

Number of scheduled tasks in a given Dag.

ti.queued

ti.queued.{queue}.{dag_id}.{task_id}

Number of queued tasks in a given Dag.

ti.running

ti.running.{queue}.{dag_id}.{task_id}

Number of running tasks in a given Dag. As ti.start and ti.finish can run out of sync this metric shows all running tis.

ti.deferred

ti.deferred.{queue}.{dag_id}.{task_id}

Number of deferred tasks in a given Dag.

ol.event.size

ol.event.size.{event_type}.{operator_name}

Size in bytes of an OpenLineage event by event type and operator.

edge_worker.status

edge_worker.status.{worker_name}

Edge worker status (expressed as Python logging level).

edge_worker.connected

edge_worker.connected.{worker_name}

Edge worker in state connected.

edge_worker.maintenance

edge_worker.maintenance.{worker_name}

Edge worker in state maintenance.

edge_worker.jobs_active

edge_worker.jobs_active.{worker_name}

Number of active jobs in an edge worker.

edge_worker.concurrency

edge_worker.concurrency.{worker_name}

Concurrency capacity in an edge worker.

edge_worker.free_concurrency

edge_worker.free_concurrency.{worker_name}

Available concurrency in an edge worker.

edge_worker.num_queues

edge_worker.num_queues.{worker_name}

Number of queues in an edge worker.

kubernetes_executor.pod_creation_batch_size

-

Number of worker pods created in one Kubernetes Executor scheduler loop.

Timers

Name

Legacy Name

Description

dagrun.dependency-check

dagrun.dependency-check.{dag_id}

Milliseconds taken to check Dag dependencies

task.duration

dag.{dag_id}.{task_id}.duration

Milliseconds taken to run a task

task.scheduled_duration

dag.{dag_id}.{task_id}.scheduled_duration

Milliseconds a task spends in the Scheduled state, before being Queued

task.queued_duration

dag.{dag_id}.{task_id}.queued_duration

Milliseconds a task spends in the Queued state, before being Running

dag_processing.last_duration

dag_processing.last_duration.{bundle_name}.{file_name}

Milliseconds taken to load the given Dag file

dagrun.duration.success

dagrun.duration.success.{dag_id}

Milliseconds taken for a DagRun to reach success state

dagrun.duration.failed

dagrun.duration.failed.{dag_id}

Milliseconds taken for a DagRun to reach failed state

dagrun.schedule_delay

dagrun.schedule_delay.{dag_id}

Milliseconds of delay between the scheduled DagRun start date and the actual DagRun start date

triggerer.batch_trigger_creation_duration

-

Time in milliseconds spent creating a batch of pending triggers.

scheduler.critical_section_duration

-

Milliseconds spent in the critical section of scheduler loop

scheduler.critical_section_query_duration

-

Milliseconds spent running the critical section task instance query

scheduler.scheduler_loop_duration

-

Milliseconds spent running one scheduler loop

scheduler.executor_heartbeat_duration

-

Milliseconds spent in executor.heartbeat() per scheduler loop iteration, tagged by executor class name so each configured executor is reported separately.

triggerer.trigger_queue_delay

-

Time in milliseconds between a trigger workload being queued and being processed by the TriggerRunner.

dagrun.first_task_scheduling_delay

dagrun.{dag_id}.first_task_scheduling_delay

Milliseconds elapsed between first task start_date and dagrun expected start

dagrun.first_task_start_delay

-

Milliseconds elapsed between dagrun queued_at and first task start_date

kubernetes_executor.adopt_task_instances.duration

-

Milliseconds taken to adopt the task instances in Kubernetes Executor

kubernetes_executor.pod_creation

-

Milliseconds taken for a Kubernetes create_namespaced_pod call from the Kubernetes Executor

kubernetes_executor.pod_creation_batch_duration

-

Milliseconds taken to create one batch of worker pods in a Kubernetes Executor scheduler loop, covering both the sequential and the concurrent (async_pod_creation) creation paths.

kubernetes_executor.pod_deletion

-

Milliseconds taken for a Kubernetes delete_namespaced_pod call from the Kubernetes Executor

kubernetes_executor.pod_patching

-

Milliseconds taken for a Kubernetes patch_namespaced_pod call from the Kubernetes Executor

batch_executor.adopt_task_instances.duration

-

Milliseconds taken to adopt the task instances in the AWS Batch Executor

ecs_executor.adopt_task_instances.duration

-

Milliseconds taken to adopt the task instances in the AWS ECS Executor

lambda_executor.adopt_task_instances.duration

-

Milliseconds taken to adopt the task instances in the AWS Lambda Executor

edge_executor.sync.duration

-

Milliseconds taken for one sync heartbeat of the Edge Executor

ol.emit.attempts

ol.emit.attempts.{event_type}.{transport_type}

Milliseconds taken by an attempt to emit an OpenLineage event.

ol.extract

ol.extract.{event_type}.{operator_name}

Milliseconds taken to extract an OpenLineage event by event type and operator.

airflow.io.load_filesystems

-

Milliseconds taken to load filesystem implementations from providers

serde.load_serializers

-

Milliseconds taken to load all serializer modules

connection_test.dispatch_duration

-

Milliseconds the scheduler spends dispatching pending connection tests to executors in a single tick.

connection_test.hook_duration

-

Milliseconds a worker spends running the hook’s test_connection for a worker-dispatched connection test.

Was this entry helpful?