Managing queues
flyteplugins-union pluginThe queue CLI commands and Python objects on this page are provided by the
flyteplugins-union package. Install it with pip install flyteplugins-union.
A queue is a named scheduling lane. It does two jobs at once: it routes work to a cluster pool (and, optionally, specific clusters within it), and it governs that work with concurrency, depth, priority, and fairness limits.
This page covers creating and managing queues administratively, from either the CLI or Python. For how workflow authors target a queue from task code, see Queues in Configure tasks.
How a queue routes
A queue lives inside one cluster pool and routes work to one or more clusters
within that pool. By default (the * selector) it spreads across the pool’s
healthy, active clusters; you can also pin it to specific clusters. It can never
reach a cluster in another pool: pools are isolation boundaries.
flowchart TD
R(["Runs & actions"])
subgraph Pdef["Cluster pool: default"]
direction TB
QD["Queue: default<br/>selector: *"]
CA["Cluster A"]
CB["Cluster B"]
QD --> CA
QD --> CB
end
subgraph Pprod["Cluster pool: prod"]
direction TB
QP["Queue: prod-queue<br/>selector: *"]
QG["Queue: gpu-queue<br/>pinned: Cluster C"]
CC["Cluster C"]
CD["Cluster D"]
QP --> CC
QP --> CD
QG --> CC
end
R --> QD
R --> QP
R --> QG
Users submit to a queue, never to a pool or a cluster directly. Each queue sits inside exactly one pool:
defaultspreads across the eligible clusters in thedefaultpool.prod-queuespreads across the eligible clusters in theprodpool.gpu-queuelives in the sameprodpool but is pinned to a single cluster.
A * selector does not mean every cluster in the pool — it means every
cluster that is both healthy and active in its
lifecycle, evaluated against the pool’s
current state. Pinned selectors are filtered the same way: an unhealthy
cluster, or one that is draining, drained, or being deleted, receives no
new work from any queue, wildcard or pinned.
Say a pool holds two clusters and the second goes unhealthy for any reason,
including a
config mismatch with the pool:
the queue routes new runs and actions to the first cluster only, and sends
nothing to the second until it becomes healthy again. This governs the placement
of new work; it does not move work that has already been dispatched. Check with
flyte get cluster <name>, which reports each cluster’s lifecycle state,
health, and unhealthy reasons.
Queues you get for free
You don’t have to create a queue to have one. Two exist without any action on your part:
- The org-wide
defaultqueue, in thedefaultpool with the*selector. Anything that doesn’t explicitly target a queue goes here. (If thedefaultqueue is drained or deleted, untargeted submissions are rejected until it is active again — a restored queue comes backdrained, so after an undelete it must also be reactivated.) - A co-named queue for every cluster: registering a cluster creates a queue
with the same name as the cluster, in that cluster’s pool, whose selector
names that one cluster explicitly rather than using
*. Registerprod-us-east-1and you get aprod-us-east-1queue that routes only toprod-us-east-1— so any cluster can be targeted by name immediately, without setting up a queue for it. See The co-named queue.
Both are ordinary queues: they show up in flyte get queue and take the same
settings and updates as queues you create yourself — with one exception. A
co-named queue’s cluster selector and pool are managed by its cluster and cannot
be edited directly: the queue follows its cluster if the cluster is
reassigned to another pool. It
also follows the cluster through drain, activate, and delete transitions, and
it cannot be activated on its own while its cluster is draining or drained
(see
How the co-named queue follows its cluster).
The cluster and queue finish their transitions separately, so they may reach
their final states at slightly different times. Its other settings (concurrency,
depth, priority, fairness) stay editable like any queue’s. Listings make the
distinction visible: cluster-managed queues are flagged in the flyte get queue
table (cluster_managed, exposed as Queue.cluster_managed in Python), and
flyte update queue --edit says so at the top of the edit buffer.
The selector (which clusters within the pool) is mutable and can be changed at
any time. The pool a queue lives in can only change once the queue is fully
drained: a pool move on an active or draining queue is rejected, because
moving a queue that still holds work would cross an isolation boundary. See
Move work to another pool.
Every queue is visible to the whole organization; a queue cannot yet be scoped
to a project or a domain. Some CLI and Python surfaces already expose project
and domain parameters, but project/domain-scoped queue creation is not
implemented yet and is rejected. Support is coming soon.
Create a queue
run_concurrency and action_concurrency are required; everything else has a
sensible default. With no cluster selector, a queue spreads work across all
healthy clusters in its pool. The name default is reserved, and a queue cannot
share a name with a cluster — every cluster already owns its
co-named queue — or with a soft-deleted queue,
whose name stays reserved until it is undeleted.
flyte create queue my-queue \
--run-concurrency 100 \
--action-concurrency 1000Create a higher-priority queue in a specific pool:
flyte create queue gpu-queue \
--cluster-pool prod \
--cluster prod-us-east-1 \
--run-concurrency 50 \
--action-concurrency 500 \
--depth 5000 \
--priority max \
--fairness round_robinfrom flyteplugins.union.remote import Queue
queue = Queue.create(
"my-queue",
run_concurrency=100,
action_concurrency=1000,
)
print(queue.to_dict())Create a higher-priority queue in a specific pool:
queue = Queue.create(
"gpu-queue",
cluster_pool="prod",
clusters=["prod-us-east-1"],
run_concurrency=50,
action_concurrency=500,
depth=5000,
priority="max",
fairness="round_robin",
)In the console, go to Settings > Queues and click New Queue. Fill in the New queue form and click Create queue. The fields map to the same settings the CLI and Python expose:
| Form field | Setting |
|---|---|
| Name | the queue name |
| Priority | priority (shown as Low / Medium / High, see below) |
| Cluster pool | cluster_pool / --cluster-pool |
| Clusters | clusters / --cluster (“All available clusters (default behavior)” routes to every cluster in the pool) |
| Depth | depth / --depth |
| Run concurrency | run_concurrency / --run-concurrency |
| Action concurrency | action_concurrency / --action-concurrency |
The console labels priority Low, Medium, and High; these are the same
levels the CLI and Python call min, medium, and max. Fairness is not in
the form, so set it from the CLI or Python if you need a value other than the
default.
Every queue is bound to a cluster pool, chosen at creation time with
cluster_pool in Python or --cluster-pool in the CLI. If you omit it, the
queue is bound to the default cluster pool.
What each setting controls
cluster_pool/--cluster-pool: the pool this queue lives in. A queue can only route to clusters in its own pool. Omit to bind the queue to thedefaultpool.clusters/--cluster: pin the queue to one or more clusters in the pool. Omit to use all clusters in the pool. In the API,["*"]means allactive, healthy clusters in the pool (see Wildcard routing), and*must be the only entry if used.run_concurrency/--run-concurrency: maximum number of runs active on the queue at once. Children of an active run aren’t counted; use this to stop a job from overlapping with a previous invocation of itself.0means no limit.action_concurrency/--action-concurrency: maximum number of actions (tasks) running at once. A cap of 1 serializes the queue; higher values bound the burst rate.0means no limit.depth/--depth: total in-flight plus waiting items the queue will hold (default10000).0means no limit.priority/--priority:min,medium(default), ormax. Among queues contending for the same pool’s capacity, higher-priority work is scheduled first. Under the hood these map to enum values 1, 50, and 100; usemaxfor a priority higher than 50. Priority controls ordering, not preemption.fairness/--fairness:round_robin(default) orshuffle_interleave. This controls how actions from different projects sharing the queue are interleaved.
Inspect queues
# List all queues
flyte get queue
# List live queues in one state: active, draining, drained, or deleting
flyte get queue --state draining
# Inspect one queue's settings and status
flyte get queue gpu-queue
# Stream live metrics — runs in-flight, actions in-flight, queue depth
flyte get queue gpu-queue --watch
# List soft-deleted queues (hidden from the plain listing)
flyte get queue --deleted--watch renders live progress bars for run concurrency, action concurrency, and
depth, so you can see a queue filling up or draining in real time. Metrics are
available while a queue is active, draining, drained, or deleting.
Watching a draining queue shows work finishing normally; watching a deleting
queue shows its cleanup progress. A deleted queue cannot be
watched until you restore it.
Fetching a queue by name works even after it has been
deleted: flyte get queue <name> returns the soft-deleted
queue carrying its deletion time instead of failing, so a queue that an old run
once targeted stays inspectable. Only the listing hides deleted queues, unless
you pass --deleted; --state deleted without --deleted returns nothing for
the same reason.
from flyteplugins.union.remote import Queue
for queue in Queue.listall(limit=100):
print(queue.name, queue.status, queue.priority, queue.cluster_pool, queue.clusters)
# Narrow the live listing to one lifecycle state: "active", "draining",
# "drained", or "deleting"
for queue in Queue.listall(state="draining"):
print(queue.name)
# Deleted queues are hidden from the listing unless you ask for them
for queue in Queue.listall(deleted=True):
print(queue.name)
queue = Queue.get("gpu-queue")
print(queue.to_dict())Queue.get works even on a deleted queue: it returns the
soft-deleted queue carrying its deletion time instead of failing, while
Queue.listall hides deleted queues.
metrics = Queue.details("gpu-queue")
print(metrics)To stream metrics:
for metrics in Queue.watch("gpu-queue"):
print(metrics)Queue.details and Queue.watch work on live queues, including draining,
drained, and deleting queues. Unlike Queue.get, neither works on a
deleted queue.
Go to Settings > Queues to see all your queues.
Queues are grouped by cluster pool (the View by Pool toggle), and each row shows its status, priority, and live Queued, Runs, and Actions counts. Use the Status filter or the search box to narrow the list.
Click a queue to open its detail view, which has three tabs.
Overview shows the queue’s live state: its cluster pool, the clusters it is connected to and their CPU, GPU, and memory capacity, and the in-flight Queued, Runs, and Actions counts.
Usage gives ready-to-copy snippets for routing work to this queue, at run level and per task, with the queue’s name already filled in. These are the same routing methods described in Queues in Configure tasks.
Settings lists the queue’s current configuration: its pool, connected clusters, and scope, plus its priority, depth, and run and action concurrency limits.
Change a queue’s settings
You can update limits, priority, fairness, or cluster pinning. The update API replaces the full queue spec; the Python wrapper handles this by reading the current queue first, changing only the fields you pass, and writing the complete spec back.
flyte update queue gpu-queue --editThis opens the queue in your $EDITOR so you can adjust the mutable settings.
from flyteplugins.union.remote import Queue
Queue.update(
"gpu-queue",
run_concurrency=75,
action_concurrency=750,
priority="max",
clusters=["prod-us-east-1"],
)Changing the cluster selector within the same pool (which clusters the queue pins to) takes effect immediately because every cluster in the pool shares the same data plane.
Changing the queue’s pool (cluster_pool in the YAML or Python) is allowed
only when the queue is drained, the destination
pool exists, and every cluster in the queue’s selector is a member of the
destination pool — see Move work to another pool.
A cluster’s co-named queue rejects selector and
pool changes entirely: those are managed by its cluster. A soft-deleted queue
cannot be updated at all; undelete it first.
Queue lifecycle
A queue is in one of five lifecycle states. The drain, activate, delete,
and undelete transitions are requested by you; the moves into drained and
deleted are made by the system once it confirms the queue holds nothing.
stateDiagram-v2
direction LR
[*] --> active: create
active --> deleting: cluster deleted (co-named queue only)
draining --> deleting: delete
deleting --> deleted: cleanup done (system)
active --> draining: drain
draining --> active: activate
draining --> drained: no work left (system)
drained --> active: activate
drained --> deleted: delete
deleted --> drained: undelete
active: accepting new submissions.draining: no new submissions; in-flight work runs to completion. The system moves the queue todrainedonce it confirms the queue holds no run or cleanup work.drained: confirmed empty. The only state from which a delete completes in one step, the only state in which the queue can change pools, and the state a restored queue comes back in.deleting: deletion requested while the queue may still hold work. The system aborts the queue’s remaining run work, waits for cleanup work, and then moves the queue todeleted. The queue stays listed meanwhile.deleted: soft-deleted. Hidden from listings and never scheduled on; the name stays reserved andflyte get queue <name>still returns it.
What each operation does from each state:
| Current state | --drain |
--activate |
flyte delete queue |
flyte undelete queue |
Pool change |
|---|---|---|---|---|---|
active |
→ draining |
no change | rejected (drain first) | rejected | rejected |
draining |
no change | → active |
→ deleting |
rejected | rejected |
drained |
rejected (already drained) | → active |
→ deleted |
rejected | allowed |
deleting |
rejected | rejected | rejected | rejected | rejected |
deleted |
rejected | rejected | rejected | → drained |
rejected |
A cluster’s co-named queue is the one exception in the --activate column:
while its cluster is draining or drained, the queue cannot be activated on
its own. Activating the cluster activates it.
Three rules follow from the table:
- Deleting a queue is a safe archive operation. A queue must always be
drained before it can be deleted: there is no path from
activeinto deletion that you can request on a queue, so deleting a queue that is accepting work always takes two calls, a drain and then a delete. A single misdirected command can never take a serving queue away, and the delete itself only archives the record. The one indirect path is deleting a cluster, which takes the cluster’s co-named queue into deletion whatever state it is in; that path is only as safe as the cluster delete that triggers it. - Deletion cannot be canceled: a
deletingqueue can only becomedeleted, and only then can it be undeleted. drainedanddeletedare never requested directly; the system transitions into them once it has confirmed the queue holds nothing.
Draining a cluster drains its co-named queue and activating the cluster activates it; see How the co-named queue follows its cluster.
Drain and reactivate a queue
Draining takes a queue out of rotation without losing in-flight work: the
queue stops admitting new submissions, work already in flight runs to
completion, and once nothing is left the queue settles into the drained state.
Use it before maintenance, deletion, or as part of
moving work to another pool.
flyte update queue gpu-queue --drain # stop new submissions; let in-flight work finish
flyte update queue gpu-queue --activate # put the queue back in rotationfrom flyteplugins.union.remote import Queue
Queue.drain("gpu-queue") # stop new submissions; let in-flight work finish
Queue.activate("gpu-queue") # put the queue back in rotationTo see where things stand across the organization, filter listings by state:
flyte get queue --state draining shows the queues still finishing in-flight
work, --state drained the ones that are done, and --state deleting the ones
being removed.
An active queue must start draining before it can be deleted,
but you do not have to wait for drained: deleting a draining queue begins its
cleanup immediately. A pool move does require the
queue to be fully drained. Deleting a cluster automatically takes its co-named
queue into deletion, whatever state that queue is in.
Any queue can be drained, including the default queue and a queue configured
as the default run queue (run.default_queue) in your organization’s settings
at any scope — draining is how you stop scheduling on it, and the prerequisite
for deleting it. A draining or drained queue still resolves
as the default: runs that land on it are rejected with an error saying the
queue is not accepting work. Only deleting a queue
referenced in settings is refused.
Delete a queue
A queue must be draining or drained before it is deleted: an active queue is
rejected, so start draining it first. This is what makes deleting a queue a
safe archive operation, in contrast to
deleting a cluster, which can interrupt work.
Once the queue is draining you can choose whether to wait:
- Deleting a
drainingqueue moves it todeleting. Queued and running actions on it are failed rather than finished, as a final attempt that is not retried; cleanup work is left to run. The system moves the queue todeletedonce nothing remains on it. - Deleting a
drainedqueue moves it directly todeleted, because the system has already confirmed that no work or cleanup remains.
Deletion cannot be canceled. A deleting queue rejects every further request:
it cannot be drained, activated, deleted again, or undeleted. Wait until it
becomes deleted before undeleting it.
A queue referenced as the default run queue (run.default_queue) in settings at
any scope can be drained but not deleted: update or unset those settings first.
A cluster’s co-named queue can be deleted on its
own while the cluster lives. Deleting the cluster also deletes its co-named queue
and is the only operation that can take that queue directly from active to
deleting; that cascade does not apply the run.default_queue check, so a
co-named queue used as a default run queue goes away with its cluster.
The default queue is no exception: drain it, then delete it, like any other
queue — deleting it is also how the default pool is emptied of live queues so
that the pool itself can be deleted. While the
default queue is deleted, submissions that name no queue and have no
run.default_queue setting to fall back on are rejected at creation. Nothing
re-creates it implicitly; flyte undelete queue default brings it back.
flyte delete queue gpu-queue
flyte delete queue gpu-queue --yes # skip the confirmation prompt
# List deleted queues, restore one
flyte get queue --deleted
flyte undelete queue gpu-queuefrom flyteplugins.union.remote import Queue
queue = Queue.delete("gpu-queue")
print(queue.status) # deleting or deleted
Queue.undelete("gpu-queue") # restore a deleted queueA deleting queue remains visible in normal listings while cleanup runs, and it
can still be inspected or watched. Once deleted, it is soft-deleted: it
disappears from normal listings and is no longer scheduled on, but its name stays
reserved. Fetching it by name still works — flyte get queue <name> and
Queue.get return the queue with its deletion time. Use
flyte get queue --deleted to list deleted queues.
Restore it with flyte undelete queue <name>. The queue’s pool and pinned
clusters must still be live. A cluster’s co-named queue undeletes like any other
queue while its cluster is live; while the cluster is deleted, the undelete is
refused and undeleting the cluster brings the
queue back with it. The restored queue comes back drained; reactivate it to
resume routing.
Move work to another pool
Moving work to a different pool crosses an isolation boundary. In-flight runs have already landed their data, containers, code, and secrets in the old pool’s data plane, and a different pool’s clusters cannot read them. So work in flight never moves: the migration is always drain-first, in one of two shapes.
Move the queue itself. A queue can change pools once it holds no work:
- Drain the queue and wait for it to reach the
drainedstate. - Update the queue’s pool —
flyte update queue <name> --editand changecluster_poolin the YAML, orQueue.update(name, cluster_pool=...)— and set its cluster selector to clusters that are members of the destination pool (or*). - Reactivate the queue.
Anything targeting the queue by name follows it to the new pool. New submissions made while it is drained are rejected, so coordinate the window with the queue’s users.
Replace the queue. Alternatively, keep the old queue’s pool binding and shift traffic over:
- Create a new queue in the destination pool.
- Update workflows, launch plans, triggers, or run overrides to target the new queue.
- Drain the old queue to shut out straggler submissions and let in-flight work finish.
- Delete the drained queue, or leave it idle — an idle queue costs nothing.
A task can override its queue at runtime
(task.override(queue=...)),
but only to another queue in the same pool as the run’s original queue. A
cross-pool override is rejected, for the same data plane reason that moving
work between pools requires a drain-and-replace migration.
See also
- Queues in Configure tasks: routing work to a queue from task code, triggers, and per-run context.
- Cluster pools and Clusters: the routing targets a queue points at.