Concept and mechanism
A Swarm service declares an intent that managers try to maintain. In a replicated service, intent includes task count; in a global service, there is one task per eligible node. Eligibility depends on conditions such as availability and placement constraints. Free CPU does not resolve a missing required label. Distinguish task execution from manager consensus: five voters require three to accept changes. Existing tasks may continue when that majority is lost, but this does not establish capacity to recover another failure. More workers do not replace voters. A green application dashboard can therefore coexist with a serious control-plane problem.
Guided application
In fictional APS maintenance, begin by inventorying managers, workers, tasks, and independent containers. Confirm health before each stop and define a pause point when quorum headroom disappears. Drain prevents new Swarm tasks and leads to rescheduling existing ones, but does not empty containers created with docker run or compose up. Those workloads need separate handling. For a Pending task, collect the scheduler message and compare constraints, labels, resources, and availability before rebuilding the image. In handover, record who decides to defer the window and what evidence demonstrates recovery. Business availability and cluster-management capability should appear as separate conditions.
Five managers, two unavailable: the next stop requires restoring headroom first.
Common pitfalls
HTTP as quorum; CPU as eligibility; Drain as stopping everything; global as fixed replicas.
Related topics: Delivery, rollback, and readiness · Reproducible images and registries · Daemon, logs, and recovery
Relate desired state, eligible nodes, and majority before intervening.
Reference: Swarm administration · DCA Study Guide v1.5 (January2025); current exam listing checked2026-09-30