clusterctl slurm node drain
clusterctl slurm node drain
Take nodes out of production
Synopsis
Set the nodes to drain so that no new job is scheduled on them. Jobs already running are left alone.
The reason is mandatory and comes first, because a drained node with no reason is a node nobody dares resume. Say what is wrong and where it is tracked, on one line of at most 200 characters and without |. A reason that names nodes is refused, since that is what a forgotten reason looks like, and so are nodes named both after the reason and with -n.
The nodes are checked against Slurm before anything is shown: slurmctld reads ALL as every node and a NodeSet name of slurm.conf as its members, so a name Slurm does not read as exactly the node named is refused.
clusterctl slurm node drain ’ticket 4711: failing DIMM’ -n exe0007
clusterctl slurm node drain REASON [NODESET] [flags]Options
-h, --help help for drainOptions inherited from parent commands
--config strings configuration file or directory to read, repeatable (default: CLUSTERCTL_CONFIG or the search path)
--context string context to act on (default: the current one)
--dry-run report what would be done and change nothing
--fanout int how many hosts to work on at once, at least 1 (default: from the configuration)
--force allow protected hosts, and nodes the inventory does not know, to be touched
-n, --nodes stringArray node set to act on, for example 'exe[1-10],@rack:R02' (default: CLUSTERCTL_NODES)
-o, --output string output format: table, wide, json, yaml, nodeset, name, jsonpath=, jq= (default "table")
--set stringArray override one configuration value as PATH=VALUE, repeatable
-y, --yes answer the confirmation prompts with yesSEE ALSO
- clusterctl slurm node - Read and change the state of the nodes