Selecting nodes
Every command that acts on nodes takes a node set. The syntax is ClusterShell’s,
so it is what you already type at nodeset, clush and sinfo. A few corner
cases are decided differently; the padding note below is the one you are most
likely to meet.
The syntax
$ clusterctl node select 'exe0001' # one host
$ clusterctl node select 'exe[1-10]' # a range
$ clusterctl node select 'exe[0001-0010]' # padded
$ clusterctl node select 'exe[1-10/2]' # with a step
$ clusterctl node select 'exe[1,5,9-12]' # several ranges
$ clusterctl node select 'rack[1-2]node[01-04]' # two dimensions
$ clusterctl node select '@compute' # a group
$ clusterctl node select '@slurm:main' # a group from a source
Combining sets
| Operator | Meaning | Example |
|---|---|---|
, or a space | union | exe[1-4],sub[1-2] |
! | difference | @compute!@static:infra |
& | intersection | @slurm:main&@rack:R02 |
^ | in exactly one | @yesterday^@today |
@a!@b&@c is ((@a minus @b) intersect @c). Write the order you
mean; there are no brackets for grouping.!, & and ^ need something on both sides. When a command substitution
prints nothing, -n "@rack:R02&$(clusterctl slurm node nodeset idle)" becomes
@rack:R02&, which is an error rather than the whole rack:
$ clusterctl node select '@rack:R02&'
clusterctl: in "@rack:R02&": the & operator has no right operand
Folding and expanding
$ clusterctl node select 'exe[1-3],exe7'
exe[0001-0003,0007]
$ clusterctl node select 'exe[1-3]' --expand
exe0001
exe0002
exe0003
$ clusterctl node select '@compute' --count
1024
exe0001, so a set
comes back under the names the site gave them even when you typed exe1.
exe1 and exe0001 are the same host: padding is not part of what a name
identifies. Every host keeps the name the inventory gave it, so a stray exe11
next to exe[0001-0010] prints as exe[0001-0010,11], never as exe0011.
ClusterShell would treat exe1 and exe0001 as two hosts.What a name may contain
A node name is used as a host name, so every name a set resolves to has to be
one: letters, digits and hyphens, in labels separated by dots, with no label
beginning or ending with a hyphen. Anything else, such as a name beginning
with - or containing :, @, /, ?, #, _ or a space, is refused
with exit code 2 before anything is sent. This holds for names that come from
a group source or the inventory as well as for names you type, and it is what
stops a name from turning into an ssh option or sending a Redfish request,
with the BMC password, to another host.
$ clusterctl node hw -n '-oProxyCommand=...'
clusterctl: "-oProxyCommand=..." is not a host name: it begins with -
One machine, one name
A name is turned into the name the inventory uses for the machine it refers to. Case and a final dot make no difference, other padding is resolved, and a node’s host name, its service processor name and the addresses the inventory records for it all name the node. Names that turn out to be one machine are one node:
$ clusterctl node select 'exe0001,EXE1,exe0001.hpc.example.org.,10.0.2.1'
exe0001
So a command runs once on a machine however many ways it was named, and a
protected host is refused however it was written. A name with a domain the
naming rules do not give it is left as it is, and a change to a name the
inventory does not know needs --force.
Groups
A group is resolved by a source. Which sources exist is configuration:
$ clusterctl node groups
SOURCE GROUP
inventory exe
inventory sub
inventory wlm
rack R01
rack R02
slurm main
slurm debug
A bare @group searches the default source first and then the others, so
@exe works without a prefix. @source:group names one explicitly, which is
what to write when two sources could both answer.
The search only moves on when a source answers that it has no such group. If a source cannot be asked, because its host does not answer or its command fails, the command stops with that source’s error rather than taking a group of the same name from another source:
$ clusterctl node select '@compute'
clusterctl: group @compute: source "slurm": login (login.hpc.example.org): ssh: connect to host login.hpc.example.org port 22: Connection timed out; the search for @compute stops at a source that cannot answer (before it: source "inventory": no node has class=compute; source "rack": no node has rack=compute)
$ echo $?
3
slurm source maps a group to a partition, so @slurm:main is every
node of the partition main. Slurm states such as idle or drained are not
groups. Ask for them with slurm node nodeset and pass the result on, as
shown under Passing a set to another tool.$ clusterctl node groups exe0007
SOURCE GROUPS
inventory exe
rack R02
slurm main
Where groups come from
groups:
defaultSource: inventory
sources:
inventory:
# One group per value of a node attribute. This is what a genders
# file provided: @exe, @wlm, @dbm.
attribute: class
rack:
attribute: rack
static:
# A table. A group may refer to other groups.
static:
infra: wlm01,dbm01
compute: "@inventory:exe"
slurm:
# Asked of the workload manager and cached for a minute.
cacheTtl: 60s
exec:
role: login
map: [sinfo, -h, -o, "%N", -p, $GROUP]
all: [sinfo, -h, -o, "%N"]
list: [sinfo, -h, -o, "%R"]
reverse: [sinfo, -h, -N, -o, "%R", -n, $NODE]An exec source is given an argument vector, and $GROUP and $NODE are
substituted as whole arguments. A group name is never interpreted by a shell.
The commands run on the host role the source names, which is required, and are
bounded by fanout.commandTimeout like every remote command.
A cached answer is kept per site, cluster and context, per host and per exact
command, so two clusters that both have a partition main never answer for
each other. The list answer is cached too, except when a bare @group search
asks whether a group is a source’s: that is asked again, so that a group added
since does not send the search on to another source.
A selection for a whole session
$ export CLUSTERCTL_NODES='@rack:R02'
$ clusterctl bmc status
$ clusterctl exec -- uptime
-n or a node set argument always wins over it. Nothing else falls back to a
set you set earlier: a command with no selection stops rather than guessing.
Give a node set once. -n together with a node set argument, or -n twice, is
refused with exit code 2 rather than one of them being ignored; write the union
into one expression instead, such as -n exe0001,exe0002.
Passing a set to another tool
$ clusterctl slurm node nodeset drain
exe[0007,0042,0511]
$ clusterctl exec -n "$(clusterctl slurm node nodeset idle)" -- uptime
$ clusterctl node select "@compute!$(clusterctl slurm node nodeset drain)" --count
When no node is in that state, the inner command prints nothing and -n is
given an empty set. That is refused with exit code 2, even with
CLUSTERCTL_NODES set: an explicit -n is never replaced by the session set.
clusterctl hands Slurm and FreeIPMI a set with at most one bracketed range per
name, rack1node[001-100],rack2node[001-100] rather than
rack[1-2]node[001-100], and with groups already resolved. You never have to
think about it.
Checking a set before you use it
$ clusterctl node select '@rack:R02&@slurm:main' --expand
$ clusterctl node list '@rack:R02'
$ clusterctl node fqdn -n '@rack:R02'
Worth doing before anything destructive. --dry-run also prints the set it
resolved before it stops.