clusterctl · for the people who run HPC clusters
Select nodes with ClusterShell syntax, fan out commands, drive service processors, reinstall nodes and administer Slurm. Anything that changes something shows you first and asks.
mise use -g github:GSI-HPC/clusterctl$ clusterctl node select '@idle&@rack:R03' --count
39
allocated idle drained downstatic binary. No interpreter, no virtual environment, no agent on the nodes.
passwords in an argument vector. Secrets are decrypted in memory and streamed.
configuration layers, one schema. config explain says which one won.
on every command that changes something, and a question before it does.
Every command is a noun then a verb, and the global flags mean the same thing in all of them.
node select '@slurm:main&@rack:R02'Node sets with groups from the inventory, the racks and Slurm, and set operations that mean what they say.
exec --dedup -- uname -rA thousand nodes answer; the ones that agree collapse into one line. Your quoting arrives intact.
bmc power on -n '@rack:R02'Redfish first, pinned certificates, and a rack powered on in batches so the breaker stays in.
provision reinstall --dry-runPXE and GRUB boot paths, DHCP checks and secrets streamed to the node, rehearsed first.
slurm node drain 'ticket 4711'A drain needs a reason, so nobody has to guess before resuming. The queue, accounting and fair share too.
mcp serveAn AI agent reads the cluster and plans a change. It cannot apply one you have not confirmed.
Above a threshold you type the number of hosts, because a y is easy to type by reflex. A protected host is refused, and so is powering off a node that runs a Slurm job, unless you say --force.
$ clusterctl bmc power off -n '@rack:R02'About to power off 10 hosts: exe[0001-0010] through RedfishThis is more than 8 hosts. Type the number of hosts to continue: 10