clusterctl · for the people who run HPC clusters

One binary.
Every node.
No surprises.

Select nodes with ClusterShell syntax, fan out commands, drive service processors, reinstall nodes and administer Slurm. Anything that changes something shows you first and asks.

$mise use -g github:GSI-HPC/clusterctl
$ clusterctl node select '@idle&@rack:R03' --count 39 allocated idle drained down
1

static binary. No interpreter, no virtual environment, no agent on the nodes.

0

passwords in an argument vector. Secrets are decrypted in memory and streamed.

7

configuration layers, one schema. config explain says which one won.

--dry-run

on every command that changes something, and a question before it does.

Nouns and verbs, not scripts.

Every command is a noun then a verb, and the global flags mean the same thing in all of them.

Asks before it breaks things.

Above a threshold you type the number of hosts, because a y is easy to type by reflex. A protected host is refused, and so is powering off a node that runs a Slurm job, unless you say --force.

$ clusterctl bmc power off -n '@rack:R02'About to power off 10 hosts: exe[0001-0010]  through RedfishThis is more than 8 hosts. Type the number of hosts to continue: 10