Manual
Manual
clusterctl administers one or more HPC clusters from an administrator’s workstation. It replaces a toolkit of shell scripts with a single binary, one layered configuration and a command line that means the same thing everywhere.
Where to start
Install it, write a configuration, run the first commands.
How to do each job: select nodes, fan out, power, reinstall, Slurm.
Configuration fields, exit codes, environment variables.
Every command and flag, generated from the binary.
The shape of a command
Every command is a noun then a verb, and the global flags mean the same thing in all of them:
$ clusterctl [--context C] [-n NODESET] [-o FORMAT] [--dry-run] [-y] NOUN VERB
| Flag | Meaning |
|---|---|
-n, --nodes | The node set to act on |
-o, --output | table, wide, json, yaml, nodeset, name, jq=… |
--context | Which cluster to act on |
--dry-run | Say what would happen and change nothing |
-y, --yes | Answer the confirmations in advance |
--force | Allow a protected host, or a node the inventory does not know, to be touched |
--set | Override one configuration value for this command |
--fanout | How many hosts to work on at once; a lower value lowers the bounds of service processors and names too |
--progress | auto, tty, counter, plain or none: how progress is shown on standard error |
--progress-log | A file to append the command’s progress events to, as JSON lines |
A note on safety
Anything that powers off, reinstalls, drains or overwrites shows you what it is
about to do and asks. Above a configured host count it asks you to type the
number of hosts, because a y is easy to type by reflex. Without a terminal to
ask on, it refuses rather than proceeding — pass -y when you mean it.
--dry-run is the first thing to reach for on a command you have not run
before. It prints what would be changed and changes nothing. Lookups still run
for real, such as asking Slurm for a group or whether a node is running a job,
so a dry run refuses what the real run would refuse.