clusterctl exec

clusterctl exec

clusterctl exec

Run one command on many nodes at once

Synopsis

Run a command on every node of a set, in parallel.

The command follows –, so that none of its options is read as one of clusterctl’s: without it, a -r or -n meant for the node would change what clusterctl does. The node set goes before –, or in -n, but not in both.

The argument vector is quoted once and reassembled by the remote shell, so it arrives exactly as it was typed. A timeout is enforced on the node with timeout(1), because killing the local ssh would leave the remote process running. A node that has not answered five seconds after the timeout, plus the time ssh.connectTimeout and ssh.connectionAttempts let reaching it take, has stopped answering: ssh is stopped here, and the node is reported as unreachable.

clusterctl exec -n @slurm:main – uptime clusterctl exec exe[1-10] –dedup – uname -r clusterctl exec -n exe[1-4] –script ‘systemctl is-active slurmd || journalctl -u slurmd -n 5’

A protected host is refused unless –force is given, on every run. exec does not ask before it runs, unless –confirm is given. With –stdin the payload takes up standard input, so the confirmation cannot be read from it and –confirm needs -y.

With –dedup the nodes that answered the same thing, and ended the same way, are collapsed into one block, which turns a thousand replies into the few worth reading. Control characters in what the nodes answered are shown as escapes rather than passed to the terminal.

clusterctl exec [NODESET] -- COMMAND... [flags]

Options

      --confirm            ask before running, as the destructive commands do; with --stdin, give -y as well
  -b, --dedup              collapse nodes that answered the same thing
  -h, --help               help for exec
  -r, --root               run as root
      --script string      shell program to run instead of a command
      --stdin              read standard input once and send it to every node
      --timeout duration   how long the command may run on a node (default: from the configuration)
  -u, --user string        remote account to run as

Options inherited from parent commands

      --config strings        configuration file or directory to read, repeatable (default: CLUSTERCTL_CONFIG or the search path)
      --context string        context to act on (default: the current one)
      --dry-run               report what would be done and change nothing
      --fanout int            how many hosts to work on at once, at least 1; caps the service processors and the names asked at once too, which fanout.max and CLUSTERCTL_FANOUT do not (default: from the configuration)
      --force                 allow protected hosts, and nodes the inventory does not know, to be touched
  -n, --nodes stringArray     node set to act on, for example 'exe[1-10],@rack:R02' (default: CLUSTERCTL_NODES)
  -o, --output string         output format: table, wide, json, yaml, nodeset, name, jq= (default "table")
      --progress string       how to show the progress of a command on standard error: auto, tty, counter, plain, none (default: CLUSTERCTL_PROGRESS, else auto, a live tree when standard error is a terminal; plain writes lines for a log)
      --progress-log string   append the progress events of a command to this file, one JSON object per line, created readable by you alone (default: CLUSTERCTL_PROGRESS_LOG; an empty one writes none)
      --set stringArray       override one configuration value as PATH=VALUE, repeatable
  -y, --yes                   answer the confirmation prompts with yes

SEE ALSO

  • clusterctl - Administer HPC clusters from one binary