Skip to main content
sind
Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Toggle Dark/Light/Auto mode Back to homepage

Node Definitions

Node roles

Role Count Required Slurm daemons Description
controller exactly 1 yes slurmctld Cluster controller
submitter 0–1 no none (clients only) Job submission node
worker 1+ yes slurmd Worker nodes

Node parameters

Parameter Scope Default Description
image global + per-node ghcr.io/gsi-hpc/sind-node:latest Container image
cpus global + per-node 1 CPU limit
memory global + per-node "512m" Memory limit
tmpSize global + per-node "256m" tmpfs size for /tmp
count worker only 1 Number of worker nodes
managed controller + worker true Worker: start slurmd and add to slurm.conf. Controller: false makes the whole cluster unmanaged (see below)
backupController controller only false Add a backup controller, controller-backup (see below)
capAdd global + per-node none Extra Linux capabilities (e.g. SYS_ADMIN)
capDrop global + per-node none Dropped Linux capabilities
devices global + per-node none Host devices to expose (e.g. /dev/fuse)
securityOpt global + per-node none Extra security options

Per-node scalar values override the defaults section. Security list fields (capAdd, capDrop, devices, securityOpt) are merged with defaults rather than replacing them.

Shorthand syntax

Nodes can be specified in a compact form when only role and count are needed:

nodes:
  - controller               # bare role string
  - submitter
  - worker: 3                # role with count

This is equivalent to:

nodes:
  - role: controller
  - role: submitter
  - role: worker
    count: 3

Shorthand and full forms can be mixed in the same configuration.

Managed vs unmanaged workers

By default, workers are managed: sind starts slurmd and adds the node to sind-nodes.conf so Slurm knows about it. This is the typical setup.

Unmanaged workers (managed: false) are created as containers but sind does not start slurmd and does not add them to the Slurm configuration. This is useful for:

  • Testing manual node registration workflows
  • Simulating nodes that join the cluster later
  • Running custom Slurm configurations
nodes:
  - role: worker
    count: 3              # 3 managed workers

  - role: worker
    count: 2
    managed: false        # 2 unmanaged workers

Unmanaged workers can also be created dynamically:

sind create worker --count 2 --unmanaged

In an unmanaged cluster every worker is unmanaged.

Backup controller

backupController: true on the controller node adds a second controller, controller-backup, and configures Slurm’s active/passive controller pair:

nodes:
  - role: controller
    backupController: true
  - role: worker
    count: 2
  • controller-backup uses the same image, resources, volumes, capabilities, devices and security options as controller.
  • slurm.conf lists both hosts (SlurmctldHost=controller, then SlurmctldHost=controller-backup) and sets SlurmctldTimeout=20 unless the main section sets it.
  • Both controllers mount the shared state volume <realm>-<cluster>-state at /var/spool/slurmctld (StateSaveLocation).
  • slurmctld runs on both. The backup takes over when the primary stops responding for SlurmctldTimeout seconds, and hands control back when the primary’s slurmctld starts again.

SlurmctldHost and StateSaveLocation cannot be set in slurm.main when the backup is enabled. See Controller Failover for checking which controller is in control and triggering a failover.

Unmanaged cluster

managed: false on the controller makes the whole cluster unmanaged: sind creates the nodes, the volumes and the munge key, but writes no Slurm configuration and starts no Slurm daemon, so your own tooling can provision Slurm.

nodes:
  - role: controller
    managed: false
  - role: worker
    count: 2
  • Every worker is unmanaged. A worker with managed: true is rejected, and so is any slurm section.
  • backupController: true still adds controller-backup and the shared state volume.
  • managed is not valid on the submitter, which runs no Slurm daemon.

See Unmanaged Cluster for provisioning Slurm on such a cluster.

Capabilities and devices

sind’s default security posture avoids extra capabilities and device access. When specific use cases require them (e.g. testing CVMFS provisioning or FUSE-based filesystems), you can grant targeted privileges per node.

nodes:
  - role: controller
  - role: worker
    count: 3
    capAdd:
      - SYS_ADMIN
    devices:
      - /dev/fuse

Capability names follow Docker convention (without the CAP_ prefix). Device strings use Docker’s format: /dev/fuse or /dev/sda:/dev/xvda:rwm.

When set in the defaults section, security fields apply to all nodes. Per-node values merge with (not replace) defaults:

defaults:
  capAdd:
    - SYS_ADMIN
  devices:
    - /dev/fuse

nodes:
  - role: controller
  - role: worker
    count: 3
    capAdd:
      - NET_ADMIN    # workers get both SYS_ADMIN and NET_ADMIN

sind logs a notice at cluster creation when extra privileges are configured, making the escalation visible.

Workers created via sind create worker also support these fields:

sind create worker --cap-add SYS_ADMIN --device /dev/fuse

Default nodes

When the nodes section is omitted entirely, sind creates a minimal cluster:

nodes:
  - role: controller
  - role: worker