Cluster Configuration
The simplest valid configuration creates a cluster with one controller and one worker using the default image:
kind: Cluster
This is equivalent to the fully expanded form:
kind: Cluster
name: default
defaults:
image: ghcr.io/gsi-hpc/sind-node:latest
cpus: 1
memory: 512m
tmpSize: 256m
nodes:
- role: controller
- role: worker
kind: Cluster
name: test-cluster
defaults:
image: ghcr.io/gsi-hpc/sind-node:latest
cpus: 1
memory: 512m
tmpSize: 256m
storage:
dataStorage:
type: volume
mountPath: /data
slurm:
main: |
SelectType=select/cons_tres
SelectTypeParameters=CR_Core_Memory
cgroup: |
ConstrainCores=yes
nodes:
- role: controller
cpus: 2
memory: 1g
tmpSize: 512m
- role: submitter
- role: worker
count: 3
cpus: 2
memory: 1g
- role: worker
count: 2
managed: false
| Field | Required | Default | Description |
|---|---|---|---|
kind |
yes | — | Must be "Cluster" |
name |
no | "default" |
Cluster name, used in resource naming; must be a valid name (see below) |
realm |
no | "sind" |
Realm namespace for resource isolation; must be a valid name (see below) |
defaults |
no | — | Default settings applied to all nodes |
storage |
no | — | Shared storage configuration |
slurm |
no | — | Slurm configuration extension |
nodes |
no | 1 controller + 1 worker | Node definitions |
Cluster and realm names end up in Docker resource names, DNS names and paths, so each must be a single DNS label: lowercase ASCII letters, digits and -, 1 to 63 characters, not beginning or ending with -. Names such as Dev, my_cluster, dev.test or ../x are rejected. The same rule applies to cluster names given on the command line and to --realm and SIND_REALM.
The defaults section sets values inherited by all nodes unless overridden at the node level.
| Field | Default | Description |
|---|---|---|
image |
ghcr.io/gsi-hpc/sind-node:latest |
Container image |
cpus |
1 |
CPU limit per container |
memory |
"512m" |
Memory limit per container |
tmpSize |
"256m" |
tmpfs size for /tmp |
capAdd |
none | Extra Linux capabilities |
capDrop |
none | Dropped Linux capabilities |
devices |
none | Host devices to expose |
securityOpt |
none | Extra security options |
Scalar fields (image, cpus, memory, tmpSize) are overridden by per-node values. List fields (capAdd, capDrop, devices, securityOpt) are merged with per-node values.
storage:
dataStorage:
type: hostPath # "hostPath" or "volume"
hostPath: ./data # host directory for type: hostPath
mountPath: /data # default: /data
| Field | Default | Description |
|---|---|---|
type |
set by --data |
"hostPath" bind-mounts hostPath; "volume" uses the Docker volume <realm>-<cluster>-data and ignores hostPath. A hostPath without type means "hostPath" |
hostPath |
— | Host directory, required with type: hostPath. A relative path is taken relative to the directory sind create cluster runs in and stored as an absolute path |
mountPath |
"/data" |
Absolute mount point inside the nodes |
If the config sets neither type nor hostPath, the --data flag of sind create cluster decides; its default, ., bind-mounts the working directory (see Data mount).
Workers added with sind create worker mount the data at the same place, sind get cluster reports it there, and sind enter and sind exec start in it.
The slurm section extends the generated Slurm configuration. Each key maps to a config file:
| Key | Config file | sind generates defaults |
|---|---|---|
main |
slurm.conf |
yes |
cgroup |
cgroup.conf |
yes |
gres |
gres.conf |
no |
topology |
topology.conf |
no |
plugstack |
plugstack.conf |
yes (always scaffolded) |
Each key supports two forms:
String form — content appended to the config file:
slurm:
main: |
SelectType=select/cons_tres
SelectTypeParameters=CR_Core_Memory
cgroup: |
ConstrainCores=yes
Map form — named fragments placed in a .conf.d/ directory:
slurm:
main:
scheduling: |
SchedulerType=sched/backfill
SchedulerParameters=bf_continue
resources: |
SelectType=select/cons_tres
Fragment validation:
- Names must be plain filenames (no path separators)
- Names and content must not be empty
See Slurm Configuration for details on the generated files.
kindmust be"Cluster"- Exactly one
controllernode is required - At most one
submitternode is allowed - At least one
workernode is required countis only valid for worker nodesmanagedis only valid for controller and worker nodes- With
managed: falseon the controller, no worker may setmanaged: trueand noslurmsection may be set backupControlleris only valid for controller nodes- With
backupController,slurm.mainmust not setSlurmctldHost(or its deprecated formsControlMachine,BackupController,BackupAddr) orStateSaveLocation countmust not be negative;0means the default, 1capAdd/capDropvalues must be recognized Linux capability names (e.g.SYS_ADMIN,NET_ADMIN,ALL)devicespaths must be absolute (start with/)storage.dataStorage.typemust bevolumeorhostPath;hostPathrequires ahostPath, andmountPathmust be absolute- Unknown keys are rejected