Collectors¶
nomad collect (usually nomad collect --once from cron, every 5 minutes) runs the collectors this host's nomad.toml turns on. Each writes its own tables; nomad sync copies every site's tables to the hub.
Which run¶
Each collector is turned on or off in its own table:
| Collector | Default | Where it belongs | Needs | Writes |
|---|---|---|---|---|
disk |
on | every host | df |
filesystems |
slurm |
on | Slurm head node | squeue, sacct |
queue_state, jobs |
job_metrics |
on | Slurm head node | sacct |
jobs, job_summary |
slurm_usage |
where slurm runs |
Slurm head node | sreport |
cluster_usage |
iostat |
on | every host | iostat (sysstat) |
iostat_device, iostat_cpu |
mpstat |
on | every host | mpstat (sysstat) |
mpstat_summary, mpstat_core |
vmstat |
on | every host | vmstat |
vmstat |
node_state |
on | Slurm head node | scontrol |
node_state |
gpu |
on | hosts with NVIDIA GPUs, or a head node reaching GPU nodes over SSH | nvidia-smi (DCGM if present) |
gpu_stats, gpu_health |
nfs |
on | hosts that mount NFS (client side) | nfsiostat (nfs-utils, in /usr/sbin) |
nfs_stats |
groups |
on | every site | getent; sacct for job accounting |
group_membership, job_accounting |
interactive |
off | hosts running RStudio or Jupyter | — | interactive_sessions, interactive_summary |
workstation |
off | a hub reaching workstations over SSH | ssh |
workstation_state, workstation_user_snapshot, ... |
per_user |
off | login nodes, shared interactive hosts | Python psutil |
per_user_sample, per_user_alert, per_user_daily |
storage |
off | a host reaching ZFS/NFS servers over SSH | ssh |
storage_state |
network_perf |
off | any host; tests paths that start there | ping; iperf3 for throughput |
network_perf |
cloud.aws |
off | sites with AWS resources | Python boto3 |
cloud_metrics |
An enabled = [...] list under [collectors], which older example configs carried, is not read.
groups reads membership with getent group: on the head node of each [clusters] entry (over SSH when it has a host), on the first reachable workstation of a type = "workstations" cluster with an ssh_user, and on this host otherwise. Job accounting needs sacct and is skipped without it. On the same host it looks up each member with getent passwd and records whether they still have an account (has_account): a hand-kept /etc/group goes on listing people whose accounts were deleted long ago. "No account" needs every run for two days to find none (getent says the same for an unreachable directory), and a run where the lookup fails keeps the last known answer. Every run stamps what it saw with one time (collected_at), so a membership no longer listed stops being stamped. A lab's members are the people listed in their cluster's latest run (give or take two days, for a failed run; a cluster not collected for 30 days while its site goes on is retired) who have an account; nomad console roles NETID says how many are not counted, and why.
nomad collect -C disk,nfs (or -C disk -C nfs) runs only those, if enabled. An unknown name is warned about and skipped; if no name is known, nothing runs and the command says so.
Is it working?¶
On a site, lists every collector: on or off and why, what it needs that this host lacks, and how its runs went over the last 7 days (from collection_log):
nfs on on by default
needs: nfsiostat
2,016 runs, never any data, last 2026-10-01T07:20
nfsiostat not installed (nfs-utils)
On the hub, nomad collectors --db ~/.local/share/nomad/combined.db shows each site's collectors and marks those that run but collect nothing.
A collector that runs but has nothing to collect logs why — nfsiostat not installed, no NFS mounts on this host, no NVIDIA GPU here, psutil not installed: nothing collected, no network_tests configured — in collection_log.error_message of a successful run. Before 1.7.13 these logged "success, 0 records" (or, for groups and per_user, "1 records" whatever they held), which looked like a working collector. A command that isn't installed fails the run at once with the reason, without retrying.
Commands are looked for on the PATH and in /usr/sbin, /sbin, /usr/local/sbin and /usr/local/bin: cron's PATH is only /usr/bin:/bin, and nfs-utils installs nfsiostat in /usr/sbin.
Jobs¶
The slurm collector reads running and pending jobs from squeue and every
user's jobs of the last job_history_days (7) from sacct --allusers
(without it, sacct run by anyone but root shows only that account's own
jobs). squeue is run with -a, so hidden partitions and partitions the
collector's account can't submit to are seen too. Jobs are recorded from every
partition. partitions limits the queue snapshot (queue_state, pending and
running counts per partition) to the partitions listed, and the old dashboard
(nomad dashboard) shows those partitions' nodes; empty means all. Before 1.7.43 it limited the jobs too,
and the setup wizard lists the partitions sinfo shows, which leaves out hidden
ones: jobs in the partitions left out reached jobs only through
job_metrics (finished, at least 10 seconds, a few states) and so every count
about jobs was short. Partitions that share nodes don't count a job twice: a
job runs in one partition, and questions about hardware go by the nodes in its
node list. A job submitted to several partitions is stored with the list
Slurm shows (a,b) while it waits, again after a requeue, and with the one it
runs in once it starts; an ended job reported with its list keeps the
partition it ran in. Jobs older than job_history_days from partitions that
were left out before 1.7.43 stay as job_metrics recorded them.
A job that ended between two runs leaves squeue; if that run's sacct
pull did not report it, it is looked up in sacct by number, and the answer
is taken only if its submit time matches (Slurm reuses job numbers after its
counter restarts). When sacct cannot account for the job yet, it is marked
UNKNOWN: it has ended, its outcome is not known, it has no exit code or
failure reason, and no success or failure count includes it. The next sacct
pull corrects it. A pending array range (123_[5-10]) that left the queue is
dropped, since each of its tasks has a row of its own. Nothing is concluded
about any job in a run where squeue did not answer.
Before 1.7.19 such jobs were marked COMPLETED, with an end time in UTC
where every other job time is local, so a failed job could count as a
success. Those rows (COMPLETED, an end time without a T) are looked up in
sacct too, 200 a run, and become UNKNOWN (end time moved to local) if
sacct no longer has them.
From sacct (1.7.42) each job also gets account, alloc_tres (what Slurm
allocated, e.g. cpu=6,gres/gpu:a40=1,gres/gpu=1,mem=32G,node=1),
alloc_gpus (the GPUs in it: gres/gpu and gres/gpu:TYPE, never
gres/gpumem or gres/gpuutil), and two parts of its working directory:
work_root, the first (/home, /scratch), and work_tail, the last two
(proj/run1). The whole path is not kept. squeue doesn't report these;
what sacct stored stays.
node_list is kept as Slurm writes it, in range form (n[01-04,07]). Code
that needs the nodes themselves reads it with nomad.hostlist.expand_hostlist
(or node_in); splitting on commas or matching with LIKE misses ranges.
Job numbers that come back¶
Slurm reuses job numbers: after its counter restarts (an outage that loses
its state) or wraps at MaxJobId, number 4711 is a new job. Before 1.7.42 the
new job's state, nodes and times were written onto the stored row, under the
earlier job's user, partition, name and submit time. Now, when a job arrives
whose number is stored with another submit time, and the stored job is
another job (another user or job name, or it ended more than 30 days before,
or it is still marked active but was submitted more than its time limit plus
30 days before -- a year without a limit -- as rows left RUNNING by an outage
that lost Slurm's state are), the earlier job moves aside to
NUMBER@SUBMIT-TIME (4711@2026-03-01T10:00:00; ~2 and so on if that is
taken) in every table with a job id (jobs, job_summary, job_metrics,
job_similarity, ...), and the number is the new job's. Rows in those
tables dated (submit_time or timestamp) from the new job's submission on
stay: they were written for the new job. A job moved aside while still marked
active becomes UNKNOWN. A requeued job (same user and name) stays one job,
and its submit time follows Slurm's (a requeue gets a new one). See
nomad.db.jobkeys.place.
Start times¶
squeue reports a pending job's start as Slurm's estimate. It is never
stored (before 16 Sep 2026 it was, and when the job's outcome then came from
job_metrics, which did not update the start, the estimate stayed, with a
wait to match). job_metrics now takes sacct's real start. Once per
database, finished jobs whose start disagrees with end − runtime by more than
two minutes are looked up in sacct and, when sacct's job is the same job,
stored as sacct has it (no start for a job cancelled before it ran); those
sacct no longer has are corrected to end − runtime only when their start is
impossible (after the end, or before submission). At most 1,000 are looked
up a run, so a backlog takes several runs; how far it got is kept in
config (repair.job_start_times.after), and when done, what was done is
recorded under repair.job_start_times. Rows from before 1.7.19 with an
assumed end time are left to the outcome lookups above.
Jobs from before nomad: nomad import sacct¶
nomad's job history starts when it was installed, and Slurm deletes its own
job records after PurgeJobAfter. An export taken in time, or sacct itself
for the months it still holds, fills that in, on the head node, into that
site's database:
nomad import sacct /var/tmp/sacct_all.psv.gz # what would change
nomad import sacct /var/tmp/sacct_all.psv.gz --apply
nomad import sacct --from 2025-10-01 --apply # ask sacct, a month at a time
An export is any sacct -P output with its header line (sacct -a -X -P -o ALL
-S DATE; not -n), plain or gzipped; fields are read by name. A record spread
over lines (a field with newlines) is put back together: a line that reads as
a record, or starts with one far enough to show a job id, starts a record, any
other line goes on the one before. A | inside free-text fields (submit line,
job name, working directory, comments) is put back where the other fields read
as they should. Job steps and pending array ranges are skipped.
A job already stored only gets what it lacks (account, allocation, working
directory, a partition instead of a list, a failure reason job_metrics left
at 0). Its outcome changes only when nomad had none -- stored as running,
pending or UNKNOWN, or an outcome assumed before 1.7.19 -- and the records
show it ended; a stored outcome is never mixed with another. A job the records
show running or pending, and nomad doesn't have, goes in as UNKNOWN with no
end (an export is a snapshot). Numbers Slurm gave out again are kept apart (an
older job as NUMBER@SUBMIT-TIME). Importing the same export twice changes
nothing. The import commits every half second and gives way a moment after
each commit, so the collectors writing the same database get in; with
--from, each month is committed before sacct is asked about the next, and if
sacct fails part way the months before stay written and the command says
where it stopped. The database must exist (nomad collect makes it). Without
--apply it is opened read-only and the counts say what would change, by
month submitted (not counting what the import itself would change, such as a
number used twice in the export).
Usage history (sreport)¶
slurm_usage keeps Slurm's monthly usage totals in cluster_usage, one row
per cluster, month and TRES (cpu, and gres/gpu where GPUs are
accounted): hours allocated, down, planned down, idle, planned (held for
waiting jobs; "Reserved" in older Slurm) and reported. Slurm usually keeps
these for good (PurgeUsageAfter=NONE) after the job records are gone, so
they are the long view of load. It runs where the slurm collector does,
unless [collectors.slurm_usage] enabled says otherwise.
Each run asks sreport -P -t hours -T cpu,gres/gpu cluster utilization for at
most months_per_run (24) months, newest first: this month and last month
again every refresh_hours (6) until a month has been over two days
(settled), and further back to the cluster's start or since = "YYYY-MM".
The start is where Slurm has had nothing for three months in a row and
nothing at all before them (one question about everything earlier, so a long
outage isn't taken for it); it is kept in config as
slurm_usage.nothing_before. Which months were asked, with which TRES and
when, is in cluster_usage_asked; cluster_usage holds only what Slurm
reported. A Slurm that says it doesn't account GPUs is asked for cpu alone
from then on (slurm_usage.cpu_only); any other sreport error fails the run
and the next tries again. A timeout ends the run without retries. A setting
that doesn't read is warned about and its default used. Where two hosts report
the same cluster, read one of them (the rows are the same).
Allocated, idle and planned are shares of the time nodes were up (reported −
down); down is a share of all reported time. Never add cpu and gres/gpu
rows together.
Disks¶
disk reads each of filesystems with df every run. Each reading also
fits the readings of the last forecast_window_hours (default 6) with a
straight line and stores the fill rate and when the filesystem will be full
(fill_rate_bytes_per_day, days_until_full); a disk that will be full
soon raises a disk_forecast alert (see docs/alerts.md).
NFS¶
On a host that mounts NFS, each run reads nfsiostat 5 2 and keeps the second report: what each mount did in those 5 seconds (sample_seconds in [collectors.nfs]). The first report averages everything since the share was mounted. Per mount: operations, read and write rates, and round-trip and execution times and retransmissions weighted over reads and writes — empty (NULL) for a mount with no reads or writes in the sample, rather than 0 ms. A host that mounts nothing says "no NFS mounts on this host"; turn nfs off there.
Network¶
Latency and packet loss from this host to the hosts you list, every run (ten pings, about two seconds per path):
[collectors.network_perf]
enabled = true
[[collectors.network_perf.network_tests]]
dest = "nas1"
path_type = "switch"
- Each path starts on the host that runs it. A
sourcenaming another host (by name or address) is skipped, and the run says so; list that path in the other host's config. - A path with no reply at all is stored as
unreachable. - Throughput is off unless
throughput = true. It then runs at most once perthroughput_interval(3600 s) per path — checked against the table, so it holds under cron — foriperf_durationseconds withiperf3, which needsiperf3 -srunning on the destination. Without iperf3, or when noiperf3 -sanswers, it times an SSH copy of 50 MB, which encryption may limit below what the network carries; if that fails too, no throughput is stored (never a 0).
Heavy use of login nodes (per_user)¶
On a login node or another host people share, per_user says who used it heavily, with what, and for how long. Turn it on there:
Each run (every 5 minutes from cron) reads every process and averages its CPU and I/O over the 5 minutes since the previous run, from counters that run left in per_user_state. A process seen for the first time gets its average since it started; if that is above a rule's threshold, the next reading confirms it is still busy, and the flag is dated from the process's start — a process that has run flat out for two days is flagged as such five minutes after per_user is turned on. Where systemd puts each user's logins in a slice (user-<uid>.slice, cgroup v2 or v1 with CPU accounting), it also reads each user's CPU there, which counts processes that started and ended between two readings — a parallel make, a loop starting short programs. Work of another account inside a user's slice (sudo, su), and nomad's own when its cron session runs in its account's slice, is not counted as that user's.
What it keeps:
- A sample (
per_user_sample, with the command line) only for a process above a floor:floor_cpu_percent(10),floor_memory_gb(2) orfloor_io_mb_per_s(10). Idle shells and daemons leave no row. Raw samples are deleted aftersample_retention_days(30; 0 keeps them), a batch of 100,000 per run. - Flags (
per_user_alert), kept. One row per episode — a process (or a user's processes together) above a rule's threshold without a break — and rule, written once the condition has held for the rule's duration and updated every run while it goes on:last_seen,sustained_for_seconds,occurrences(readings), and the episode's peaks. - Daily totals (
per_user_daily), kept: for each user, host and day (local date), CPU seconds used, time busy (the user's processes together above the CPU floor), peaks of CPU and memory, and I/O where readable. An interval is added to the day it ends; one longer than two hours (the collector was stopped) is left out.
The rules (a duration is a minimum: readings come every 5 minutes, so a 2-minute CPU rule fires on the first 5-minute average above it):
| Rule | Over | Condition | Severity |
|---|---|---|---|
cpu_10pct_5min |
one process | CPU ≥ 10% of a core for 5 min | informational |
cpu_50pct_2min |
one process | CPU ≥ 50% for 2 min | actionable |
memory_4gb_10min |
one process | resident memory ≥ 4 GB for 10 min | informational |
memory_16gb_2min |
one process | ≥ 16 GB for 2 min | actionable |
user_cpu_200pct_10min |
a user's processes | together ≥ 2 cores for 10 min | actionable |
user_memory_32gb_10min |
a user's processes | together ≥ 32 GB for 10 min | actionable |
io_50mbs_10min |
one process | reads + writes ≥ 50 MB/s for 10 min | informational |
A user's memory is the resident memory of their processes added up, so memory that forked workers share counts once for each.
[[collectors.per_user.rules]] tables replace them (see nomad.toml.example). System accounts (uid below min_uid), users, user_commands and processes started from parent_paths are never flagged; they are stored, marked, when above the floor.
I/O is what a process read and wrote through system calls — files on any filesystem, NFS included, pipes and sockets — so a large copy to NFS or a transfer through the host shows. Reading another user's I/O counters, and the executable path that parent_paths matches, needs root: run as another account, only that account's own processes have them. For a script run by an interpreter (python3 /usr/local/sw/backup/backup.py), parent_paths and user_commands also match the script, read from the command line, which needs no root: user_commands = [["backupuser", "backup.py"]].
If /proc is mounted with hidepid, an account other than root sees only its own processes; the run says so (nomad collectors).
nomad per-user # this host's database
nomad per-user --db ~/.local/share/nomad/combined.db --days 30 # every site, on the hub
nomad per-user --mask # user and command names replaced, for sharing
Each flagged line is a process (or a user's processes together) and a stretch of time: from when the condition began to when it was last seen, the rules it broke (! actionable, i informational), and its peaks.
Workstations over SSH¶
workstation runs its commands and probes on each machine over SSH
(BatchMode: a key, no password prompt), as the user ~/.ssh/config gives
for that host, and runs them directly on a machine listed under the
collecting host's own name. Some lab machines print a banner on every login,
interactive or not, from a shell startup file. Since 1.7.27 nomad prints a
marker before each command and reads only what follows it, so such a banner
is never read as the machine's figures. To keep the banner off
non-interactive logins anyway, wrap it in if [[ $- == *i* ]]; then ... fi.
Workstation mounts¶
Each workstation run also checks every mount on each machine (NFS, and
local filesystems other than the system's own). A small probe goes over the
same SSH connection to the machine's own python3 (3.6 or later), so
nothing is installed on the workstations. It checks all mounts at once, each
in its own thread with a 3-second limit, and records whether stat()
answered (workstation_mount_state: is_responsive, response_ms). A call
on a dead NFS server cannot be interrupted, so a thread still waiting at the
limit is abandoned, the mount recorded as not responding, and the probe ends
anyway: a dead NAS costs one limit, and the other mounts are still reported.
Since 1.7.24 it also records the size of the filesystem behind each mount,
as df shows it: total_bytes, used_bytes and avail_bytes (what users
can still write). These are empty for a mount that did not answer, and in
rows from before 1.7.24. For an NFS mount they are the server's figures for
that export, so the lab machines report their NAS's space without nomad
reaching the NAS. "Used" is the whole export's, whoever wrote it.
Exports that are datasets of one ZFS pool each report their own used space
but share the pool's free space, and each one's total is its own used space
plus that shared free space. So their totals overlap and must not be added
up. When two exports of one server are read with the same free space
(within 0.1%, at most 1 GiB) in one run, whether by the same lab machine or
by two (an export mounted on a single machine is read by no other),
nomad console roles says "free space shared with ...", leaves out each
one's total, and adds a line with their space together: their used space
summed, the free space counted once. Exports on different server addresses
are never put together, even if the same NAS answers on both.
Only the collecting host needs 1.7.24 for the sizes; the hub's combined database gains the columns at its next sync.
Storage servers¶
[collectors.storage]
enabled = true
[[collectors.storage.storage_devices]]
hostname = "nas1"
type = "zfs"
[[collectors.storage.storage_devices]]
hostname = "nfs-home"
type = "nfs"
paths = ["/export/home"]
Reached over SSH with BatchMode (a key, no password prompt) unless hostname is this host. type = "zfs" reads pool health and capacity: the space users have, from each pool's root dataset (zfs list), so RAIDZ parity isn't counted as it is in zpool list, and a boot pool (boot-pool, freenas-boot, bpool) isn't counted at all. It works on TrueNAS CORE (FreeBSD) as on Linux. Any type reads NFS exports and connected clients. A login banner is skipped, as for workstations. nomad lab add-nas adds a NAS (see docs/config.md). Without ZFS, capacity comes only from the paths listed (df) — never from the server's root disk. An unreachable server is stored as offline with its capacity unknown (NULL), not as an empty server.
[[storage_devices]] and [[network_tests]] at the top level of nomad.toml, where older examples put them, are still read, with a warning to move them.