Alerts¶
NØMAÐ supports both threshold-based and predictive alerts.
Alert Types¶
Threshold Alerts¶
Trigger when metrics exceed configured limits:
[alerts.thresholds.disk]
used_percent_warning = 80
used_percent_critical = 95
[alerts.thresholds.nfs]
retrans_percent_warning = 1.0
avg_rtt_ms_critical = 100
Before 1.7.16 [alerts.thresholds] was not read: every site alerted at the
built-in values. The flat keys the example config used to show for disks
(disk_warning_percent, disk_critical_percent) are read too.
Disk fill forecasts¶
Each disk reading fits the readings of the last forecast_window_hours
([collectors.disk], default 6) with a straight line. When the filesystem is
filling, the reading stores the rate and when it will be full
(filesystems.fill_rate_bytes_per_day, days_until_full), and an alert
(disk_forecast) says so:
[WARNING] NØMAÐ spydur: Disk /home will be full in about 45 hours (filling 2.0 TB/day over the last 6 hours; 84% used)
[alerts.thresholds.disk]
full_within_hours_warning = 72
full_within_hours_critical = 24
full_within_hours_past_critical = 6 # once past used_percent_critical
[alerts.predictive]
enabled = true # false: no forecasts
The forecast fits the space taken (size less free space), so a dataset on a
shared pool is forecast as the pool fills. It needs four readings over at
least an hour and growth of at least 0.1% of the filesystem over the window;
readings of another filesystem (the path unmounted, df reading the one
beneath: a different size and usage) are left out, and a disk already full gets the threshold
alert instead. Past its critical threshold a disk has that alert, reminded
of daily; a forecast is added only when the rest will be gone within
full_within_hours_past_critical (6) hours.
days_until_full_* under [alerts.predictive] and
disk_fill_days_warning, which older configs carry for a forecast that never
ran, are not read.
Backends¶
Email¶
[mail] # shared by everything that sends mail
host = "smtp.example.edu"
port = 587
starttls = "required" # required | if-offered | off
from = "hpc@example.edu" # a sender your mail server accepts
[alerts.email]
enabled = true
recipients = ["admin@example.edu", "hpc-team@example.edu"]
Alert email goes through [mail]; see nomad.toml.example for its other
options (a certificate name, a missing intermediate, a login). Any of [mail]'s
keys set in [alerts.email] itself override it, for alerts that must use a
different server. nomad test-alerts --email checks the connection.
Each message names its cluster -- the name [clusters] gives, else the host
name -- in the subject with the problem itself, so an inbox shared by several
sites reads at a glance:
Slack¶
[alerts.slack]
enabled = true
webhook_url = "https://hooks.slack.com/services/T00/B00/xxx"
channel = "#hpc-alerts"
Webhook¶
[alerts.webhook]
enabled = true
url = "https://your-service.example.edu/alerts"
headers = { Authorization = "Bearer xxx" }
CLI¶
# View recent alerts
nomad alerts
# Unresolved only
nomad alerts --unresolved
# Test alert backends
nomad alerts test
Repeats¶
Each alert is about a condition: its source, its host, what on the host it
is about -- the disk path, the NFS mount, the GPU, the node -- and the metric. /home and
/scratch on one head node are two conditions. A condition is raised (stored
and sent):
- when it appears,
- when it gets worse (warning → critical),
- and once every
reminder_hourswhile it lasts.
What has been raised is kept in the alert_state table, so this holds across
nomad collect --once runs from cron. A condition not seen for
episode_gap_minutes has ended; if it comes back it is raised again (keep the
gap longer than the collection interval). A short dip back (critical →
warning → critical) raises nothing new; after episode_gap_minutes below its
worst, getting worse again is raised. Two metrics of one thing (an NFS
mount's latency and its retransmissions) are two conditions. If every backend
fails to send, the alert is sent again after 15 minutes, then 30, an hour and
so on up to reminder_hours (stored once). If the database cannot be written
(nomad's own database on the disk that filled), what was sent is remembered
in a small file in the local temporary directory instead, and the readings
are checked even though they could not be stored. nomad test-alerts
always sends its test.
Before 1.7.16 an alert was "the same" when its source, host and severity
matched, whatever the path, with a 15-minute cooldown: a disk at 87% for
weeks was stored and sent every 15 minutes, and a second disk's warning
waited behind it. cooldown_minutes now applies only where alerts have no
database.
nomad alerts lists the conditions active now (seen within the episode gap)
before the alerts raised.