User Tools

Site Tools


wiki:ai:slurm_servicenow
DraftReady for approval 2026/09/11 20:08 by bgourley | Approver: @ai-us-principals

Slurm Job Failure → ServiceNow Incident Alerts

This documents how to automatically create a ServiceNow incident the moment a Slurm job fails, using BCM's Slurm epilog chain.

Prerequisite: a working ServiceNow OAuth2 client-credentials setup (client ID/secret, incident-table write access) — this reuses the same auth pattern as any other BCM-to-ServiceNow alert.


Why this approach

BCM's built-in monitoring metrics (JobsRunning, FailedJobs, job_metadata_*) are aggregate queue-health numbers, not per-job failure events — there's no stock BCM measurable/trigger for “this specific job failed.” NVIDIA's Autonomous Job Recovery (a separate Mission Control add-on) does this natively, but on a basic BCM license, hooking directly into Slurm's own job-completion mechanism is the practical path.

Slurm runs an epilog script automatically the instant any job reaches a terminal state (success or failure) — lower latency and more reliable than polling sacct on a schedule. BCM already has an epilog mechanism wired in; you don't hand-edit EpilogSlurmctld= into slurm.conf (BCM manages and can overwrite that file) — instead you drop a script into a specific directory BCM already chains together.


The script

sudo nano /cm/local/apps/cmd/scripts/actions/slurm-epilog-servicenow.sh
#!/bin/bash
# Runs automatically via BCM's Slurm epilog chain on every job completion.
# Only creates a ServiceNow incident if the job's final state indicates failure.
 
export SLURM_CONF=/cm/shared/apps/slurm/etc/slurm/slurm.conf
export PATH=$PATH:/cm/local/apps/slurm/25.05/bin
 
SN_INSTANCE="https://YOUR_INSTANCE.service-now.com"
SN_CLIENT_ID="your_client_id"
SN_CLIENT_SECRET="your_client_secret"
 
LOG=/var/log/cmd-servicenow-slurm-status.log
JOBID="${SLURM_JOB_ID}"
 
if [ -z "$JOBID" ]; then
  echo "$(date) - No SLURM_JOB_ID set, exiting" >> "$LOG"
  exit 0
fi
 
# sacct accounting records can lag slightly behind the epilog firing —
# retry briefly rather than failing silently on a race condition.
for i in 1 2 3 4 5; do
  JOBINFO=$(/cm/local/apps/slurm/25.05/bin/sacct -j "$JOBID" --noheader --parsable2 \
    --format=JobID,JobName,User,Partition,State,ExitCode,NodeList,End 2>>"$LOG" \
    | grep -E "^${JOBID}\|")
  [ -n "$JOBINFO" ] && break
  sleep 2
done
 
if [ -z "$JOBINFO" ]; then
  echo "$(date) - Could not find sacct record for job ${JOBID} after retries" >> "$LOG"
  exit 0
fi
 
IFS='|' read -r JOBID JOBNAME USER PARTITION STATE EXITCODE NODELIST END <<< "$JOBINFO"
 
# Only alert on real failure states — not COMPLETED, CANCELLED (user-initiated), etc.
case "$STATE" in
  FAILED|TIMEOUT|NODE_FAIL|OUT_OF_MEMORY)
    ;;
  *)
    exit 0
    ;;
esac
 
get_token() {
  curl -s -X POST "${SN_INSTANCE}/oauth_token.do" \
    -H 'Content-Type: application/x-www-form-urlencoded' \
    --data-urlencode 'grant_type=client_credentials' \
    --data-urlencode "client_id=${SN_CLIENT_ID}" \
    --data-urlencode "client_secret=${SN_CLIENT_SECRET}" | jq -r '.access_token'
}
 
ACCESS_TOKEN=$(get_token)
 
if [ -z "$ACCESS_TOKEN" ] || [ "$ACCESS_TOKEN" = "null" ]; then
  echo "$(date) - Failed to get ServiceNow token for job ${JOBID}" >> "$LOG"
  exit 0
fi
 
curl -s -w "\nHTTP_STATUS:%{http_code}\n" \
  -H "Authorization: Bearer ${ACCESS_TOKEN}" \
  -H "Content-Type: application/json" \
  -X POST "${SN_INSTANCE}/api/now/table/incident" \
  -d "{
    \"short_description\": \"Slurm job ${JOBID} (${JOBNAME}) failed: ${STATE}\",
    \"description\": \"User: ${USER}, Partition: ${PARTITION}, State: ${STATE}, ExitCode: ${EXITCODE}, Nodes: ${NODELIST}, Ended: ${END}\",
    \"category\": \"Infrastructure\",
    \"urgency\": \"2\",
    \"impact\": \"2\"
  }" >> "$LOG" 2>&1
sudo chown root:root /cm/local/apps/cmd/scripts/actions/slurm-epilog-servicenow.sh
sudo chmod 700 /cm/local/apps/cmd/scripts/actions/slurm-epilog-servicenow.sh

Why the state filter matters: the case block is what separates a real failure from a normal completion. COMPLETED and CANCELLED fall through to exit 0 before any ServiceNow call is made — only FAILED, TIMEOUT, NODE_FAIL, and OUT_OF_MEMORY reach the incident-creation code.

Why the script sets SLURM_CONF/PATH explicitly: when slurmctld/slurmd invoke this script, it runs in a minimal daemon environment — it does not inherit your interactive shell's PATH or SLURM_CONF. Without these two lines, sacct fails silently (or errors with a DNS-SRV config lookup failure) and the script can't tell “no failed jobs” apart from “sacct itself couldn't run.” This cost significant debugging time — don't skip it.


Wiring it into BCM's Slurm epilog chain

BCM does not use a raw EpilogSlurmctld= line for this — it manages slurm.conf itself and generates a chain of numbered symlinks in a dedicated directory, each run in order by a generic epilog wrapper.

Enable post-job processing (off by default):

% wlm
% use slurm
% set enablepostjob yes
% commit

Check the epilog directory — BCM auto-populates its own validation script here once post-job is enabled:

ls -la /cm/local/apps/slurm/var/epilogs/

Expect to see something like:

01-wlm-post-job-validation -> /cm/local/apps/cmd/scripts/wlm-post-job-validation

Add your script with the next number in sequence (so it runs after BCM's own validation):

sudo cp /cm/local/apps/cmd/scripts/actions/slurm-epilog-servicenow.sh /cm/local/apps/slurm/var/epilogs/02-servicenow-incident
sudo chmod 700 /cm/local/apps/slurm/var/epilogs/02-servicenow-incident
sudo chown root:root /cm/local/apps/slurm/var/epilogs/02-servicenow-incident

The equivalent prejob directory (for reference/symmetry) is /cm/local/apps/slurm/var/prologs/, with 01-wlm-pre-job-validation as BCM's own script.


Debugging tools specific to this chain

BCM logs every prolog/epilog run, per job, regardless of your own script's logging:

sudo cat /var/log/slurm-prologs.log
sudo cat /var/log/slurm-epilogs.log

These show each numbered script's exit code per job — extremely useful for telling apart “my script didn't run,” “my script ran but errored,” and “a script earlier in the chain blocked the rest.”


Testing

# Should create an incident
sbatch --wrap="exit 1"
 
# Should NOT create an incident
sbatch --wrap="exit 0"

For each, confirm the actual Slurm-recorded state (don't assume the wrapped exit code is what Slurm reports):

sacct -j <jobid> --format=JobID,State,ExitCode

Then check:

sudo cat /var/log/cmd-servicenow-slurm-status.log

and the ServiceNow Incident table.


Known gotchas hit during setup

  • rogueprocess health check blocks job launches on a combined head/compute node. If BCM's default Failing health checks trigger is active and rogueprocess is selected as a prejob check, your own interactive SSH session gets flagged as a “rogue” process every time, draining the node before any job can start. On a node that's also used interactively (as opposed to a dedicated compute node), disable this check:
% monitoring measurable
% use rogueprocess
% set disabled yes
% commit
This only matters because our test setup used one node as both head node and compute client — on a real multi-node cluster with dedicated compute nodes, this wouldn't be an issue.
  • A drained node needs an explicit resume, and scontrol needs both SLURM_CONF and sudo's stripped PATH/env accounted for:
sudo SLURM_CONF=/cm/shared/apps/slurm/etc/slurm/slurm.conf /cm/local/apps/slurm/25.05/bin/scontrol update NodeName=<node> State=RESUME
  • Never set a cmsh list-type field to an empty string to “clear” it — e.g. set grestypes “” wrote a literal GresTypes= (no value) into slurm.conf, which Slurm's parser rejects outright as a fatal error, taking down both slurmctld and slurmd cluster-wide. Look for a dedicated clear <field> command instead, and if a bad value already broke the config, fix slurm.conf directly first to restore service, then correct the source in cmsh afterward.
  • A GRES type (e.g. gpu) can be configured in Slurm/BCM even on a node with no actual GPU, if the wizard's defaults or an earlier config pass left it set. This is worth checking (wlm[slurm]; showGres types) on any node that unexpectedly fails jobs immediately regardless of content.
  • An unresolved issue from this session, left open: jobs were still failing immediately with JobState=FAILED Reason=RaisedSignal:53(Real-time_signal_19) on every job (including trivial exit 0 ones) even after removing the phantom GPU GRES config. The cgroup v2 device-constraint plugin (Constrain devices: yes from the cgroups setup screen) is the next suspect, untested at the time of writing. If you hit this signature, check cgroup.conf, dmesg around the job's timestamp, and try toggling constraindevices off as a diagnostic before assuming it's something else.
wiki/ai/slurm_servicenow.txt · Last modified: by bgourley