Skip to content

Latest commit

 

History

History
187 lines (147 loc) · 7.71 KB

File metadata and controls

187 lines (147 loc) · 7.71 KB

Server status lights

These scripts drive the status table at the top of Computing Hardware.

Why there are two probers

The three servers are not reachable from the same place:

Host Public DNS How it is checked
raven.fish.washington.edu no TCP 8787 (RStudio Server), falling back to 22 — UW network only
gannet.fish.washington.edu yes HTTPS GET /
klone.hyak.uw.edu yes TCP 22 — there is no web port to probe

Raven is not in public DNS, so nothing outside the UW network can see it. Klone has no HTTP port at all, and browsers refuse to connect to port 22, so no purely client-side check can work either.

So there are two probers writing to the same place:

  • internal — cron on a machine inside the UW network. The only source that can see raven. This is the primary.
  • external — the server-status GitHub Action. Sees gannet and klone from the public internet, so a dead internal prober does not take out all three lights at once.

Each writes status/<profile>.json to the orphan server-status branch. The page fetches both from raw.githubusercontent.com (which sends access-control-allow-origin: *, so there is no CORS problem) and takes the most recent fresh reading per host.

Raven's disk/CPU snapshot

Raven also drops a daily df + CPU-load snapshot at https://gannet.fish.washington.edu/v1_web/owlshell/bu-github/ghr.log, generated by a cron on raven around 7am and synced to gannet's public web root. Because that URL is on gannet (public), check_servers.py fetches and parses it (fetch_raven_stats) under either profile, unlike the live TCP/8787 check, which only the internal profile can reach.

The parsed result is written as a top-level raven_stats field (not nested under hosts), separate from the up/down lights:

"raven_stats": {
  "generated": "2026-08-18T14:01:02Z",
  "cpu_percent": 0.58,
  "disks": [
    {
      "filesystem": "/dev/sdd1",
      "mount": "/home/shared/8TB_HDD_02",
      "size": "7.3T",
      "used": "6.3T",
      "available": "602G",
      "use_percent": 92
    }
  ]
}

It has to stay out of hosts because hosts entries are merged by picking the freshest whole document, and the external prober can run more often than the internal cron. If raven_stats lived under hosts.raven, an external run with no real port-check result would periodically overwrite a correct, live up/down reading. Instead, docs/javascripts/server-status.js merges raven_stats on its own generated timestamp.

It renders into its own "Raven Disk / CPU" table on the page ([data-raven-stats-base], one row per drive, fullest first) rather than a column in the status table — the snapshot is daily rather than live, and covers only raven, so mixing it into the live up/down table misrepresented both. CPU load and the snapshot time go in the .ss-raven-meta line above the table, flagged stale if the snapshot is older than 30 hours. Drives at 90%+ full are marked critical and 75%+ warn, in the text as well as the bar color.

Gannet's daily health report

Gannet runs gannet_health.sh once a day and writes the result to https://gannet.fish.washington.edu/v1_web/owlshell/latest.txt, plus a dated copy, gannet_health_YYYY-MM-DD.txt, in the same directory. The report covers uptime and load, df and inode use, memory, RAID (/proc/mdstat), services and failed systemd units, kernel errors, and a summary of alerts.

Gannet does not send CORS headers, so the page cannot fetch the report itself. Instead check_servers.py fetches and parses it (fetch_gannet_stats) under either profile and writes a top-level gannet_stats field. It is kept out of hosts for the same reason as raven_stats. The field holds the parsed sections, the raw text (raw), and a history list with one small entry per dated report for the last 30 days (load, per-mount use %, alert count). That list is built by reading the directory index and fetching each dated file.

docs/javascripts/gannet-health.js merges gannet_stats on its own generated timestamp and renders it in three places:

  • the Daily health cell for gannet in the status table ([data-health="gannet"]),
  • the Gannet Health section on Computing Hardware ([data-gannet-summary]: alerts and the disk table), and
  • the Gannet Dashboard page (docs/Gannet-Dashboard.md, [data-gannet-dashboard]), with the full report and history charts.

A report older than 30 hours is flagged stale. If the report format changes, parse_gannet_report treats every section as optional, so a partial parse still renders.

The branch is rewritten as a single root commit on every run. At one check every 10 minutes an append-only branch would add roughly 50,000 commits a year to a repo that everyone clones.

Files

  • check_servers.py — runs the probes, prints or writes the JSON. Stdlib only.
  • publish_status.sh — runs the checker and pushes the result to the server-status branch. Used by both the cron job and the Action.

Check without publishing anything:

./scripts/check_servers.py --profile internal

Off the UW network, raven will report DNS lookup failed — that is expected, and is exactly why the internal prober has to run inside.

Setting up the in-network cron

Run this on a machine inside the UW network that is up continuously. Gannet is the natural choice.

1. Give the machine push access. Generate a deploy key on that host:

ssh-keygen -t ed25519 -f ~/.ssh/robertslab_status -C "roberts-lab status bot" -N ""

Add the public key (~/.ssh/robertslab_status.pub) at https://github.com/RobertsLab/resources/settings/keys as a deploy key with write access checked. A deploy key is scoped to this one repo, so it cannot be used to touch anything else in the org.

Then point git at it in ~/.ssh/config:

Host github-robertslab
  HostName github.com
  User git
  IdentityFile ~/.ssh/robertslab_status
  IdentitiesOnly yes

2. Clone the repo somewhere on that host, e.g. ~/robertslab-resources.

3. Add the cron entry with crontab -e:

*/10 * * * * ~/robertslab-resources/scripts/publish_status.sh --profile internal --repo github-robertslab:RobertsLab/resources.git >> ~/status-cron.log 2>&1

publish_status.sh keeps its own scratch clone under ~/.cache/robertslab-server-status, so it will not disturb the checkout it is run from.

4. Confirm that status/internal.json appears on the server-status branch and that the lights on the handbook page go green within a few minutes.

Notes and limits

  • A green light means the port answered. It says nothing about Slurm health or whether jobs are running. Raven's disk space and CPU load, and gannet's daily health report, are covered separately as described above; klone has no equivalent.
  • raw.githubusercontent.com caches for about 5 minutes, so the page can lag the actual check by that much on top of the check interval.
  • GitHub's scheduled workflows are best-effort: delayed under load, minimum 5-minute interval, and disabled automatically after 60 days without repo activity. That is why the in-network cron is primary and the Action is only a backstop. Measured on this repo, the Action asks for */15 but lands every 30-71 minutes (median 46). The 90-minute stale threshold in docs/javascripts/server-status.js is sized around that; if you tighten it, make sure the in-network cron is actually running first, or the lights will spend most of each cycle showing unknown.
  • To add a host: add a probe in check_servers.py and a <tr data-host="..."> row in docs/Computing-Hardware.md. The JavaScript matches the two by name and needs no change.