Monitoring and alerts
What Prometheus watches, the nine alerts that ship with the stack, and how to make them reach Slack.
Three containers in deploy/platform: Prometheus scrapes, Grafana draws,
Alertmanager routes. All three are provisioned from files under
monitoring/, so what you see running is what is in git.
Prometheus
Four scrape jobs, in monitoring/prometheus.yml:
| Job | Target | What it gives you |
|---|---|---|
litellm | the gateway's /metrics | request rate, errors, latency, and the guardrail decision counters |
postgres | postgres-exporter | connections, database sizes |
redis | redis-exporter | memory, hit rate |
prometheus | itself |
Retention is 15 days (--storage.tsdb.retention.time=15d in the compose
file). The UI is on host port 9090, meant for an SSH tunnel rather than a
public hostname.
Rules reload without a restart:
$EDITOR monitoring/rules/litellm.rules.yml
curl -X POST http://localhost:9090/-/reloadThe alerts that ship
monitoring/rules/litellm.rules.yml watches the gateway:
| Alert | Fires when | Severity |
|---|---|---|
LiteLLMDown | up{job="litellm"} == 0 for 1 minute | critical |
LiteLLMHighErrorRate | 4xx and 5xx are more than 5 % of requests for 5 minutes | warning |
LiteLLMHighLatencyP95 | p95 latency above 10 seconds for 5 minutes | warning |
monitoring/rules/guardrails.rules.yml watches whether the security
controls are actually running, which no request metric can tell you:
| Alert | Fires when | Severity |
|---|---|---|
GuardrailMetricsAbsent | the guardrail pipeline is not loaded, for 5 minutes | critical |
MandatoryControlMissing | a control the policy marks mandatory is absent, for 5 minutes | critical |
MandatoryControlStoppedEnforcing | a mandatory control has dropped to logging only, for 5 minutes | critical |
GuardrailDegraded | nufi_guardrail_degraded > 0 for 2 minutes: a control is failing open | critical |
GuardrailSilent | the injection control has stopped scanning while traffic continues, for 15 minutes | warning |
GuardrailBlockRateHigh | guardrails are blocking more than 5 % of requests, for 10 minutes | warning |
A critical guardrail alert means the gateway is passing traffic without the checks it promises. Treat it as an incident, not a warning.
Routing to Slack
Out of the box the default receiver is noop, but the route already sends
every severity: critical alert to the slack receiver, which reads its
webhook from a file that is gitignored. Wire it up:
cp monitoring/secrets/slack-webhook.example monitoring/secrets/slack-webhook
$EDITOR monitoring/secrets/slack-webhook # the incoming-webhook URL, one line, no quotes
$EDITOR monitoring/alertmanager.yml # the channel (#npuops-alerts), and route.receiver if warnings should go too
docker compose restart alertmanagerAlerts are grouped by name and severity, repeat every 12 hours while firing, and send a resolved message when they clear.
Grafana
Host port 3030 (3000 is Langfuse). Sign in with GRAFANA_ADMIN_USER and
GRAFANA_ADMIN_PASSWORD from .env; sign-up is off in the compose file.
One dashboard is provisioned, LiteLLM Overview
(monitoring/grafana/dashboards/litellm-overview.json): requests per
second, error rate, p95 latency and the rest of the gateway's request
metrics. Drop another dashboard's JSON into the same directory and restart
Grafana to add it. Dashboards edited in the UI are not written back to the
file; export them if you want to keep them.
Without a browser
curl -s http://localhost:9090/api/v1/alerts | jq # what is firing now
curl -s http://localhost:9090/api/v1/targets | jq # every scrape target and whether it is upWhat to add
Nothing in the stack watches disk, and disk is what runs out first: ClickHouse and MinIO grow with every trace. Add a node exporter on the host and alert on the data volume at 70 %. If you run on Kubernetes or a cloud host, use its own disk alerts instead. See Infra sizing for the growth rates and Backup and restore for what is worth keeping.
Agent egress
Keeping the agent products' model traffic on the gateway — what is measured today, what is only configured, and how to prove which is which.
Operating the guardrails
The one file that decides what the gateway blocks, what needs a rebuild, and the two checks that catch a control that is silently absent.