Grafana
The one dashboard that ships, the nine alerts behind it, and how to read them.
Where Langfuse shows what the model said, Grafana shows how the platform held up while it said it: request rate, errors, latency, spend, and whether the security controls were running.
NuFi's own copy is at grafana.codechi.me; on a compose stack it is host
port 3030. Sign in with GRAFANA_ADMIN_USER and GRAFANA_ADMIN_PASSWORD
from the stack's .env; sign-up is off.
The dashboard
One dashboard is provisioned, LiteLLM Overview. Its panels:
| Panel | What it shows |
|---|---|
| Requests / sec (5m) | request rate at the gateway |
| Error rate (5m) | 4xx and 5xx as a share of requests |
| p95 latency (5m), Latency percentiles (5m) | p50, p95 and p99 |
| Total spend (cumulative), Spend rate by model (5m) | cost, by the model name requested |
| RPS by requested model | which models are being used |
| Token throughput (input vs output) | tokens per second |
| Guardrail decisions by control, Decisions in this window | what G1 to G4 did |
| Controls enforcing, Guardrail degraded (failing open) | whether each control is on and healthy |
| Guardrail latency p95 by control | what the checks cost in time |
| Postgres connections by database, Redis memory | the two stores the gateway depends on |
There is no per-user panel; Langfuse is where per-person figures live. The dashboard is provisioned from a file in git, so edits made in the UI are not kept across restarts: Save as a copy to work on one.
The alerts
Nine rules, in two files in deploy/platform/monitoring/rules/:
| Alert | Fires when | Severity |
|---|---|---|
LiteLLMDown | the gateway is not scraped for a minute | critical |
LiteLLMHighErrorRate | errors above 5 % for five minutes | warning |
LiteLLMHighLatencyP95 | p95 above ten seconds for five minutes | warning |
GuardrailMetricsAbsent | the guardrail pipeline is not loaded | critical |
MandatoryControlMissing | a mandatory control is absent | critical |
MandatoryControlStoppedEnforcing | a mandatory control fell back to logging | critical |
GuardrailDegraded | a control is failing open | critical |
GuardrailSilent | the injection control stopped scanning while traffic continues | warning |
GuardrailBlockRateHigh | more than 5 % of requests blocked | warning |
The critical ones already route to Slack; they reach a channel once the operator has put the webhook in place. Until then they fire and nobody is told. Monitoring and alerts is the operator's page.
Retention
Fifteen days. Prometheus deletes older data; nothing is kept at lower resolution. Langfuse keeps the per-request history.
Grafana or Langfuse?
Grafana first, when you noticed something: is the platform up, is it slow, is there an error storm, did a control stop enforcing. Then Langfuse, for the request that shows what happened.