NuFiDocs

Grafana

The one dashboard that ships, the nine alerts behind it, and how to read them.

Where Langfuse shows what the model said, Grafana shows how the platform held up while it said it: request rate, errors, latency, spend, and whether the security controls were running.

NuFi's own copy is at grafana.codechi.me; on a compose stack it is host port 3030. Sign in with GRAFANA_ADMIN_USER and GRAFANA_ADMIN_PASSWORD from the stack's .env; sign-up is off.

The dashboard

One dashboard is provisioned, LiteLLM Overview. Its panels:

PanelWhat it shows
Requests / sec (5m)request rate at the gateway
Error rate (5m)4xx and 5xx as a share of requests
p95 latency (5m), Latency percentiles (5m)p50, p95 and p99
Total spend (cumulative), Spend rate by model (5m)cost, by the model name requested
RPS by requested modelwhich models are being used
Token throughput (input vs output)tokens per second
Guardrail decisions by control, Decisions in this windowwhat G1 to G4 did
Controls enforcing, Guardrail degraded (failing open)whether each control is on and healthy
Guardrail latency p95 by controlwhat the checks cost in time
Postgres connections by database, Redis memorythe two stores the gateway depends on

There is no per-user panel; Langfuse is where per-person figures live. The dashboard is provisioned from a file in git, so edits made in the UI are not kept across restarts: Save as a copy to work on one.

The alerts

Nine rules, in two files in deploy/platform/monitoring/rules/:

AlertFires whenSeverity
LiteLLMDownthe gateway is not scraped for a minutecritical
LiteLLMHighErrorRateerrors above 5 % for five minuteswarning
LiteLLMHighLatencyP95p95 above ten seconds for five minuteswarning
GuardrailMetricsAbsentthe guardrail pipeline is not loadedcritical
MandatoryControlMissinga mandatory control is absentcritical
MandatoryControlStoppedEnforcinga mandatory control fell back to loggingcritical
GuardrailDegradeda control is failing opencritical
GuardrailSilentthe injection control stopped scanning while traffic continueswarning
GuardrailBlockRateHighmore than 5 % of requests blockedwarning

The critical ones already route to Slack; they reach a channel once the operator has put the webhook in place. Until then they fire and nobody is told. Monitoring and alerts is the operator's page.

Retention

Fifteen days. Prometheus deletes older data; nothing is kept at lower resolution. Langfuse keeps the per-request history.

Grafana or Langfuse?

Grafana first, when you noticed something: is the platform up, is it slow, is there an error storm, did a control stop enforcing. Then Langfuse, for the request that shows what happened.