Langfuse
Every request as a trace: who, which model, what it cost, and where to look when a reply was wrong.
Langfuse records every request the gateway handled: the prompt and the reply, the account id the app attached, the model name and the hardware that served it, tokens, latency, cost. It is what you open when someone says "the model got this wrong" or "where did the budget go".
NuFi's own copy is at langfuse.codechi.me; on a compose stack of your
own it is host port 3000.
Sign in
The stack seeds an organisation (npuops), a project (npuops-default)
and one admin at first start, from LANGFUSE_INIT_USER_EMAIL and
LANGFUSE_INIT_USER_PASSWORD in .env. Sign in as that admin, then
invite others from Settings → Members. Langfuse's accounts are its own,
separate from the app's, which keeps the set of people who can read other
people's prompts small.
What a trace is
A trace is one request. Inside it, observations are the calls it
made, usually one model call. Each carries its own timing, model, input,
output and cost. NuFi adds two fields to every trace, hardware_id and
backend_type, from the model entry that served it.
Find one user's requests
- Traces, filter by user id and a time range. The id is the app's account id, which the app sends with each request.
- Open the trace; the observation shows the system prompt, the messages, the reply, and the model that answered.
A guardrail block does not appear as a reply here: the gateway refused before or after the model, and the refusal is an audit event with a reference code. The person who was refused can quote that code; the gateway's log holds the event.
Compare, score, keep
Langfuse's datasets, experiments and scores work as its documentation describes; NuFi does not change them. Scores you add to traces show up in its dashboards.
Cost
Langfuse's dashboards break spend down by model and by user. Because NuFi's gateway answers several model names with the same provider, the per-model figure follows the name asked for; the split by actual hardware is the Spend rate by model panel in Grafana, which reads the gateway's own counters.
Retention
Nothing is deleted by default. The traces live in ClickHouse and MinIO on the compose stack, and those are the volumes that grow; Backup and restore and Infra sizing say how fast. Ask the operator for a retention policy when the disk alert first fires, not after.