Data flow
Follow a single chat message from the browser through NuFi and back.
A user types a message into the app and presses Enter. This is what happens, top to bottom.
1. The browser to the app
The user opened chat.nufi.me earlier and signed in; the browser holds
the session cookie. They pick a model from the dropdown, which the app
filled from the gateway's model list, and send. The app's backend receives
the message with the conversation it belongs to.
2. The app to the gateway
The app forwards the conversation to the gateway through its configured
endpoint: http://litellm-proxy:4000/v1 inside the compose stack,
https://api.codechi.me/v1 from Railway. It authenticates with the
deployment's endpoint key and attaches the user's account id to the
request.
The gateway checks the key, its budget and its rate limit, and resolves the model name to a provider entry: which provider, which credentials, which hardware label to record.
3. Before the model
The text is first rewritten into a canonical form: Unicode normalised, invisible characters stripped, lookalike letters folded, encoded spans decoded. Then two controls run:
- G1, prompt injection. A classifier and a pattern list score every part of the prompt by who wrote it. Text a person typed, and text the model or a tool produced earlier, needs both detectors to agree; text that arrived inside a retrieved document is stopped on one. A hit blocks the request.
- G2a, personal data in the prompt. Records what it found in the audit trail. It never edits the prompt: masking a value in a question makes the model answer about the placeholder.
The prompt goes to the provider as the user wrote it.
4. The provider
The gateway calls the provider behind the model name. On NuFi's own gateway that is Google Gemini for every shipped name; on yours, whatever you registered, including a model server on your own hardware. Tokens stream back as they are generated.
5. After the model
Three controls read the answer:
- G2b, personal data in the answer. Emails, phone numbers, card and bank identifiers, IP addresses and Korean registration numbers are hidden, unless the answer is grounded in a document the asker already has access to and the key is allowed to say so.
- G3, system prompt echo. An answer that repeats the system prompt is blocked.
- G4, exfiltration. Tracking images and script URLs the model was talked into including are removed; the rest of the answer is delivered.
Streaming survives this: the gateway forwards tokens as they arrive and holds back only what a control has to see whole. The app relays the stream to the browser, and the reply appears word by word.
6. The record
While the reply streams, the gateway writes one trace to Langfuse: the account id the app sent, the model name, the hardware label of the entry that served it, tokens in and out, latency, cost. Any guardrail decision writes one audit event with a reference code, the control and the verdict, and never the text that triggered it.
Prometheus scrapes the gateway continuously: request rate, error rate, latency, and per-control guardrail decisions. Grafana shows them; Alertmanager routes the nine alerts, to Slack once a webhook is configured.
Where the identity goes
The app knows the user from the cookie. The gateway sees the deployment's endpoint key plus the account id the app attached, so it can name the user in the trace even though the key is shared; budgets and rate limits are on the key, so on chat traffic they cover the deployment, not the person. A key the same person created in the console is theirs alone: its spend, its limits and its traces are attributed to them. "Everything this user did yesterday" is one search in Langfuse either way.
When something fails
| What happened | What the user sees | What is recorded |
|---|---|---|
| the provider did not answer | an error in the conversation | a trace marked failed, with the provider error |
| a guardrail blocked the request or the answer | the gateway's refusal, which names the reason and carries a reference code | an audit event; the nufi_guardrail_decisions_total counter |
| the key is over budget | the gateway's budget error | the trace, and the key's spend in the console |
| the key is rate-limited | a request to slow down | the trace |
The refusal is machine-readable too: a response of type
nufi_guardrail_blocked with a risk code such as LLM01_INJECTION, so a
program calling the gateway can tell policy from malfunction. Error rate
and guardrail decisions each have a panel on the LiteLLM Overview
dashboard.