NuFiDocs

Gateway admin

Register models, issue keys, set budgets and rate limits, and know what the gateway does not bound.

The NuFi AI Gateway is the one gate every model call passes. It holds the model list, the keys, the budgets and the rate limits, runs the security controls, and writes the trace. Its admin UI is LiteLLM's: at api.codechi.me/ui on NuFi's stack, http://localhost:4000/ui on a compose stack of your own. Sign in with the gateway master key (LITELLM_MASTER_KEY in the stack's .env).

Models

Two ways to register one.

In the configuration file, deploy/platform/litellm/config.yaml, through add-model.sh. Version-controlled, reproduced on every deploy, and the one to use for anything that stays. The file is baked into the gateway image, so the gateway serves the new model only after docker compose build litellm-proxy && docker compose up -d litellm-proxy; Add or change a model has the whole procedure.

In the admin UI, saved to the gateway's database. Right for trying a model for a day; not in version control and not seen by the script.

Either way, the app caches the gateway's model list when it starts. A new model reaches the dropdown after the app restarts (docker compose restart librechat, or a redeploy on Railway).

Every entry carries two labels: backend_type (gpu, npu, cloud) and hardware_id (gemini-cloud, mac-local, whatever names the box). Traces and the Grafana spend panels aggregate on hardware_id; a model without them exists but is invisible to reporting. add-model.sh refuses to register without them; the UI lets you leave them blank. Fill them in.

NuFi's own gateway ships with Google Gemini behind every model name it offers, including names that look like other vendors' models; the hardware_id on each entry (gemini-cloud) is the truthful record.

Keys

A key is what a person or a program presents to the gateway. LiteLLM's key form gives it an owner, a budget with a refresh period, per-minute token and request limits, a model allow-list, and optionally a team.

Members issue their own keys in NuFi Console, which creates them in the gateway with the deployment's defaults (DEFAULT_USER_BUDGET and its siblings). Issue a key here yourself when it belongs to no person: a CI job, an integration, an agent that runs under a service identity.

What the gateway bounds, and what it does not

Budgets and rate limits attach to keys. Every key a member creates is bounded per member. The app's own traffic, every conversation in the chat, goes through one key for the whole deployment, so its budget bounds the app as a whole and not any one person. Per-person figures for chat traffic come from Langfuse, which records the account id the app attaches to each request. NuFi Works keeps its own budgets per company and per agent, outside the gateway: Agent roles and budgets.

Teams

LiteLLM's teams give a set of keys a shared budget, limits and model list. The gateway on NuFi's stack uses them to decide which models a key may see; no tiers are preconfigured, and nothing attaches a new member to a team automatically. Create teams in the UI when you need a shared cap.

Watching spend

The UI's Usage page is the live view by key and model. For history, per-request detail and the account id, use Langfuse; for rates and totals over time, the spend panels in Grafana.

Routing

The shipped configuration routes by model name and shuffles between entries that share one. There are no fallbacks and no traffic splitting configured; LiteLLM can do both, and the file has a note about switching to weighted routing when a canary split is wanted, but neither is turned on today.

Guardrails

The security controls run here: deploy/platform/litellm/guardrails/policy.yaml decides what is blocked, logged or redacted, and a check script confirms the controls are wired. The gateway UI does not show them; the Grafana dashboard and the alerts do. Operating the guardrails is the operator's page; Security explains the model.