Gateway admin
Register models, issue keys, set budgets and rate limits, and know what the gateway does not bound.
The NuFi AI Gateway is the one gate every model call passes. It holds
the model list, the keys, the budgets and the rate limits, runs the
security controls, and writes the trace. Its admin UI is LiteLLM's: at
api.codechi.me/ui on NuFi's stack, http://localhost:4000/ui on a
compose stack of your own. Sign in with the gateway master key
(LITELLM_MASTER_KEY in the stack's .env).
Models
Two ways to register one.
In the configuration file, deploy/platform/litellm/config.yaml,
through add-model.sh. Version-controlled, reproduced on every deploy,
and the one to use for anything that stays. The file is baked into the
gateway image, so the gateway serves the new model only after
docker compose build litellm-proxy && docker compose up -d litellm-proxy;
Add or change a model has the whole procedure.
In the admin UI, saved to the gateway's database. Right for trying a model for a day; not in version control and not seen by the script.
Either way, the app caches the gateway's model list when it starts. A new
model reaches the dropdown after the app restarts (docker compose restart librechat, or a redeploy on Railway).
Every entry carries two labels: backend_type (gpu, npu, cloud) and
hardware_id (gemini-cloud, mac-local, whatever names the box). Traces
and the Grafana spend panels aggregate on hardware_id; a model without
them exists but is invisible to reporting. add-model.sh refuses to
register without them; the UI lets you leave them blank. Fill them in.
NuFi's own gateway ships with Google Gemini behind every model name it
offers, including names that look like other vendors' models; the
hardware_id on each entry (gemini-cloud) is the truthful record.
Keys
A key is what a person or a program presents to the gateway. LiteLLM's key form gives it an owner, a budget with a refresh period, per-minute token and request limits, a model allow-list, and optionally a team.
Members issue their own keys in NuFi Console,
which creates them in the gateway with the deployment's defaults
(DEFAULT_USER_BUDGET and its siblings). Issue a key here yourself when it
belongs to no person: a CI job, an integration, an agent that runs under a
service identity.
What the gateway bounds, and what it does not
Budgets and rate limits attach to keys. Every key a member creates is bounded per member. The app's own traffic, every conversation in the chat, goes through one key for the whole deployment, so its budget bounds the app as a whole and not any one person. Per-person figures for chat traffic come from Langfuse, which records the account id the app attaches to each request. NuFi Works keeps its own budgets per company and per agent, outside the gateway: Agent roles and budgets.
Teams
LiteLLM's teams give a set of keys a shared budget, limits and model list. The gateway on NuFi's stack uses them to decide which models a key may see; no tiers are preconfigured, and nothing attaches a new member to a team automatically. Create teams in the UI when you need a shared cap.
Watching spend
The UI's Usage page is the live view by key and model. For history, per-request detail and the account id, use Langfuse; for rates and totals over time, the spend panels in Grafana.
Routing
The shipped configuration routes by model name and shuffles between entries that share one. There are no fallbacks and no traffic splitting configured; LiteLLM can do both, and the file has a note about switching to weighted routing when a canary split is wanted, but neither is turned on today.
Guardrails
The security controls run here: deploy/platform/litellm/guardrails/policy.yaml
decides what is blocked, logged or redacted, and a check script confirms
the controls are wired. The gateway UI does not show them; the Grafana
dashboard and the alerts do. Operating the guardrails
is the operator's page; Security explains the
model.