FAQ
Questions operators ask before and after deploying.
Does NuFi host AI models?
No. The gateway routes to whatever you point it at: a model server of your own (vLLM, TGI, Ollama), or OpenAI, Anthropic, Google, Together, Groq, any provider with an OpenAI-compatible API. Bring your own model.
Can I run only the app?
Yes. deploy/railway is the app on its own: the chat, a MongoDB, and a
small RAG service for file uploads, pointed at any OpenAI-compatible
endpoint through BACKEND_BASE_URL and BACKEND_API_KEY. No gateway, no
guardrails, no observability. It is what NuFi's Railway staging runs, and
it works on one host with ./bootstrap.sh.
Can I run several app instances behind a load balancer?
Yes. The app supports Redis-backed resumable streams for multi-instance
deployments; set its REDIS_* variables and put a load balancer in front.
The compose stack runs one instance.
How do users get accounts?
With ALLOW_REGISTRATION=true anyone who reaches the URL can register; the
first account becomes ADMIN. For production, set it to false and create
users in the admin panel, or connect your identity provider: the app
supports OpenID Connect, and the admin panel signs in against the same
accounts.
Which models should I offer?
Whatever you have. The gateway's shipped configuration serves gemini,
claude-sonnet-4-5, claude-haiku-4-5, gpt-5 and gpt-5-mini given
the provider keys, and a local Ollama model such as qwen2.5:7b for
laptops. Expose several; per-user and per-team allow-lists decide who sees
which.
How do I limit what someone spends?
Two layers, both in the gateway: a budget on the user, and a budget on
each key the user issues in the console. The console shows both to the
user. Defaults for new users come from DEFAULT_USER_BUDGET and its
siblings, on the console.
Can I show certain models only to certain people?
Yes. In the gateway, per-key and per-team allowed_models decide what a
call may use (policy). In the app, the admin panel's per-role and
per-group overrides decide what appears in the dropdown (presentation).
Use both: present only what you have allowed.
How do I add a provider?
./scripts/add-model.sh with the base URL, the key's variable name and the
model id, then rebuild the gateway image.
Add or change a model.
Where do conversations live?
In the app's MongoDB. Only the prompt goes to the provider. A user deletes a conversation in the app; an admin deletes a user, and their conversations and files, in the admin panel.
Is every prompt logged?
Yes: the gateway sends each request and reply to Langfuse, which is how
"what did this user see?" gets answered and how spend is attributed. A
temporary chat in the app is not persisted by the app, but still passes
through the gateway. To stop the recording, remove the LANGFUSE_*
variables from the gateway's environment; you lose observability and the
console's usage figures.
Why does the trace store grow so fast?
It keeps every request in full. Plan on roughly 50 MB per thousand requests and see Backup and restore for retention.
How do I rotate the gateway master key?
Generate a new one, set LITELLM_MASTER_KEY in .env, docker compose up -d (the gateway and the console read it), and update anything else
that calls the gateway with it. Keys users issued keep working; only the
master key changed.
Does NuFi work behind Cloudflare Access or a corporate proxy?
Yes. It is HTTPS. Behind Cloudflare Access, the Access sign-in comes first, then the app's own.
Where do I report a problem?
The repository, dudaji-vn/nufi-app, holds every component; open an issue
there and name the surface (app, console, admin panel, Studio, Works,
stack).