NuFiDocs

Troubleshooting

The failures operators hit on this stack, each with the command that tells you which one you have.

Three commands cover most incidents. Run them from deploy/platform:

docker compose ps                                   # anything not (healthy)?
docker compose logs --tail 200 <service>            # the last lines of the one that is not
curl -s http://localhost:9090/api/v1/alerts | jq    # what Prometheus thinks is wrong

The gateway restarts in a loop

docker compose logs --tail 50 litellm-proxy
The log ends withCauseDo
unknown threshold key(s) [...]the locally built gateway image is older than the checkout; docker compose up never rebuilds itdocker compose build litellm-proxy nufi-scanner && docker compose up -d
a Gemini or GEMINI_API_KEY errorlitellm/config.yaml ships two Gemini entries and LiteLLM refuses to start with an empty keyset the key in .env, or remove both entries and rebuild
a model entry named in the errora typo in litellm/config.yaml, or a ${VAR} it references is unsetfix the entry, rebuild
connection refused to PostgresPostgres was still startingwait 30 seconds; if it persists, docker compose logs postgres

The model dropdown is empty, or shows _no-model-registered

The app fills the dropdown from the gateway. Either the gateway has no models or the app cannot reach it.

curl -s http://localhost:4000/health/liveliness                              # "I'm alive!"
curl -s -H "Authorization: Bearer $LITELLM_MASTER_KEY" http://localhost:4000/v1/models | jq '.data | length'
docker compose exec librechat wget -qO- http://litellm-proxy:4000/health/liveliness   # from inside the network

Zero models: register one (Add or change a model) and rebuild the gateway image. bad address from the third command: the Docker network is gone; docker compose down && docker compose up -d.

A model was registered and does not appear

add-model.sh writes litellm/config.yaml, but the gateway runs the copy baked into its image. Until you rebuild, the script's "is now registered" is not true:

docker compose build litellm-proxy && docker compose up -d litellm-proxy

no matching manifest for linux/arm64/v8

The published images are amd64-only. On an arm64 host add the platform override; see Run the stack locally.

denied: denied when pulling

Not signed in to the registry, or the token expired. GHCR images.

Everything runs, but with the example secrets

grep -c replace-me .env

Anything but 0 on a server means the databases were initialised with placeholder passwords and the gateway's master key is sk-replace-me. Nothing in the stack refuses to run like this. Fixing it after the fact is not a matter of editing .env: Postgres and MongoDB keep the password they were created with. On a fresh install, docker compose down -v, fill the secrets (bootstrap.sh does it), bring it up again. On an install with data, change the passwords inside each database first, then in .env.

nufi-scanner stays unhealthy for minutes

First start downloads a 700 MB model. The health check allows five minutes; docker compose logs -f nufi-scanner shows the download.

Signed in to the app, the console says /unauthorized

The console needs the chat's session cookie. On one origin (localhost, or one hostname with path routing) sign in to the app first and reload. On subdomains, all of these, in order:

  1. the chat service has COOKIE_DOMAIN=.example.com and COOKIE_SAMESITE=lax; on the compose stack that means two lines in the librechat service's environment: (the wrapper stack in deploy/railway already passes them from .env)
  2. the reverse proxy forwards Set-Cookie unchanged
  3. both hosts are HTTPS
  4. JWT_REFRESH_SECRET is identical in the chat and console environments

See SSO and reverse proxy.

Entering Studio or Works fails, or creates a user called external-<hash>

The console could not ask the chat who the member is. CHAT_BASE_URL on the console must reach the chat the member is signed in to; Works then refuses with email_is_missing, Studio provisions a nameless account. See Single sign-on for the agent apps.

network npuops_npuops not found

The wrapper stack in deploy/railway is in shared-network mode and the platform stack is not running, or its network has another name. docker network ls, then fix SHARED_DOCKER_NETWORK in the wrapper's .env, or remove its docker-compose.override.yml to leave shared mode.

Langfuse stops showing traces

docker compose logs --tail 200 langfuse-worker

Usually ClickHouse or MinIO out of disk, or the gateway's LANGFUSE_* variables changed. The smoke test's step 6 checks this path end to end.

Memory pressure, OOM kills

ClickHouse, the scanner, the two Presidio containers and the Langfuse worker are the heavy ones. The compose file declares no limits; add deploy.resources.limits.memory to the service you want contained, and lower ClickHouse's max_memory_usage if you cap it.

Disk filling

ClickHouse and MinIO grow with every trace, roughly 50 MB per thousand. Backup and restore for retention, Infra sizing for when to move them off the host.

The admin panel keeps signing you out

Its own idle timeout is 30 minutes (ADMIN_SESSION_IDLE_TIMEOUT_MS), and it revalidates the session against the app every minute; if the app is briefly unhealthy the panel signs you out. docker compose logs librechat.

"Conversation not found" after an upgrade

A schema change the running version cannot read. Roll the image back and restore the pre-upgrade MongoDB dump: Upgrade the NuFi app.

The console cannot provision a user

docker compose logs --tail 100 console | grep -i provision

The console's LITELLM_MASTER_KEY does not match the gateway's, or the gateway is down. Provisioning is retried on the next visit once it is back.

The end-to-end smoke test fails at step 3

./scripts/e2e-smoke-test.sh calls /api/ask/custom, a route of an older chat version; the current app answers 404 there, so the test cannot pass until it is updated. ./scripts/smoke-test.sh (gateway, trace, metrics, guardrail) still passes, and the four manual checks in Upgrade the NuFi app cover the chat leg.