Troubleshooting
The failures operators hit on this stack, each with the command that tells you which one you have.
Three commands cover most incidents. Run them from deploy/platform:
docker compose ps # anything not (healthy)?
docker compose logs --tail 200 <service> # the last lines of the one that is not
curl -s http://localhost:9090/api/v1/alerts | jq # what Prometheus thinks is wrongThe gateway restarts in a loop
docker compose logs --tail 50 litellm-proxy| The log ends with | Cause | Do |
|---|---|---|
unknown threshold key(s) [...] | the locally built gateway image is older than the checkout; docker compose up never rebuilds it | docker compose build litellm-proxy nufi-scanner && docker compose up -d |
a Gemini or GEMINI_API_KEY error | litellm/config.yaml ships two Gemini entries and LiteLLM refuses to start with an empty key | set the key in .env, or remove both entries and rebuild |
| a model entry named in the error | a typo in litellm/config.yaml, or a ${VAR} it references is unset | fix the entry, rebuild |
| connection refused to Postgres | Postgres was still starting | wait 30 seconds; if it persists, docker compose logs postgres |
The model dropdown is empty, or shows _no-model-registered
The app fills the dropdown from the gateway. Either the gateway has no models or the app cannot reach it.
curl -s http://localhost:4000/health/liveliness # "I'm alive!"
curl -s -H "Authorization: Bearer $LITELLM_MASTER_KEY" http://localhost:4000/v1/models | jq '.data | length'
docker compose exec librechat wget -qO- http://litellm-proxy:4000/health/liveliness # from inside the networkZero models: register one (Add or change a model)
and rebuild the gateway image. bad address from the third command: the
Docker network is gone; docker compose down && docker compose up -d.
A model was registered and does not appear
add-model.sh writes litellm/config.yaml, but the gateway runs the copy
baked into its image. Until you rebuild, the script's "is now registered"
is not true:
docker compose build litellm-proxy && docker compose up -d litellm-proxyno matching manifest for linux/arm64/v8
The published images are amd64-only. On an arm64 host add the platform override; see Run the stack locally.
denied: denied when pulling
Not signed in to the registry, or the token expired. GHCR images.
Everything runs, but with the example secrets
grep -c replace-me .envAnything but 0 on a server means the databases were initialised with
placeholder passwords and the gateway's master key is sk-replace-me.
Nothing in the stack refuses to run like this. Fixing it after the fact is
not a matter of editing .env: Postgres and MongoDB keep the password they
were created with. On a fresh install, docker compose down -v, fill the
secrets (bootstrap.sh does it), bring it up again. On an install with
data, change the passwords inside each database first, then in .env.
nufi-scanner stays unhealthy for minutes
First start downloads a 700 MB model. The health check allows five
minutes; docker compose logs -f nufi-scanner shows the download.
Signed in to the app, the console says /unauthorized
The console needs the chat's session cookie. On one origin (localhost,
or one hostname with path routing) sign in to the app first and reload. On
subdomains, all of these, in order:
- the chat service has
COOKIE_DOMAIN=.example.comandCOOKIE_SAMESITE=lax; on the compose stack that means two lines in thelibrechatservice'senvironment:(the wrapper stack indeploy/railwayalready passes them from.env) - the reverse proxy forwards
Set-Cookieunchanged - both hosts are HTTPS
JWT_REFRESH_SECRETis identical in the chat and console environments
Entering Studio or Works fails, or creates a user called external-<hash>
The console could not ask the chat who the member is. CHAT_BASE_URL on
the console must reach the chat the member is signed in to; Works then
refuses with email_is_missing, Studio provisions a nameless account. See
Single sign-on for the agent apps.
network npuops_npuops not found
The wrapper stack in deploy/railway is in shared-network mode and the
platform stack is not running, or its network has another name.
docker network ls, then fix SHARED_DOCKER_NETWORK in the wrapper's
.env, or remove its docker-compose.override.yml to leave shared mode.
Langfuse stops showing traces
docker compose logs --tail 200 langfuse-workerUsually ClickHouse or MinIO out of disk, or the gateway's LANGFUSE_*
variables changed. The smoke test's step 6 checks this path end to end.
Memory pressure, OOM kills
ClickHouse, the scanner, the two Presidio containers and the Langfuse
worker are the heavy ones. The compose file declares no limits; add
deploy.resources.limits.memory to the service you want contained, and
lower ClickHouse's max_memory_usage if you cap it.
Disk filling
ClickHouse and MinIO grow with every trace, roughly 50 MB per thousand. Backup and restore for retention, Infra sizing for when to move them off the host.
The admin panel keeps signing you out
Its own idle timeout is 30 minutes (ADMIN_SESSION_IDLE_TIMEOUT_MS), and
it revalidates the session against the app every minute; if the app is
briefly unhealthy the panel signs you out. docker compose logs librechat.
"Conversation not found" after an upgrade
A schema change the running version cannot read. Roll the image back and restore the pre-upgrade MongoDB dump: Upgrade the NuFi app.
The console cannot provision a user
docker compose logs --tail 100 console | grep -i provisionThe console's LITELLM_MASTER_KEY does not match the gateway's, or the
gateway is down. Provisioning is retried on the next visit once it is back.
The end-to-end smoke test fails at step 3
./scripts/e2e-smoke-test.sh calls /api/ask/custom, a route of an older
chat version; the current app answers 404 there, so the test cannot pass
until it is updated. ./scripts/smoke-test.sh (gateway, trace, metrics,
guardrail) still passes, and the four manual checks in
Upgrade the NuFi app
cover the chat leg.