Docker Compose
The compose stack in deploy/platform, service by service, with the volumes, ports and the one rule about rebuilding.
The platform is one compose file, deploy/platform/docker-compose.yml,
project name npuops, one bridge network of the same name. Services reach
each other by name (http://litellm-proxy:4000); only the ones a person
opens have host ports.
Services
postgres ─┬─► langfuse-worker ─┐
├─► langfuse-web ────┼─► litellm-proxy ─► librechat (the NuFi app)
└─► litellm-proxy │ └─► console
│
clickhouse ──► langfuse-worker │
redis ───────► litellm-proxy │
minio ──┬─► minio-init ────────┘
└─► langfuse-{worker,web}
presidio-analyzer ─┐
presidio-anonymizer ┼─► litellm-proxy (guardrails)
nufi-scanner ───────┘
prometheus → grafana
prometheus → alertmanager
postgres-exporter, redis-exporter → prometheus| Service | Image | Note |
|---|---|---|
librechat | ghcr.io/dudaji-vn/nufichat:main | the NuFi app; librechat.yaml is mounted from the platform root |
console | ghcr.io/dudaji-vn/nufi-console:${NUFI_CONSOLE_TAG:-main} | |
litellm-proxy | built from litellm/, nufi/litellm:local | config, callbacks and guardrail code are baked in; litellm/guardrails/policy.yaml is mounted |
nufi-scanner | built from scanner/, nufi/scanner:local | the injection classifier; 700 MB model on first start |
presidio-analyzer, presidio-anonymizer | mcr.microsoft.com/presidio-*:2.2.362 | PII detection for the guardrails |
langfuse-web, langfuse-worker | langfuse/langfuse:3, langfuse-worker:3 | |
postgres, redis, clickhouse, minio, minio-init | pinned images | minio-init creates the bucket and exits |
mongodb | mongo:4.4 | the last version that runs without AVX, for the target VM |
prometheus, grafana, alertmanager, postgres-exporter, redis-exporter | pinned images | |
e2e-test | built from scripts/e2e/, profile e2e | never starts by default; see the note below |
The admin panel is not in this file. It runs from its own image
(ghcr.io/dudaji-vn/nufichat-admin-panel) wherever you put it, pointed at
the app's URL; NuFi's own copy runs on Railway.
Host ports
| Host port | Service | In production |
|---|---|---|
| 3080 | librechat | behind the reverse proxy |
| 3001 | console | behind the reverse proxy |
| 3000 | langfuse-web | behind the reverse proxy, or admin-only |
| 3030 | grafana | behind the reverse proxy, or admin-only |
| 4000 | litellm-proxy | behind the reverse proxy, as the API host |
| 9090 | prometheus | admin-only: firewall it, reach it over SSH |
| 9093 | alertmanager | admin-only |
The databases, the scanner, the Presidio sidecars and the exporters have no host ports.
Volumes
| Volume | Holds | Back up |
|---|---|---|
npuops_postgres-data | gateway keys, budgets, spend; Langfuse metadata | yes |
npuops_mongodb-data | the app's users and conversations | yes |
npuops_clickhouse-data | Langfuse traces; the largest | yes, weekly |
npuops_minio-data | Langfuse payloads | yes, weekly |
npuops_redis-data | rate-limit counters, cache | no |
npuops_prometheus-data | 15 days of metrics | no |
npuops_grafana-data | Grafana state; the dashboard is provisioned from git | no |
Backup and restore has the commands.
Start order and health
The databases, the gateway, the app, the console, the guardrail sidecars
and the monitoring services have health checks (the Langfuse pair and the
two exporters do not), and depends_on uses service_healthy, so the
gateway waits for Postgres, Redis and all three guardrail sidecars, and the
app waits for MongoDB and the gateway. When docker compose up -d seems stuck, docker compose ps
names the service that is not healthy; docker compose logs -f <service>
says why. The scanner is the slow one on first start.
The one rule: what needs a rebuild
Two services are built on the host, and docker compose up -d never
rebuilds them. After a git pull that touches litellm/ or scanner/,
and after add-model.sh:
docker compose build litellm-proxy nufi-scanner
docker compose up -dWhat a change needs:
| You changed | Then |
|---|---|
litellm/guardrails/policy.yaml | docker compose restart litellm-proxy (it is mounted) |
litellm/config.yaml, litellm/callbacks/*, litellm/guardrails/*.py | rebuild, as above |
librechat.yaml | docker compose restart librechat |
.env | docker compose up -d recreates the services whose environment changed |
| an image tag | docker compose pull <service> && docker compose up -d <service> |
monitoring/rules/*.yml | curl -X POST http://localhost:9090/-/reload |
Day to day
docker compose ps
docker compose logs -f litellm-proxy
docker compose restart librechat
docker compose down # stop, keep the volumes
docker compose down -v # stop and delete every volumeThe e2e profile
./scripts/e2e-smoke-test.sh runs the e2e-test service, which registers
a user in the app and sends a message through the gateway. It targets a
chat route the current app no longer serves and fails at its third step;
until it is updated, ./scripts/smoke-test.sh plus the manual checks in
Upgrade the NuFi app
are the working verification.
Layout
deploy/platform/
├── docker-compose.yml
├── .env # created by bootstrap.sh from .env.example; gitignored
├── librechat.yaml # the app's configuration, mounted into librechat
├── litellm/ # the gateway image: config.yaml, Dockerfile, guardrails/, callbacks/
├── scanner/ # the injection classifier image
├── monitoring/ # prometheus.yml, alertmanager.yml, rules/, grafana/
└── scripts/ # bootstrap.sh, add-model.sh, smoke-test.sh, e2e/The wrapper stack for the app alone, with its own MongoDB and a RAG
service, is deploy/railway/; see FAQ.