NuFiDocs

Backup and restore

What is worth backing up, the exact commands, how to restore, and the drill that proves the backups are real.

The state that matters lives in four volumes. Everything else (metrics, cache, provisioned dashboards) is rebuilt on start and not worth keeping.

What to back up

VolumeWhat is in itFrequency
npuops_postgres-datatwo databases: npuops (gateway keys, budgets, spend) and langfuse (trace metadata)nightly
npuops_mongodb-datathe NuFi app: users, conversations, agents, files metadatanightly
npuops_clickhouse-dataLangfuse tracesweekly
npuops_minio-dataLangfuse blob payloadsweekly

Postgres and MongoDB are the business-critical stores; without them you lose accounts, keys and conversation history. ClickHouse and MinIO are forensic; losing them costs observability history, not operation.

The commands below run from deploy/platform. Credentials come from the containers' own environment, so nothing has to be pasted on the host.

Postgres

docker compose exec -T postgres sh -c 'pg_dumpall -U "$POSTGRES_USER"' \
  | gzip > "/backups/pg_$(date +%F).sql.gz"

Restore into a fresh Postgres (the dump recreates both databases):

gunzip < /backups/pg_2026-09-07.sql.gz \
  | docker compose exec -T postgres sh -c 'psql -U "$POSTGRES_USER"'

MongoDB

The stack's MongoDB has a root user, so mongodump must authenticate:

docker compose exec -T mongodb sh -c \
  'mongodump --archive --gzip -u "$MONGO_INITDB_ROOT_USERNAME" -p "$MONGO_INITDB_ROOT_PASSWORD" --authenticationDatabase admin' \
  > "/backups/mongo_$(date +%F).gz"

Without the three auth flags it stops at command listDatabases requires authentication. Restore:

docker compose exec -T mongodb sh -c \
  'mongorestore --archive --gzip --drop -u "$MONGO_INITDB_ROOT_USERNAME" -p "$MONGO_INITDB_ROOT_PASSWORD" --authenticationDatabase admin' \
  < /backups/mongo_2026-09-07.gz

--drop replaces every collection with the archive's copy. Use it on a fresh target, or when you mean it.

ClickHouse and MinIO

Too large for logical dumps. Snapshot the volumes with the services stopped, so the files are consistent:

docker compose stop clickhouse minio
docker run --rm -v npuops_clickhouse-data:/data -v "$PWD/backups:/out" alpine \
  tar czf "/out/clickhouse_$(date +%F).tar.gz" /data
docker run --rm -v npuops_minio-data:/data -v "$PWD/backups:/out" alpine \
  tar czf "/out/minio_$(date +%F).tar.gz" /data
docker compose start clickhouse minio

Restore by extracting into the volume while the service is stopped, the same way in reverse.

Schedule it and ship it

Put the dumps in a cron job or a systemd timer, keep 30 days, and end the script with a copy to somewhere that is not the same disk:

rsync -av /backups/ backup-host:/srv/nufi/      # another machine
aws s3 sync /backups s3://nufi-backups/nufi/    # object storage
restic backup /backups                           # deduplicated, encrypted

Restore drill

Once a quarter:

  1. Bring up a second stack on another host with bootstrap.sh --backend skip.
  2. Stop it. Restore the Postgres and MongoDB dumps into it; extract the two volume snapshots.
  3. Start it. Sign in, open a recent conversation, send a message, generate a key in the console.

Untested backups are not backups. The drill is where you find out that a schema moved, or that the ClickHouse snapshot was taken while it was writing.

What you do not back up

  • Prometheus data (npuops_prometheus-data): 15 days of metrics, cheap to lose.
  • Redis (npuops_redis-data): rate-limit counters and cache, reset on restart.
  • Grafana (npuops_grafana-data): the dashboard is provisioned from monitoring/grafana/dashboards/ in git; a fresh container has it.
  • The MinIO bucket itself: minio-init recreates it on every start.
  • .env: not a volume, and not in git. Keep a copy in your secrets manager; without it a restored stack cannot decrypt anything.

If the Postgres dump keeps growing

Langfuse keeps metadata in Postgres and the traces themselves in ClickHouse, so Postgres stays small, a few gigabytes at moderate use. A dump past ten gigabytes almost always means the gateway's spend logs are not being pruned; LiteLLM has a retention setting for them.