NuFiDocs

Infra sizing

The reference host for a single-server pilot, what grows, and when to split.

Reference host

The whole compose stack on one machine, sized for a pilot of ten to fifty people:

ResourceSizingWhy
CPU12 vCPUClickHouse, the injection scanner and the two Presidio sidecars are the busy ones
RAM32 GBClickHouse and the guardrail models dominate
Disk256 GB SSDtraces grow fastest; alert at 60 %
OSUbuntu 22.04 LTSanything with a current Docker works

This carries roughly fifty active users at a hundred prompts a day each. The internal record for NuFi's own host is deploy/platform/docs/deployment-infra.md.

Per-service caps

Guidance, not configuration: the compose file declares no deploy.resources limits, so nothing enforces these until you add them.

ServiceCPUMemory
ClickHouse24 GB
nufi-scanner12 GB
Presidio (each)0.51 GB
the gateway12 GB
the NuFi app12 GB
Langfuse web and worker1 each2 GB each
Postgres, MongoDB1 each2 GB each
the console0.5512 MB
Prometheus12 GB

Add them under a service's deploy.resources.limits when one process must not be allowed to take the host down. If you cap ClickHouse, lower its max_memory_usage to match.

Disk

VolumeStartGrowth
npuops_postgres-data10 GBslow
npuops_mongodb-data20 GBwith conversations
npuops_clickhouse-data50 GBfastest: about 50 MB per thousand traces
npuops_minio-data30 GBwith ClickHouse
npuops_prometheus-data20 GBbounded by the 15-day retention
images, logs, headroom50 GB

Alert at 60 %. Expand to 500 GB when ClickHouse and MinIO together pass 100 GB; to 1 TB past about a hundred users or a million traces a month.

When to split

  • ClickHouse and MinIO past 100 GB: move both to a host with a bigger disk; they only talk to Langfuse.
  • More than a hundred concurrent users: a second app instance behind a load balancer, with the app's Redis-backed resumable streams turned on.
  • Gateway CPU saturated: more gateway containers behind a load balancer, not more workers in one. The gateway runs with --num_workers 1 on purpose: its guardrail metrics live in one process, and a second worker would report a second, disagreeing set. Postgres and Redis are shared; the guardrail sidecars scale with the gateway.
  • Trace ingestion lagging: more Langfuse workers.

Network

Outbound to the model providers you use; inbound only 80 and 443 for the reverse proxy, or nothing at all with a Cloudflare tunnel. Everything else stays on the npuops Docker network.

Licences to know about

Three components carry licences that constrain redistribution, not internal use:

ComponentLicenceConstraint
MongoDBSSPLcannot be offered as a hosted database service
MinIOAGPLv3copyleft if you redistribute it modified
Redis 7.4RSALv2 / SSPLcannot be offered as a hosted service

Running the stack for your own organisation is fine under all three. Distributing NuFi to others would mean swapping in FerretDB, SeaweedFS and Valkey, which speak the same protocols.

Cost

The reference host at cloud prices is about $120 to $180 a month, plus a few dollars of egress and backups. The tunnel and Access are free at this size. Model inference is on top and depends entirely on the backend.