Infra sizing
The reference host for a single-server pilot, what grows, and when to split.
Reference host
The whole compose stack on one machine, sized for a pilot of ten to fifty people:
| Resource | Sizing | Why |
|---|---|---|
| CPU | 12 vCPU | ClickHouse, the injection scanner and the two Presidio sidecars are the busy ones |
| RAM | 32 GB | ClickHouse and the guardrail models dominate |
| Disk | 256 GB SSD | traces grow fastest; alert at 60 % |
| OS | Ubuntu 22.04 LTS | anything with a current Docker works |
This carries roughly fifty active users at a hundred prompts a day each.
The internal record for NuFi's own host is
deploy/platform/docs/deployment-infra.md.
Per-service caps
Guidance, not configuration: the compose file declares no deploy.resources
limits, so nothing enforces these until you add them.
| Service | CPU | Memory |
|---|---|---|
| ClickHouse | 2 | 4 GB |
nufi-scanner | 1 | 2 GB |
| Presidio (each) | 0.5 | 1 GB |
| the gateway | 1 | 2 GB |
| the NuFi app | 1 | 2 GB |
| Langfuse web and worker | 1 each | 2 GB each |
| Postgres, MongoDB | 1 each | 2 GB each |
| the console | 0.5 | 512 MB |
| Prometheus | 1 | 2 GB |
Add them under a service's deploy.resources.limits when one process must
not be allowed to take the host down. If you cap ClickHouse, lower its
max_memory_usage to match.
Disk
| Volume | Start | Growth |
|---|---|---|
npuops_postgres-data | 10 GB | slow |
npuops_mongodb-data | 20 GB | with conversations |
npuops_clickhouse-data | 50 GB | fastest: about 50 MB per thousand traces |
npuops_minio-data | 30 GB | with ClickHouse |
npuops_prometheus-data | 20 GB | bounded by the 15-day retention |
| images, logs, headroom | 50 GB |
Alert at 60 %. Expand to 500 GB when ClickHouse and MinIO together pass 100 GB; to 1 TB past about a hundred users or a million traces a month.
When to split
- ClickHouse and MinIO past 100 GB: move both to a host with a bigger disk; they only talk to Langfuse.
- More than a hundred concurrent users: a second app instance behind a load balancer, with the app's Redis-backed resumable streams turned on.
- Gateway CPU saturated: more gateway containers behind a load
balancer, not more workers in one. The gateway runs with
--num_workers 1on purpose: its guardrail metrics live in one process, and a second worker would report a second, disagreeing set. Postgres and Redis are shared; the guardrail sidecars scale with the gateway. - Trace ingestion lagging: more Langfuse workers.
Network
Outbound to the model providers you use; inbound only 80 and 443 for the
reverse proxy, or nothing at all with a
Cloudflare tunnel. Everything else
stays on the npuops Docker network.
Licences to know about
Three components carry licences that constrain redistribution, not internal use:
| Component | Licence | Constraint |
|---|---|---|
| MongoDB | SSPL | cannot be offered as a hosted database service |
| MinIO | AGPLv3 | copyleft if you redistribute it modified |
| Redis 7.4 | RSALv2 / SSPL | cannot be offered as a hosted service |
Running the stack for your own organisation is fine under all three. Distributing NuFi to others would mean swapping in FerretDB, SeaweedFS and Valkey, which speak the same protocols.
Cost
The reference host at cloud prices is about $120 to $180 a month, plus a few dollars of egress and backups. The tunnel and Access are free at this size. Model inference is on top and depends entirely on the backend.